AI Potluck
Model components / Fine-tuning code

OpenAI Fine-Tuning API

OpenAI

Managed fine-tuning for GPT-4o, GPT-4o-mini, GPT-4.1, and o-series via the OpenAI API; supports supervised fine-tuning, DPO (added late 2024), and vision fine-tuning. Picked when the use case demands GPT-4-tier quality at lower latency/cost on a narrow task; the only way to fine-tune OpenAI's frontier weights. Training priced at ~$3/M tokens for GPT-4o-mini and GPT-4o (GPT-4.1 family ~$0.80-3/M); models served at standard inference rates.

OpenAI model-optimization / fine-tuning API. Methods: SFT, vision fine-tuning, DPO, reinforcement fine-tuning (RFT, reasoning models only). Tunable models incl. gpt-4.1 / -mini / -nano (SFT,DPO), gpt-4o (vision), o4-mini (RFT). Proprietary managed service. FRESHNESS: OpenAI is winding down this fine-tuning platform (per primary deprecations page): new/never-before users blocked from creating jobs as of May 7 2026; orgs inactive 60d lose job creation Jul 2 2026; ALL customers lose ability to create new fine-tuning jobs Jan 6 2027. Inference on existing fine-tuned models persists until base-model deprecation. Product still exists and functions for existing users at scoring (June 2026) but is on a sunset path.

Openness

1 high confidence
1.0
source
closed
training-code
closed
runs-on
OpenAI-platform-only
license
Proprietary(paid API service)

Proprietary managed fine-tuning service, no source, runs only on OpenAI's platform against OpenAI's closed models.

Adoption

not assessed

No disclosed standalone fine-tuning-API usage figure located this run (developer count / jobs run / tokens). The OpenAI platform broadly serves ~1B+ ChatGPT WAU, but that surface is not the fine-tuning API and cannot be honestly attributed to this SKU. Declining to assign a level rather than borrow the umbrella surface number.

Capability

4 medium confidence
4.0

Strong managed method coverage (SFT+DPO+RFT) over frontier base models, but black-box: no parallelism/precision/scale knobs exposed, no MLPerf-Training basis. Scored 4 on method breadth + frontier base quality; not 5 since the user cannot control training scale/internals.

Unchanged since 2026-06-24 (last edited, not re-checked)