How was Chatgpt fine-tuned?
Asked "How was Chatgpt fine-tuned?", Gemini, Google AI Mode and Perplexity named 9 distinct names across 5 answers on August 22, 2026, and 2 of them in two or more answers, and all 3 engines agreed on OpenAI, the only name every engine named.
2 of 9 names confirmed · named in 2 or more of 5 answers · asked August 22, 2026 · 3 engines
ChatGPT was fine-tuned using reinforcement learning from human feedback (RLHF), which involves a multi-stage process including data preparation, supervised fine-tuning, and reinforcement learning to align outputs with human preferences. The fine-tuning process relies on three core steps: supervised fine-tuning, reward model training, and reinforcement learning via proximal policy optimization.
- 1OpenAInamed in 3 of 5 answers
- 2Proximal Policy Optimizationnamed in 2 of 5 answers
- 3Relinns Technologiesnamed in 1 of 5 answersone answer
- 4YouTubenamed in 1 of 5 answersone answer
- 5Facebooknamed in 1 of 5 answersone answer
- 6Mediumnamed in 1 of 5 answersone answer
- 7IntuitionLabsnamed in 1 of 5 answersone answer
- 8Didanamed in 1 of 5 answersone answer
- 9Unslothnamed in 1 of 5 answersone answer
The full measurement
- The position each of the 3 engines gave all 9 names.
- How many of the 5 answers named each of them.
- 12 sampled observations behind this ranking, and where the engines disagree.
- Fan-out — the query each engine actually searched.
- Every citation, and the sources nobody cited.