How was Chatgpt fine-tuned?

    Asked "How was Chatgpt fine-tuned?", Gemini, Google AI Mode and Perplexity named 9 distinct names across 5 answers on August 22, 2026, and 2 of them in two or more answers, and all 3 engines agreed on OpenAI, the only name every engine named.

    Measured acrossGeminiGoogle AI ModePerplexity

    2 of 9 names confirmed · named in 2 or more of 5 answers · asked August 22, 2026 · 3 engines

    ChatGPT was fine-tuned using reinforcement learning from human feedback (RLHF), which involves a multi-stage process including data preparation, supervised fine-tuning, and reinforcement learning to align outputs with human preferences. The fine-tuning process relies on three core steps: supervised fine-tuning, reward model training, and reinforcement learning via proximal policy optimization.

    • 1OpenAInamed in 3 of 5 answersGemini logo
    • 2Proximal Policy Optimizationnamed in 2 of 5 answersGemini logo
    • 3Relinns Technologiesnamed in 1 of 5 answersone answerGemini logo
    • 4YouTubenamed in 1 of 5 answersone answerGemini logo
    • 5Facebooknamed in 1 of 5 answersone answerGemini logo
    • 6Mediumnamed in 1 of 5 answersone answerGemini logo
    • 7IntuitionLabsnamed in 1 of 5 answersone answerGemini logo
    • 8Didanamed in 1 of 5 answersone answerGemini logo
    • 9Unslothnamed in 1 of 5 answersone answerGemini logo
    LOCKED

    The full measurement

    • The position each of the 3 engines gave all 9 names.
    • How many of the 5 answers named each of them.
    • 12 sampled observations behind this ranking, and where the engines disagree.
    • Fan-out — the query each engine actually searched.
    • Every citation, and the sources nobody cited.
    UNLOCK