Who should I follow to learn about Model Reliability in machine learning?

    Updated September 13, 2026
    CONSENSUS
    Percy Liang

    Answer summary

    To learn about model reliability, follow experts such as Shiori Sagawa, Aleksander Madry, Percy Liang, Chip Huyen, and D. Sculley, who are leading voices in robust machine learning, adversarial robustness, and MLOps.

    Your presence

    Sign in to check whether you are mentioned or cited in this prompt.

    Where each engine placed them

    #BrandChatGPTGemini logoGeminiAI ModeCopilot logoCopilotMention rateBrand Share of VoiceCiteRankCitabilitySentiment
    1PLPercy Liang - 8% -
    ChatGPTGeminiAI ModeCopilot
    Mention rate - Share of Voice 8%CiteRank - Citability Sentiment
    2CHChip Huyen - 8% -
    ChatGPTGeminiAI ModeCopilot
    Mention rate - Share of Voice 8%CiteRank - Citability Sentiment

    What the models searched first

    The sub-queries each engine actually searched before answering.

    Unlock the full picture

    The sub-queries each engine actually searched before answering.

    Your tracked entities

    Save a voice, brand or tool and see where it stands on this question.

    Full responses

    4 of 4 engines that answered. Full stored text as captured, with dashes shown as hyphens. Nothing is summarised.

    ChatGPT4,197 characters · September 13, 2026

    If you mean **AI/ML model reliability** (making models dependable, measurable, robust, and safe), I’d follow a mix of researchers, labs, and practitioners across a few subfields: evaluation, robustness, interpretability, uncertainty, and AI safety. ## Core researchers to follow - Dawn Song - trustworthy AI, security, robustness, and evaluation. - Percy Liang - foundation model evaluation, transparency, and the Stanford Center for Research on Foundation Models ecosystem. - Arvind Narayanan - empirical evaluation of AI systems, reliability, and responsible deployment. His group has worked on reliability benchmarks such as HAL for AI agents. [HAL Reliability Dashboard](https://hal.cs.princeton.edu/reliability/?utm_source=chatgpt.com) - Soheil Feizi - one of the researchers most directly focused on reliable AI: robustness, interpretability, hallucination detection, provenance, and dependable foundation models. [Soheil Feizi - Reliable AI Lab, University of Maryland](https://www.cs.umd.edu/~sfeizi/?utm_source=chatgpt.com) - Yarin Gal - uncertainty estimation and knowing when models are likely to fail. - Been Kim - interpretability and methods for understanding model behavior. - Chris Olah - mechanistic interpretability and understanding neural networks. - Dan Hendrycks - AI safety evaluations, robustness, and benchmarks. - Stuart Russell - foundational work on reliable and aligned AI systems. - Beth Barnes - practical evaluations of frontier model capabilities and risks through METR. [An ex-OpenAI researcher's influential nonprofit is now smack in the middle of AI's doom debate](https://www.businessinsider.com/metr-beth-barnes-ai-talent-shortage-safety-research-openai-2026-8?utm_source=chatgpt.com) ## Labs and organizations worth tracking - METR - frontier model evaluations, especially around whether models can be trusted in difficult tasks. [An ex-OpenAI researcher's influential nonprofit is now smack in the middle of AI's doom debate](https://www.businessinsider.com/metr-beth-barnes-ai-talent-shortage-safety-research-openai-2026-8?utm_source=chatgpt.com) - Stanford Center for Research on Foundation Models - foundation model transparency, benchmarks, and evaluation. - Google DeepMind - robustness, interpretability, alignment, and evaluation research. - Anthropic - model behavior research, interpretability, and safety evaluations. - MLCommons - practical ML benchmarks and reliability measurement. - Partnership on AI - responsible AI practices and deployment standards. ## If you care specifically about LLM reliability Follow people working on: - **Evals and benchmarks** - Ethan Perez - John Schulman - Jekaterina Novikova - LLM evaluation, consistency, multilingual evaluation, and trustworthy AI. [Jekaterina Novikova](https://jeknov.github.io/?utm_source=chatgpt.com) - **Agents and long-horizon reliability** - Stephan Rabanser - uncertainty, robustness, and reliable AI agent evaluation. [Stephan Rabanser -Towards a Science of AI Agent Reliability | Center for Information Technology Policy](https://citp.princeton.edu/events/2026/stephan-rabanser-towards-science-ai-agent-reliability?utm_source=chatgpt.com) - Sayash Kapoor - AI evaluation and reliability measurement. [HAL Reliability Dashboard](https://hal.cs.princeton.edu/reliability/?utm_source=chatgpt.com) - **Interpretability** - Neel Nanda - Lilian Weng ## Papers/topics to search for A good reading path: 1. **Calibration & uncertainty** - “calibration in neural networks” - “selective prediction” - “out-of-distribution detection” 2. **Robustness** - adversarial robustness - distribution shift - stress testing 3. **LLM evaluation** - HELM - BIG-bench - TruthfulQA - SWE-bench - agent reliability benchmarks 4. **Interpretability** - mechanistic interpretability - activation patching - sparse autoencoders A concise “starter follow list” would be: **Percy Liang, Soheil Feizi, Arvind Narayanan, Dan Hendrycks, Chris Olah, Neel Nanda, Beth Barnes, and Yarin Gal**. That group covers most of the modern reliability stack: *measure failures → understand failures → reduce failures → know when not to trust the model*.

    Copilot3,792 characters · September 13, 2026

    **Follow these experts and institutions to learn about model reliability: MIT’s Navid Azizan, Stanford’s Mykel J. Kochenderfer, Polytechnique Montréal’s Thibaut Vidal, CMU’s Grace Lewis and Ipek Ozkaya, and University of Tokyo’s Yuheng Huang. They are leading voices in reliability, trustworthy AI, and quality assurance for machine learning systems.** ## 🔑 Key Figures in Model Reliability Expert/Institution| Focus Area| Why Follow Them ---|---|--- **Navid Azizan (MIT)**| Reliability of foundation models; uncertainty quantification| Developed methods to assess reliability of large AI models before deployment, especially in safety-critical contexts.[How to assess a general-purpose AI model’s reliability before it’s deployed | MIT News | Massachusetts Institute of Technology](https://news.mit.edu/2024/how-assess-general-purpose-ai-models-reliability-its-deployed) **Mykel J. Kochenderfer (Stanford)**| Reliability in high-stakes ML (aviation, healthcare, defense)| Co-authored frameworks for holistic reliability assessment, including calibration, adversarial robustness, and distribution-shift resilience.[arXiv:2307.10586v2 [cs.LG] 29 Jul 2023](https://arxiv.org/pdf/2307.10586) **Thibaut Vidal (Polytechnique Montréal)**| Trustworthy ML via combinatorial optimization| Explores optimization-based approaches for robustness, fairness, and certification in ML systems.[Trustworthy Machine Learning through the Lens of Combinatorial Optimization: Survey and Research Perspectives](https://arxiv.org/html/2607.07762v1) **Grace Lewis & Ipek Ozkaya (CMU SEI)**| Quality models for ML components| Work on structured frameworks (ISO 25059) for assessing ML quality beyond accuracy, including robustness and integration.[A Quality Model for Machine Learning Components | CMU Software Engineering Institute](https://www.sei.cmu.edu/blog/a-quality-model-for-machine-learning-components/) **Yuheng Huang (University of Tokyo)**| Systemic assurance for AI-enabled systems| Researches runtime monitoring and defenses to enhance reliability of large language models and complex AI systems.[Towards Trustworthy AI-Enabled Systems: From Model-Level Reliability to Systemic Assurance](https://dl.acm.org/doi/epdf/10.1145/3803437.3804867) ## 📌 Why These Voices Matter * **Practical impact**: Their work spans both theoretical frameworks and applied reliability testing in domains like healthcare, transportation, and defense. * **Standards & frameworks**: CMU SEI and ISO initiatives are shaping how reliability is formally measured in AI systems. * **Cutting-edge methods**: MIT and Stanford researchers are pioneering techniques for uncertainty quantification and robustness evaluation. ## ⚠️ Risks & Challenges in Following the Field * **Fragmentation** : Reliability research is spread across multiple disciplines (optimization, software engineering, safety-critical systems). Following diverse voices ensures balance. * **Rapid evolution**: Techniques for reliability (e.g., calibration, adversarial robustness) evolve quickly; staying updated with recent publications is essential. * **Scalability issues**: Many reliability methods work in theory but struggle with large-scale deployment. Researchers like Vidal and Huang are tackling this challenge. ## ✅ Next Steps for You * Follow these experts on **Google Scholar, arXiv, and institutional blogs** (MIT News, CMU SEI, ACM Digital Library). * Track **reliability-focused conferences**: NeurIPS, ICML, AAAI, and FSE (Foundations of Software Engineering). * Engage with **standards bodies** (ISO/IEC JTC 1/SC 42 on AI) to see how reliability is being codified globally. Would you like me to curate a list of their **Twitter/X or LinkedIn profiles** so you can follow them directly for updates?

    Gemini3,604 characters · September 13, 2026

    Learning about **Model Reliability** (often intersecting with trustworthy AI, robust machine learning, ML observability, and MLOps) requires looking at both academic researchers pushing the boundaries of theory and practitioners solving production-level failures. The key experts, researchers, and organizations to follow depend on whether you are looking for academic rigor or production engineering: ### 1\. Academic & Research Pioneers * **Shiori Sagawa (Stanford University):** Her research focuses heavily on robust machine learning, distribution shift, and ensuring models don't fail when deployed in the real world where test data differs from training data. * **Aleksander Madry (MIT / OpenAI):** A leading voice in **adversarial robustness** and model reliability. His lab focuses on understanding how and why neural networks fail and how to make them resilient to edge cases and manipulation. * **Percy Liang (Stanford University):** As the director of the Stanford Center for Research on Foundation Models (CRFM), his work deals extensively with the evaluation, transparency, and reliability of complex AI systems (such as the Holistic Evaluation of Language Models or HELM framework). ### 2\. Industry Practitioners & Applied Reliability (MLOps/Observability) If you want to learn how to keep models reliable _in production_ , follow leaders from companies building monitoring and reliability infrastructure: * **Chip Huyen:** A prominent author, instructor, and consultant specializing in MLOps and production machine learning systems. Her work and writing focus intensely on why machine learning models break in production and how to build reliable data/model pipelines. * **Barr Moses & Lior Gavish (Monte Carlo / Data Reliability):** While focused broadly on data reliability, their work and writing are foundational for understanding upstream data errors - which are the number one cause of silent model failures. * **Emmanuel Ameisen:** Author of _Building Machine Learning Powered Applications_ , he has written extensively on validation, failure modes, and ensuring that a model's metrics actually translate to business and system reliability. ### 3\. Institutional Research Hubs * **Microsoft Research (Reliable ML Project):** Microsoft has dedicated a major cross-disciplinary initiative to **Reliable Machine Learning** , focusing on the intersection of model stability, fairness, and explainability. Their published papers, toolkits, and whitepapers are a gold standard. [Reliable Machine Learning - Microsoft Research](https://www.microsoft.com/en-us/research/project/reliable-machine-learning/#:~:text=However%2C%20these%20three%20goals%E2%80%94fairness%2C%20stability%20and%20explanation%E2%80%94are,fairness%2C%20and%20explainability%20of%20an%20ML%20model.) * **Google's People + AI Research (PAIR):** Focuses heavily on human-AI interaction, evaluation, and building dependable systems that users can trust. ### Where to Start Reading: * **Newsletters & Communities:** Look into _The Batch_ (by DeepLearning.AI) for weekly curation of reliability and safety updates, and MLOps-focused communities like _MLOps Community_ where production engineers share real-world failure post-mortems. * **Key Frameworks to Study:** Look into documentation and research surrounding **RobustBench** (for standardized model robustness) and **HELM** (for comprehensive model evaluation). Are you approaching model reliability more from a software/production engineering angle (MLOps) or a statistical/research perspective (robustness and distribution shift)?

    Google AI Mode2,604 characters · September 13, 2026

    To master **Model Reliability** - the intersection of machine learning, system stability, data drift management, and Site Reliability Engineering (SRE) - you should follow the engineers, researchers, and authors who literally wrote the playbook on keeping production ML systems stable. The top experts, authors, and thought leaders to follow include: * * * ** _D. Sculley_** - Former Google researcher and co-author of the seminal paper _"Hidden Technical Debt in Machine Learning Systems"_ as well as the comprehensive O'Reilly book __Reliable Machine Learning__. He is one of the definitive voices on production ML anti-patterns and reliability. * * * **_Niall Richard Murphy_** - A pioneer of Site Reliability Engineering (SRE) at Google and co-author of _Reliable Machine Learning_. Following his work bridges traditional software/system reliability with modern ML pipelines. * * * **_Cathy Chen_** - Co-author of _Reliable Machine Learning: Applying SRE Principles to ML in Production_ , focusing heavily on operationalizing ML systems so they don't silently fail in real-world enterprise environments. * * * **_Todd Underwood_** - An expert in reliability, security, and data science who co-authored the _Reliable Machine Learning_ text, offering deep insights into measurement, monitoring, and failure modes in large-scale infrastructure. * * * **_Andrej Karpathy_** - Former Director of AI at Tesla and co-founder of OpenAI. His technical breakdowns, educational repository styling, and writings on "Software 2.0" offer invaluable perspectives on why neural networks fail and how to reason about their non-deterministic behavior. * * * **_Chip Huyen_** - Author of _Designing Machine Learning Systems_. She focuses extensively on MLOps, data drift, monitoring, and building robust, scalable ML infrastructure that withstands production realities. [Reliable Machine Learning \[Book\] - O'Reilly](https://www.oreilly.com/library/view/reliable-machine-learning/9781098106218/)[Reliable Machine Learning: Applying SRE Principles to ML in](https://books.google.com/books/about/Reliable_Machine_Learning.html?id=c4iKEAAAQBAJ)[Top 10 AI & ML accounts to follow for an edge - LinkedIn](https://www.linkedin.com/posts/bhanu-paliwal-626223156_ai-machinelearning-deeplearning-activity-7377045553206976512-oG6L) If you want to narrow this down, tell me: * Are you focusing more on **LLM reliability** (hallucinations, eval, guardrails) or **traditional ML reliability** (data drift, pipelines, MLOps)? I can give you a more tailored list of newsletters, repos, or specific papers to check out.