Who should I follow to learn about AI Observability?

    Updated September 13, 2026
    CONSENSUS
    Chip Huyen

    Answer summary

    To learn about AI Observability, follow experts such as Chip Huyen, Shreya Shankar, and the teams behind leading AI observability platforms like Arize AI and LangChain, as well as researchers and practitioners in the field.

    Your presence

    Sign in to check whether you are mentioned or cited in this prompt.

    Where each engine placed them

    #BrandChatGPTGemini logoGeminiAI ModeCopilot logoCopilotMention rateBrand Share of VoiceCiteRankCitabilitySentiment
    1CHChip Huyen - 14% -
    ChatGPTGeminiAI ModeCopilot
    Mention rate - Share of Voice 14%CiteRank - Citability Sentiment
    2SSShreya Shankar - 7% -
    ChatGPTGeminiAI ModeCopilot
    Mention rate - Share of Voice 7%CiteRank - Citability Sentiment

    What the models searched first

    The sub-queries each engine actually searched before answering.

    Unlock the full picture

    The sub-queries each engine actually searched before answering.

    Your tracked entities

    Save a voice, brand or tool and see where it stands on this question.

    Full responses

    4 of 4 engines that answered. Full stored text as captured, with dashes shown as hyphens. Nothing is summarised.

    ChatGPT4,152 characters · September 13, 2026

    If you want to get good at **AI Observability** (LLM/agent tracing, evaluation, monitoring, reliability, cost, and debugging), I’d follow a mix of **practitioners, researchers, open-source builders, and platform teams**. ## People worth following - Chip Huyen - Excellent for production ML systems, reliability, and AI engineering principles. Her writing connects model behavior with real-world deployment challenges. - Hamel Husain - Strong focus on LLM evaluation, testing, and building reliable AI applications. - Shreya Shankar - Great perspective on data quality, evaluation, and practical ML systems. - Made With ML - A strong resource for production ML foundations: monitoring, pipelines, testing, and deployment. - Eugene Yan - Writes deeply about ML systems, experimentation, and operating models in production. ## AI observability companies and engineering blogs - [Arize AI](https://arize.com?utm_source=chatgpt.com) - One of the leading voices in AI observability and evaluation. Their open-source project Arize AI Phoenix is especially useful for learning traces, evaluations, and debugging workflows. [Arize vs. LangSmith](https://arize.com/compare/arize-vs-langsmith/?utm_source=chatgpt.com) - [LangChain](https://www.langchain.com?utm_source=chatgpt.com) / [LangSmith](https://smith.langchain.com?utm_source=chatgpt.com) - Follow for agent tracing, evaluation workflows, and production debugging patterns. LangSmith is widely used for tracing and evaluating LLM applications, especially in LangChain/LangGraph ecosystems. [LangSmith vs Arize: AI agent observability, evals, and deployment compared](https://www.langchain.com/resources/langsmith-vs-arize?utm_source=chatgpt.com) - [Langfuse](https://langfuse.com?utm_source=chatgpt.com) - Good open-source perspective on LLM tracing, prompt management, and evaluation. - [Weights & Biases](https://wandb.ai?utm_source=chatgpt.com) - Follow if you care about experiment tracking, model lifecycle management, and extending ML workflows into LLM applications. [LLM Observability Tools: Weights & Biases, Langsmith](https://aimultiple.com/llm-observability?utm_source=chatgpt.com) - [OpenTelemetry](https://opentelemetry.io?utm_source=chatgpt.com) - Essential foundation. AI observability is increasingly built around traces, spans, metrics, and vendor-neutral telemetry standards. [LLM Observability Tools: Weights & Biases, Langsmith](https://aimultiple.com/llm-observability?utm_source=chatgpt.com) ## Topics to study (in roughly this order) 1. **Distributed tracing fundamentals** - Spans, traces, context propagation - OpenTelemetry concepts 2. **LLM application telemetry** - Prompt + response capture - Token usage and latency - Retrieval traces (RAG) - Tool calls and agent trajectories 3. **Evaluation systems** - Offline eval datasets - Human feedback loops - LLM-as-a-judge - Regression testing 4. **Production reliability** - Hallucination detection - Prompt injection monitoring - Drift detection - Cost optimization 5. **Agent observability** - Multi-step traces - Tool failures - Planning/reasoning behavior - Safety boundaries ## Newsletters / communities - [MLOps Community](https://mlops.community?utm_source=chatgpt.com) - Good for production ML and LLMOps discussions. - [Hamel Husain's Substack](https://hamel.dev/blog?utm_source=chatgpt.com) - Practical AI engineering and evaluation. - [Chip Huyen's blog](https://huyenchip.com/blog?utm_source=chatgpt.com) - Production AI systems and engineering. ## A practical “follow list” if you only have 30 minutes/week 1. Arize AI blog + Phoenix project 2. LangSmith/LangChain engineering posts 3. OpenTelemetry GenAI work 4. Chip Huyen’s writing 5. Hamel Husain’s evaluation content 6. MLOps Community discussions The field is moving quickly, but the durable concepts are: **telemetry → evaluation → feedback loops → automated improvement**. Most AI observability tools are variations on that core loop. [ML Observability Hub: A Guide to Production LLM Monitoring](https://sentryml.com/posts/ml-observability-hub/?utm_source=chatgpt.com)

    Copilot5,281 characters · September 13, 2026

    **If you want to learn about AI Observability, follow experts who bridge AI, monitoring, and enterprise operations - people like Cassie Kozyrkov (decision intelligence), Andrew Ng (AI implementation), and leaders from observability platforms such as Datadog, Dynatrace, and Splunk. These voices combine technical depth with practical insights into how AI reshapes monitoring and reliability.** ## 🔑 Key Experts to Follow Expert| Focus Area| Why Follow Them ---|---|--- **Cassie Kozyrkov**| Decision intelligence, enterprise AI| She explains how organizations can make better decisions with AI, including observability of AI systems.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Andrew Ng**| AI education, practical ML deployment| His emphasis on operational AI makes him valuable for understanding observability in production.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Fei-Fei Li**| Human-centered AI| Offers perspective on responsible monitoring and observability of AI systems.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Demis Hassabis**| AI in science (DeepMind)| His work highlights how complex AI systems are monitored in research environments.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Ethan Mollick**| AI for business leaders| Shares practical frameworks for integrating AI into workflows, including observability concerns.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Yoshua Bengio**| Responsible AI research| Advocates for transparency and monitoring in AI deployments.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Melanie Mitchell**| AI development & complexity| Helps frame observability challenges in complex adaptive AI systems.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Rana el Kaliouby**| Human-aware AI| Focuses on monitoring emotional AI systems responsibly.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Jeremy Howard**| Democratizing AI education| Practical insights into monitoring AI models at scale.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) **Allie K. Miller**| Enterprise AI adoption| Shares strategies for operationalizing AI, including observability practices.[10 Trustworthy AI Thought Leaders Worth Following in 2026](https://mondo.com/insights/trustworthy-ai-thought-leaders-worth-following/) ## 🛠️ Observability Platform Leaders Beyond individual thought leaders, **vendors are shaping the AI observability landscape**: * **Datadog** → Leading innovation in AI-powered monitoring, including GenAI observability for large language models.[AI Is Reshaping Observability: What Enterprise Leaders Say](https://research.etr.ai/etr-data-drop/ai-is-reshaping-observability-what-enterprise-leaders-say) * **Dynatrace** → Strong in automated root-cause analysis and end-to-end session monitoring.[AI Is Reshaping Observability: What Enterprise Leaders Say](https://research.etr.ai/etr-data-drop/ai-is-reshaping-observability-what-enterprise-leaders-say) * **Splunk** → Widely used for log aggregation and customizable dashboards, critical for AI observability.[AI Is Reshaping Observability: What Enterprise Leaders Say](https://research.etr.ai/etr-data-drop/ai-is-reshaping-observability-what-enterprise-leaders-say) * **Grafana** → Known for flexible dashboards, often used to visualize AI system metrics.[AI Is Reshaping Observability: What Enterprise Leaders Say](https://research.etr.ai/etr-data-drop/ai-is-reshaping-observability-what-enterprise-leaders-say) ## ⚠️ Risks & Considerations * **Hype vs. Reality**: Many “AI observability” discussions are marketing-heavy. Prioritize experts with _real-world deployment experience_. * **Vendor Lock-in**: Observability tools differ in scope - Datadog may innovate faster, but Splunk remains entrenched in enterprise environments. * **Security Overlap**: Observability increasingly blends with security monitoring; following leaders in both domains is wise.[AI Is Reshaping Observability: What Enterprise Leaders Say](https://research.etr.ai/etr-data-drop/ai-is-reshaping-observability-what-enterprise-leaders-say) 👉 To deepen your learning, I’d suggest starting with **Cassie Kozyrkov’s talks on decision intelligence** and **Andrew Ng’s courses on operational AI**, while also tracking **Datadog and Dynatrace product blogs** for hands-on observability practices. Would you like me to curate a **short list of podcasts, newsletters, or blogs** specifically focused on AI observability so you can follow them regularly?

    Gemini2,411 characters · September 13, 2026

    To master AI observability, tracking the founders, engineers, and researchers building the leading platforms and open-source standards is the most effective approach. **Key Industry Leaders and Where to Follow Them** * **Shreya Shankar** (Researcher & PhD candidate at UC Berkeley) - Known for her deep research into MLOps, data reliability, and debugging machine learning systems. Follow her on **X/Twitter** and **LinkedIn** for academic-meets-practical insights on production AI issues. * **Chip Huyen** (Author of _Designing Machine Learning Systems_) - While broadly focused on MLOps and LLMops, she frequently writes and speaks about monitoring, evaluation, and production reliability for AI architectures. Follow her on **Substack** and **LinkedIn**. * **Creators of Arize AI (Jason Lopatecki & Aparna Dhinakaran)** - As the founders of Arize (creators of Phoenix), they regularly share case studies, telemetry frameworks, and best practices for tracing LLMs and embeddings. Follow them on **LinkedIn**. * **Haseeb Qureshi & The Braintrust Team** - Braintrust focuses heavily on evaluation and observability infrastructure for AI apps. Their engineering blog and founder posts offer tactical breakdowns of monitoring agentic workflows. * **The OpenTelemetry GenAI Semantic Conventions Working Group** - Rather than a single person, following key contributors to the **OpenTelemetry (OTel)** project on GitHub provides a front-row seat to how standardized AI logging, tracing, and metrics are being defined globally. **Platforms and Technical Blogs to Watch** * **Arize AI Blog & Phoenix Docs:** Excellent for practical guides on evals and tracing open-source models. * **LangChain / LangSmith Blog:** Great for learning how application-layer tracing integrates with orchestration frameworks. * **Full-Stack Vendors (Datadog & New Relic Engineering Blogs):** Useful for tracking how traditional enterprise observability scales into monitoring tokens, latency, and model drift. [Guide to AI Observability: Core Components, Tools, and Best Practices](https://newrelic.com/blog/ai/ai-in-observability#:~:text=AI%20systems%20generate%20massive%20telemetry%20volumes.%20Focus,*%20Integrate%20AI%20observability%20into%20existing%20workflows.) Whose specific angle - research-heavy, infrastructure-focused, or application-developer oriented - are you most interested in exploring further?

    Google AI Mode3,505 characters · September 13, 2026

    To truly understand **AI Observability** , **LLMOps** , and production-grade AI systems, you should follow the founders, engineers, and researchers building the telemetry layers, tracing open-source tools, and defining best practices. The top experts, practitioners, and organizations to follow span several key areas: 1\. The Practitioners & Educators (Hands-On Engineering) * * * **_Jason Liu_** (Developer Experience at OpenAI, Creator of `Instructor`): Widely followed for his deep-dive practical advice on structured outputs, data validation, and building reliable, observable interactions with LLMs. Follow him on LinkedIn (_Jason Liu_) or X/Twitter. * * * **_Chip Huyen_** (AI/MLOps Author and Researcher): Author of _Designing Machine Learning Systems_ and prominent writer on LLMOps. Her technical blogs and guides offer a foundational framework for understanding how production machine learning and LLM pipelines fail and how to monitor them. Check out her insights on Chip Huyen's Blog. * * * **_Charity Majors_** (Co-founder & CTO of Honeycomb): A legendary voice in traditional observability who frequently writes and speaks about how high-cardinality debugging, tracing, and production telemetry must evolve to handle complex AI and agentic software behavior. Follow her via Honeycomb's Blog. [MLOps guide - Chip Huyen](https://huyenchip.com/mlops/)[Jason Liu](https://www.linkedin.com/posts/jxnlco_heres-what-you-need-to-know-if-youre-an-activity-7221937592722677763-v2rj)[What is LLMOps and how does it work? - Weights & Biases](https://wandb.ai/site/articles/what-is-llmops-and-how-does-it-work/) 2\. Founders & Teams Behind Leading AI Observability Platforms Following the core teams building specialized tools gives you real-time access to production case studies, prompt tracing architectures, and evaluation metrics: * **Arize AI & Phoenix Team:** Arize specializes heavily in LLM tracing, evaluation, and root-cause analysis. Their open-source project, _Phoenix_ , is a staple for AI observability. Following the Arize AI Blog provides granular breakdowns of hallucination tracing and embedding drift. [15 Best AI Observability Tools for Production Teams in 2026](https://www.honeycomb.io/blog/best-ai-observability-tools) * **LangChain / LangSmith Team:** As the creators of core orchestration frameworks, they heavily influence how developers trace agent execution graphs. Tracking updates from LangSmith gives a window into multi-step agent debugging. * **Braintrust & Galileo Teams:** Both platforms focus on evaluation, experimentation, and production observability. Their engineering blogs routinely publish benchmarks and deep dives on tracking LLM performance drift. [www.honeycomb.io] 3\. OpenTelemetry & Cloud-Native Ecosystem * **Grafana Labs & OpenTelemetry Communities:** As enterprise observability expands to cover tokenomics, GPU utilization, and agent latencies, the OpenTelemetry AI semantic conventions are shaped heavily by contributors here. Keep an eye on sessions from events like Grafana Labs ObservabilityCON. [Observability in the Age of AI with Jeetu Patel, Atin Sanyal and](https://www.youtube.com/watch?v=WoBjCyl2akE&t=326) If you want to narrow this down, let me know: * Are you focusing more on **infrastructure/GPU metrics** or **LLM output tracing/evals**? * Do you prefer **open-source tools** or **enterprise SaaS platforms**? I can recommend specific repositories, newsletters, or starter guides tailored to your focus.