Answer summary
The AI models with the lowest hallucination rates are Zhipu AI's GLM-4-9B-Chat and Google's Gemini-2.0-Flash-Exp, both with a rate of 1.3%. However, the ranking can vary depending on the benchmark and task. Anthropic's Claude models and OpenAI's GPT models are also considered to have low hallucination rates in certain contexts.
Sign in to check whether you are mentioned or cited in this prompt.
Where each engine placed them
What the models searched first
The sub-queries each engine actually searched before answering.
Unlock the full picture
The sub-queries each engine actually searched before answering.
Your tracked entities
Save a voice, brand or tool and see where it stands on this question.
Full responses
5 of 5 engines that answered. Full stored text as captured, with dashes shown as hyphens. Nothing is summarised.
ChatGPT2,242 characters · September 7, 2026
There isn’t one universally “least-hallucinating” AI - the answer depends heavily on the benchmark and task. As of **September 2026**, the best-supported picture is: - **For grounded summarization:** Vectara’s HHEM benchmark currently has **GPT-5.4 nano at 3.1%** and **Gemini 2.5 Flash-Lite at 3.3%** among its reported models; a specialized model, Finix S1 32B, is lower at 1.8%. [hallucination-leaderboard/README.md at main · vectara/hallucination-leaderboard · GitHub](https://github.com/vectara/hallucination-leaderboard/blob/main/README.md?plain=1&utm_source=chatgpt.com) - **For broad factual accuracy:** Google’s newer **FACTS benchmark** currently ranks **Gemini 3.7 Flash first**, followed by Gemini 3.1 Pro Preview and GPT-5.5. [FACTS Benchmark Suite Leaderboard | Kaggle](https://www.kaggle.com/benchmarks/google/facts?utm_source=chatgpt.com) - **For real-world open-ended questions:** results are much less consistent. Hallucination rates can change substantially depending on whether the model is allowed to browse, the subject matter, and whether the benchmark penalizes refusing to answer. [3.2 ASSESSING RESPONSIBLE AI | RESPONSIBLE AI | AI INDEX REPORT 2026](https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf?utm_source=chatgpt.com) ### My practical ranking If your priority is **“give me the most reliable answer and don't make things up”**, I'd currently put: 1. **Gemini 3.7 Flash** - strongest recent factuality benchmark result. 2. **GPT-5.5 / GPT-5.6-class models** - very strong, particularly when browsing/citation is available. 3. **Claude's latest models** - generally excellent at careful, qualified answers, though benchmark results vary. 4. **Perplexity** - particularly useful when you want answers grounded in web sources rather than relying purely on model memory. The important distinction is that **“hallucination rate” is not a single universal statistic**. A model can be #1 on document-grounded summarization while being worse on obscure factual questions. If you tell me what you mainly use AI for - **research, coding, medical information, legal questions, writing, or everyday factual questions** - I can give you a much more meaningful “least hallucination” ranking for that use case.
Copilot2,713 characters · September 7, 2026
### 📊 Current Findings on AI Models with the Lowest Hallucination Rates (2025) Hallucination rates vary depending on the benchmark used, but several recent studies highlight which AI models are best at minimizing false or fabricated outputs. #### 🔹 **Vectara HHEM Benchmark (Document Summarization)** * **Zhipu AI GLM-4-9B-Chat** → **1.3% hallucination rate** * **Google Gemini-2.0-Flash-Exp** → **1.3%** * **OpenAI o1-mini** → **1.4%** * **OpenAI GPT-4o** → **1.5%** * **GPT-4o-mini / GPT-4-Turbo** → **1.7%** These results suggest that smaller, specialized models often hallucinate less than very large ones.[Which AI Models Are Leading the Way in Reducing Hallucinations and Improving Accuracy?](https://www.digitalinformationworld.com/2025/01/which-ai-models-are-leading-way-in.html) #### 🔹 **Omniscience Benchmark (Knowledge Calibration & Factual Recall)** * **Claude 4.1 Opus (Anthropic)** → **Lowest hallucination rate overall**, with strong calibration (knows when to abstain rather than guess). * **Claude 4.5 Haiku** → **26% hallucination rate**, significantly lower than most frontier models. * **GPT-5.1 and Grok 4** showed high accuracy but much higher hallucination rates (64–81%).[Anthropic Models Take Top 3 Spots With Lowest Hallucination Rates In New Omniscience Benchmark](https://officechai.com/ai/anthropic-models-take-top-3-spots-with-lowest-hallucination-rates-in-new-omniscience-benchmark/) #### 🔹 **HKU Hallucination Control Capability Report** * **GPT-5(Thinking)** and **GPT-5(Auto)** ranked **top globally** for hallucination control. * **Claude 4 Opus series** also performed strongly. * Among Chinese models, **ByteDance Doubao 1.5 Pro** was best, though still behind leading international systems.[A New In-Depth Report of AI Large Language Models: Hallucination Control | The HKU Lead](https://lead.hku.hk/en/article/2025/09/a-new-in-depth-report-of-ai-large-language-models-hallucination-control/) ### ⚖️ Key Takeaways * **Lowest hallucination rates (≈1–2%)** are seen in **GLM-4-9B-Chat, Gemini-2 Flash, and OpenAI o1-mini/GPT-4o** under document summarization tests. * **Anthropic’s Claude models** excel in **knowledge calibration benchmarks**, showing restraint and lower hallucination rates in open-domain Q&A. * **GPT-5 variants** lead in **hallucination control capability** in professional evaluations. * **No single model dominates across all benchmarks** - performance depends on the task (summarization vs. open-domain Q&A vs. reasoning). 👉 Would you like me to create a **comparison table** showing the top models across different benchmarks (Vectara, Omniscience, HKU) so you can see which excels in each type of test?
Gemini3,555 characters · September 7, 2026
Determining which AI has the lowest hallucination rate depends heavily on how the evaluation is measured (e.g., grounded document summarization vs. open-domain trivia and reasoning). Independent leaderboards like Vectara’s HHEM and enterprise testing benchmarks highlight distinct frontrunners: [Which AI Has the Lowest Hallucination Rate? (2026 Data) - Seekr](https://www.seekr.com/resource/ai-lowest-hallucination-rate/#:~:text=Different%20benchmarks%20test%20very%20different%20things%2C%20which,a%20document%20and%20asks%20it%20to%20summarize.) * **Grounded Summarization and Data Extraction:** Specialized smaller-footprint models or specific checkpoint variations often score the lowest hallucination rates when provided with source text. For instance, lightweight models like **GPT-5.4 Nano** and specialized enterprise models (such as **Finix-S1** or **Phi-4**) routinely register hallucination rates under 4% on Vectara’s factual consistency benchmarks because they stick tightly to provided source text. [Vectara Hallucination Leaderboard - GitHub](https://github.com/vectara/hallucination-leaderboard#:~:text=Table_title%3A%20Hallucination%20Leaderboard%20Table_content%3A%20%7C%20Model%20%7C,%25%20%7C%2099.8%20%25%20%7C%2070.6%20%7C) * **General-Purpose Frontier Models:** Among major flagship conversational models, Anthropic's **Claude** series (such as Claude 4.6 and its variants) frequently ranks at the top for epistemic calibration. They tend to exhibit higher rates of declining to answer or stating uncertainty ("I don't know") rather than fabricating details, which lowers overall factual hallucination. [Which AI Hallucinates the Least? 5 Models Tested (2026) - Talkory.ai](https://www.talkory.ai/blog/ai-models-lowest-hallucination-rate-2026#:~:text=AI%20Models%20with%20the%20Lowest%20Hallucination%20Rate,GPT%2D5.4%20is%20a%20close%20second%20at%20~6%25.)[Which AI Hallucinates the Least? 5 Models Tested (2026) - Talkory.ai](https://www.talkory.ai/blog/ai-models-lowest-hallucination-rate-2026#:~:text=Claude%204.6%3A%20Lowest%20Hallucination%20Rate%20in%202026.,says%20so%20rather%20than%20fabricating%20an%20answer.) * **The "Reasoning" Trade-Off:** Intensive reasoning models (such as advanced reasoning iterations from OpenAI, DeepSeek, and xAI) excel at logic, math, and coding. However, multi-step chain-of-thought processing can paradoxically act as a hallucination amplifier on open-domain factual retrieval, sometimes increasing fabrication rates when relying purely on internal training memory. [AI Model Hallucination Rates 2026: The Definitive Honesty Rankings](https://codingfleet.com/blog/ai-model-hallucination-rates-2026/#:~:text=*%20Reasoning%20mode%20is%20a%20hallucination%20amplifier.,provider%20endpoint%2C%20not%20just%20the%20model%20card.) Ultimately, no single AI has a 0% hallucination rate. For high-stakes factual or legal tasks, the lowest effective error rates are achieved by using **grounded retrieval-augmented generation (RAG)** or multi-model consensus frameworks. [Which AI Has the Lowest Hallucination Rate? (2026 Data) - Seekr](https://www.seekr.com/resource/ai-lowest-hallucination-rate/#:~:text=Different%20benchmarks%20test%20very%20different%20things%2C%20which,a%20document%20and%20asks%20it%20to%20summarize.)[Which AI Hallucinates the Least? 5 Models Tested (2026) - Talkory.ai](https://www.talkory.ai/blog/ai-models-lowest-hallucination-rate-2026#:~:text=Grok%204.20%20and%20Gemini%203.1%20lag%20significantly,AI%20hallucination%20rate%20by%2070%25%2B%20in%20production.)
Google AI Mode4,306 characters · September 7, 2026
# Shared 0 files **Anthropic’s Claude (specifically Claude 4.6 Sonnet) and OpenAI’s GPT models (such as GPT-5.4)** consistently record the lowest overall hallucination rates among major AI models. [LLM Hallucination Rates 2026: Best and Worst Models - ModelsLab](https://modelslab.com/blog/llm/llm-hallucination-rates-2026)[Which AI Hallucinates the Least? 5 Models Tested (2026) - Talkory.ai](https://www.talkory.ai/blog/ai-models-lowest-hallucination-rate-2026)[How to avoid AI hallucinations | Tim Burke, CTO](https://www.youtube.com/shorts/B3FaYMNIoDE) However, there is no single "perfect" AI. How much an AI hallucinates depends entirely on the **type of task** you give it. A model that rarely hallucinates when summarizing a document might frequently make things up when answering trivia from memory. [Which AI Has the Lowest Hallucination Rate? (2026 Data)](https://www.seekr.com/resource/ai-lowest-hallucination-rate/) The performance varies across different benchmarks: 1\. Grounded Summarization (Lowest Rates) When you give an AI a specific document and ask it to summarize it, hallucination rates are at their lowest because the answer is right in front of the model. [Seekr] * **Top Performers:** Anthropic's Claude 4.6, OpenAI’s GPT-5.4, and specialized models like Ant Group's Finix. * **The Score:** These models maintain an exceptionally low **0.6% to 4% hallucination rate** on strict summarization leaderboards like Vectara's HHEM. [Latest AI Hallucination Rates & Benchmarks for New AI Models 2026](https://suprmind.ai/hub/ai-hallucination-rates-and-benchmarks/) 2\. Open-Domain Factuality (Moderate to High Rates) When you ask an AI a factual question without providing text - forcing it to rely entirely on its own training data - hallucinations skyrocket across the board. [Seekr] * **Top Performers:** Claude 4.5/4.6 Sonnet and OpenAI's flagship models usually perform best here because they are optimized to recognize uncertainty and say "I don't know" rather than guessing. [Which AI Models Hallucinate the Most? : r/charts - Reddit](https://www.reddit.com/r/charts/comments/1p5mcxw/which_ai_models_hallucinate_the_most/)[AI Models with Low Hallucination Rates Revealed | Majd Y. posted on](https://www.linkedin.com/posts/majdyafi_which-ai-models-hallucinate-the-most-hallucination-activity-7409283057637146624-Ne8M) * **The Score:** Even for advanced reasoning models, hallucination rates on open-domain factuality tests (like SimpleQA or PersonQA) jump significantly, ranging from **15% to 33%+**. [Seekr +1] 3\. High-Stakes Domains (Highest Rates) In specialized fields like law or medicine, general AI models struggle deeply with precision without specific guardrails. [Seekr] * **The Score:** Studies from institutions like Stanford's RegLab show that standard AI models can hallucinate on **69% to 88% of complex legal queries**. In clinical medical tests, baseline hallucination rates frequently exceed **50% to 65%**. [HALC-Bench: LLM Hallucination on Long-Context Retrieval Benchmark](https://aimultiple.com/ai-hallucination) * * * How to Actually Stop AI Hallucinations If you want an AI that doesn't hallucinate, the solution isn't just picking a better model - it is changing how you use it. You can drastically reduce false information by applying these strategies: [Seekr +1] * **Use RAG (Retrieval-Augmented Generation):** Connect the AI directly to a verified database, database systems like [MongoDB Atlas](https://www.mongodb.com/resources/basics/ai-hallucinations), or the live web so it fetches real data before answering. [What Are AI Hallucinations | How to Account for Them - MongoDB](https://www.mongodb.com/resources/basics/ai-hallucinations) * **Force Chain-of-Thought Prompting:** Explicitly instruct the AI to explain its step-by-step reasoning and cite its sources before outputting a final answer. [YouTube·Arch Systems] * **Give it an "Out":** Always include a line in your prompt like: _"If you do not know the answer based strictly on the text, state 'I do not know' - do not guess."_ [Reddit +1] Are you looking for an AI to use for a **specific task** (like coding, medical research, or drafting legal papers)? Let me know your **use case** so I can recommend the safest model and prompting strategy for your project!
Perplexity1,539 characters · September 7, 2026
Short answer: There isn’t a single AI model that universally has the lowest hallucination rate across all benchmarks and tasks. Different studies and benchmarks report different leaders depending on the test set and metrics used. For example, some analyses have highlighted Claude variants as having very low hallucination rates on knowledge tasks, while others point to Gemini-2.0-Flash-Esp or smaller specialized models as among the lowest in certain benchmarks. Keep in mind that “lowest hallucination rate” is highly benchmark-dependent and can vary with task type (fact-checking, summarization, open-ended QA) and how hallucinations are defined and measured.[Which AI Hallucinates the Least? 5 Models Tested (2026)](https://www.talkory.ai/blog/ai-models-lowest-hallucination-rate-2026)[AI Hallucination Rates, Statistics & Benchmarks in 2026](https://suprmind.ai/hub/ai-hallucination-rates-and-benchmarks/)[Which AI Has the Lowest Hallucination Rate? (2026 Data)](https://www.seekr.com/resource/ai-lowest-hallucination-rate/)[AI Hallucination Rates Across Different Models 2026](https://www.aboutchromebooks.com/ai-hallucination-rates-across-different-models/) If you’re evaluating for a specific use case (e.g., factual QA in tech, medical, or general knowledge), I can help compare how different models perform on that task, including the benchmark definitions and caveats. Would you share your target domain and the kind of hallucinations you want to minimize (fabricated facts, misattributed sources, or unsupported conclusions)?