01
What is an AI hallucination?
An AI hallucination is a statement from an AI system that is fluent and confident but false, or that isn’t supported by the sources it was given. Language models produce them because they generate likely text rather than look up checked facts, so a wrong answer can read exactly like a right one.
Researchers sort hallucinations by what they contradict. A survey of hallucination in language generation separates intrinsic errors, which contradict the source material, from extrinsic ones, which add claims the source can’t confirm 1. The distinction comes from work on summarisation, where a model’s summary can be checked against its document 2. For chat assistants, which often have no source document, a later survey separates factuality hallucinations, which conflict with real-world facts or invent them, from faithfulness hallucinations, which drift from the user’s instructions or from the context provided 3.
Some researchers prefer ‘confabulation’ for part of the problem. Farquhar and colleagues use it for answers that are both wrong and arbitrary: ask the same question again and the model may give a different wrong answer 4. Typical cases are an invented statistic, a quotation nobody said, a real person given someone else’s job, and references to papers that don’t exist, a failure notable enough to be singled out in theoretical work on why models hallucinate 5.
02
Why do language models hallucinate?
Language models hallucinate because they are trained to predict likely text, not to verify it, and because the way they are trained and graded rewards a confident guess over an honest ‘I don’t know’.
A language model learns by predicting the next token, a word or part of a word, across a very large body of text. What it learns is which continuations are probable, and probable is not the same as true. Asked a question, it produces the most plausible answer it can, whether or not reliable knowledge stands behind it.
Some errors are statistically unavoidable for a model that predicts well. Kalai and Vempala showed that for ‘arbitrary’ facts, which can’t be worked out from patterns in the data, a model that meets a statistical calibration condition must hallucinate at a rate close to the share of such facts that appear exactly once in its training data, even if that data contains no errors 5. Facts that appear many times, and systematic ones such as arithmetic, carry no such floor.
Training and testing then reinforce the habit. Kalai, Nachum, Vempala and Zhang argue that most benchmarks grade an answer as simply right or wrong, with no credit for admitting uncertainty, so a model that guesses scores higher than one that abstains, much as a student gains by guessing on an exam 6. They propose changing how existing benchmarks are scored, rather than adding new tests for hallucination.
Models also learn human mistakes. TruthfulQA, a benchmark of 817 questions across 38 categories, was written so that some people would answer falsely because of a common misconception. In its original 2021 tests, the best model was truthful on 58% of questions, against 94% for people, and the largest models were generally the least truthful: false answers that are common in text are exactly what a good imitator learns 7.
03
What is a knowledge cutoff?
A knowledge cutoff is the date after which a model has seen no training data. The model knows nothing that happened later, and unless it is told today’s date, it can’t tell how out of date its knowledge is.
The cutoff is fixed when the training data is collected, before the model is released, and the model is then used for months or years. Google’s model card gives January 2025 as the cutoff date of Gemma 4’s pre-training data 8. Brello Pro and Brello Vision, two of the three models in Brello 1.0, are based on Gemma 4 E4B and E2B, so on 4 October 2026, the date of Brello 1.0.0, their pre-training data was more than 20 months old.
A cutoff is also less precise than it sounds. Cheng and colleagues found that a model’s effective cutoff, the point its knowledge actually reflects, often differs from the date its developers report, partly because new web crawls contain a good deal of old text 9. The practical failure is quiet. Asked about something that changes, such as a price, a software version or who holds an office, a model answers from older knowledge as though it were current.
Giving the model the current date lets it reason about time, and retrieving fresh sources supplies facts newer than its training, but an answer about recent events is only as current as the pages it was given.
04
What is calibration in a language model?
A model is well calibrated when its confidence matches its accuracy: of the answers it gives with 80% confidence, about 80% are right. Calibration is what would let a model say ‘I’m not sure’ at the right moments, and only then.
Neural networks are not calibrated by default. Guo and colleagues found that modern networks, unlike those of a decade earlier, are poorly calibrated, and that a one-parameter adjustment called temperature scaling corrects much of the problem in classification tasks 10.
Large language models do better under narrow conditions. Kadavath and colleagues found that larger models are well calibrated on multiple-choice and true-or-false questions presented in the right format. The models could also estimate the probability that their own proposed answer was correct, and could be trained to predict whether they knew the answer to a question at all, a probability that rose when relevant source material was placed in the context. That last prediction was less well calibrated on new kinds of task 11.
The difficulty is that a model’s internal probabilities don’t appear in its prose. A sentence written at 30% confidence looks the same as one written at 99%, so an assistant has to be instructed or trained to put its uncertainty into words, and a reader can’t see how reliably it does so. That is why calibration is tested in its own right, as our explainer on AI safety evaluations describes.
05
What reduces hallucinations?
Grounding answers in retrieved sources, citing those sources, checking answers for consistency and rewarding honest uncertainty all reduce hallucinations. None of them removes hallucinations, so the realistic goal is fewer errors that are easier to catch.
Grounding means retrieving relevant passages and asking the model to answer from them rather than from memory. Retrieval-augmented generation, introduced by Lewis and colleagues in 2020, pairs a language model with a searchable index of documents 12, and applied to conversation, retrieval substantially reduced hallucinated knowledge in human evaluations 13. Our explainer on retrieval-augmented generation covers the mechanism. Figure 1 shows the difference it makes to a single answer.
- One question, two ways to answer it. From memory, the model has only the question; from sources, passages are retrieved first. The model is the same in both.
- From memory, the model writes the most likely answer. It is fluent and specific, and nothing in it shows which parts the model learned and which it guessed.
- One number is wrong. 8,516 m is the height of Lhotse, a neighbouring peak; Kangchenjunga is 8,586 m. The ranking is right, which makes the error easy to miss.
- Grounding puts evidence in the context first. A search returns passages about the mountain, and the model is asked to answer from them.
- The model answers from the passages and cites them. Each claim carries the number of the passage it came from.
- A reader can check each claim against its source. A citation shows where a claim came from, not that it is true: sources can be wrong, and models sometimes cite passages that don’t support the claim.
Grounding moves the problem rather than ending it. A model can misread a passage, blend it with what it already believes, or fall back on memory when the passages don’t cover the question. The sources can themselves be wrong or out of date, and retrieved text can carry instructions planted by someone else, as our explainer on prompt injection describes.
Citations make errors checkable, not impossible. In ALCE, a benchmark for answers with citations, even the best models lacked complete citation support about half the time on ELI5, a dataset of open-ended questions 14. An audit of four generative search engines found that only 51.5% of generated sentences were fully supported by their citations, and only 74.5% of citations supported the sentence they were attached to 15.
Consistency checks use the model’s own variability. If a model knows a fact, answers sampled several times tend to agree; if it is guessing, they diverge. SelfCheckGPT detects hallucinations this way without access to the model’s internals 16, and semantic entropy, which measures disagreement in meaning rather than in wording, detects confabulations across datasets and tasks 4. A related method, chain-of-verification, has the model draft an answer, plan questions that would check it, answer those independently and then revise; it reduced hallucinations across several tasks 17.
A system prompt can also ask a model to say when it is unsure, and training can reward it for doing so, the change Kalai and colleagues argue benchmarks should make 6. An instruction is the weakest of these measures: a model follows it only as reliably as it follows anything else.
06
Why do small models need extra care?
Small models hold fewer facts, so more questions fall near the edge of what they know, which is where hallucinations concentrate. Grounding and short, plain instructions matter more for them as a result.
How much a model knows about something depends on how often it saw it. Kandpal and colleagues found that a model’s accuracy on a factual question tracks the number of training documents related to that question; larger models learn more of this long tail, but would need to grow by many orders of magnitude to answer rarely mentioned facts well 18. Mallen and colleagues found the same pattern across 14,000 questions about subjects of varying popularity: scaling barely improved recall of less popular facts, while retrieval-augmented models outperformed far larger models working unaided 19.
A phone-sized model has a few billion parameters, far fewer than frontier cloud models, so it reaches that edge sooner. Brello 1.0’s three models are based on Qwen3 1.7B, Gemma 4 E2B and Gemma 4 E4B, and our explainer on small language models describes what their size means. Small models also tend to copy the shape of long instructions, so Brello 1.0 keeps its system prompt to four plain sentences, for reasons set out in ‘Small models copy the shape of their instructions’.
Sampling adds variation. Brello Pro and Brello Vision sample at a temperature of 1.0 and Brello Core at 0.7, or 0.6 for any of them with Think harder on, so the same question can produce different answers. That variation is also a signal: if two attempts at a factual answer disagree, at least one of them is wrong.
07
What does Brello 1.0 do about hallucinations?
Brello 1.0 tells its model to admit uncertainty, offers to search the web for time-sensitive questions, and asks the model to cite each source it uses. These measures are meant to reduce errors; we have not measured how far they do, and Brello 1.0 doesn’t check that its answers are true.
Brello 1.0 doesn’t verify that a cited passage supports the sentence it is attached to, and it doesn’t measure its own confidence. The repetition stopper and markup clean-up listed on our safety page catch malformed output, not false statements. Brello’s system prompt tells the model to admit uncertainty, but that is no guarantee. How web answers are assembled is described in ‘Answering from the open web, without a server’, and how we intend to test honesty and citation accuracy in Brello Super Intelligence, which is in development, is set out in ‘Evaluate first, then ship’.
08
How can you check an AI answer?
Check the specific claims an answer depends on against a source you can open. Where the answer cites sources, read the cited passage; where it doesn’t, treat numbers, names, dates, quotations and references as unverified until you have found them elsewhere.
- Start with the specifics. A fluent guess can get a number, a name or a date wrong while the sentence around it still reads well.
- Open the citation. Find the sentence in the source that supports the claim. A citation shows where a claim came from, not that it was reported faithfully: in the audit above, about a quarter of citations didn’t support their sentence.
- Search for references by title. A paper, book or court case cited by an AI should be findable. If it isn’t, assume it was invented.
- Ask again, differently. If rewording the question, or regenerating the answer, changes a fact, the model is probably guessing.
- Check the date. For anything that changes, such as prices, software versions, schedules or office holders, assume the model’s own knowledge is out of date unless the answer cites a current source.
- Don’t read confidence as evidence. A confident tone is a habit of the model’s writing, not a measure of how reliable a claim is.
How AI systems are tested for these failures before release is covered in AI safety evaluations, explained. The glossary defines hallucination, grounding and calibration, and the Brello 1.0 product page describes the app as a whole.
References
Reviewed . Web pages were checked on that date.
- Ji, Z. et al. (2023). “Survey of Hallucination in Natural Language Generation.” ACM Computing Surveys. arxiv.org/
abs/ 2202.03629 - Maynez, J. et al. (2020). “On Faithfulness and Factuality in Abstractive Summarization.” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020). arxiv.org/
abs/ 2005.00661 - Huang, L. et al. (2023). “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.” arXiv preprint. arxiv.org/
abs/ 2311.05232 - Farquhar, S. et al. (2024). “Detecting hallucinations in large language models using semantic entropy.” Nature 630, 625–630. doi.org/
10.1038/ s41586-024-07421-0 - Kalai, A. T. and Vempala, S. S. (2024). “Calibrated Language Models Must Hallucinate.” Proceedings of the 56th Annual ACM Symposium on Theory of Computing (STOC 2024). arxiv.org/
abs/ 2311.14648 - Kalai, A. T. et al. (2025). “Why Language Models Hallucinate.” arXiv preprint. arxiv.org/
abs/ 2509.04664 - Lin, S. et al. (2022). “TruthfulQA: Measuring How Models Mimic Human Falsehoods.” Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022). arxiv.org/
abs/ 2109.07958 - Google DeepMind. “Gemma 4 model card.” Google AI for Developers. Accessed 5 October 2026. ai.google.dev/
gemma/ docs/ core/ model_card_4 - Cheng, J. et al. (2024). “Dated Data: Tracing Knowledge Cutoffs in Large Language Models.” Conference on Language Modeling (COLM 2024). arxiv.org/
abs/ 2403.12958 - Guo, C. et al. (2017). “On Calibration of Modern Neural Networks.” Proceedings of the 34th International Conference on Machine Learning (ICML 2017). arxiv.org/
abs/ 1706.04599 - Kadavath, S. et al. (2022). “Language Models (Mostly) Know What They Know.” arXiv preprint. arxiv.org/
abs/ 2207.05221 - Lewis, P. et al. (2020). “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” Advances in Neural Information Processing Systems 33 (NeurIPS 2020). arxiv.org/
abs/ 2005.11401 - Shuster, K. et al. (2021). “Retrieval Augmentation Reduces Hallucination in Conversation.” Findings of the Association for Computational Linguistics: EMNLP 2021. arxiv.org/
abs/ 2104.07567 - Gao, T. et al. (2023). “Enabling Large Language Models to Generate Text with Citations.” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023). arxiv.org/
abs/ 2305.14627 - Liu, N. F., Zhang, T. and Liang, P. (2023). “Evaluating Verifiability in Generative Search Engines.” Findings of the Association for Computational Linguistics: EMNLP 2023. arxiv.org/
abs/ 2304.09848 - Manakul, P., Liusie, A. and Gales, M. J. F. (2023). “SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023). arxiv.org/
abs/ 2303.08896 - Dhuliawala, S. et al. (2023). “Chain-of-Verification Reduces Hallucination in Large Language Models.” arXiv preprint. arxiv.org/
abs/ 2309.11495 - Kandpal, N. et al. (2023). “Large Language Models Struggle to Learn Long-Tail Knowledge.” Proceedings of the 40th International Conference on Machine Learning (ICML 2023). arxiv.org/
abs/ 2211.08411 - Mallen, A. et al. (2023). “When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.” Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023). arxiv.org/
abs/ 2212.10511
Version history
- 1.0First published.



