THE MACHINE THRESHOLD
AI / EXPLAINER / 2 MIN READ + OPTIONAL DEEP DIVE

An AI repeats its answer. Would another model agree?

Consistency within one model and disagreement between models reveal different kinds of uncertainty.

AI-assisted synthesis · Published 2026-09-21 · Updated & sources checked 2026-09-21
How we research and correct our work

Repeating the same answer and reaching the right answer are different achievements.

Consistency is only one clue

Repeating an answer does not establish that it is correct. In a March 2026 research report, MIT describes an uncertainty method that adds comparisons with other language models to checks of one model’s consistency. The extra question is whether different models produce answers with similar meanings. [1]

Two different consistency questions
  1. Within one model: does it vary?
  2. Across models: do meanings differ?
  3. Then inspect the source evidence
Disagreement is a clue, not an answer key

Compare meaning, not identical wording

The reported method compares a target model with an ensemble: a small group of other models. Their diversity and credibility matter. It combines disagreement between models with a measure of uncertainty within the target model, rather than simply counting which wording wins a vote. [1]

Keep the result inside its scope

MIT reports that the disagreement signal was more useful for tasks with a single correct answer than for open-ended tasks. This article explains that reported idea; the original paper could not be read through its verification screen. We do not claim an independently confirmed accuracy gain or a universal hallucination detector. [1]

Go a little deeper

Optional reading · about 1 more minute

A harmless comparison

Hypothetical example: Two assistants name different publication years for a paper. Put both claims beside the paper’s own date. The disagreement tells you where to investigate; the original record settles this particular question. Asking a third assistant without checking that record would still leave the evidence missing.

A useful stopping question

Our interpretation: Before collecting more answers, decide what evidence could resolve the uncertainty. For a date it may be a publication record; for a contested interpretation it may require several sources and an explicit account of their differences. Agreement can be a clue, but it should not become your definition of truth.

Original sources

Attributed synthesis, not original reporting. Examples labeled hypothetical or illustrative are explanatory. Reviewing a source does not independently validate its findings.

  1. MIT: identifying overconfident language models ↗

    March 19, 2026 institutional report, read September 21. Original OpenReview paper blocked by browser verification; methods/results not independently reviewed.

Suggest a correction

Know someone who would find this interesting?

Share this story on Facebook ↗ ·

Follow on Facebook ↗ for story highlights and questions to explore next.

Where this question leads next

Follow new explainers and updates →