Do AI models agree about you? What 175 readings show
The Global AI Index runs the same reading on the same people every week: three prompts, seven models, four scored dimensions. That produces something a single ChatGPT lookup can't — a record of how much seven models disagree about the same person at the same moment.
Everything below is computed from our own snapshots for the week of 29 June 2026: 25 subjects from the Index cohort, seven models each, 175 individual model readings. That week is the sample because it is the most recent one where all seven models returned a reading for every subject in it. These 25 are well-documented public figures, so read the agreement below as the good case, not the typical one.
Do different AI models agree about you?
Less than a single composite score suggests. Across the 25 subjects, the median gap between a person's highest-scoring and lowest-scoring model was 11 points. The mean gap was 16. The narrowest was 2 points and the widest was 72. One number averaged over seven models hides all of that.
The extremes are the instructive part. Cleo Abram (journalist, technology) scored 82 from Perplexity and 10 from Meta AI in the same week. Ramez Naam (investor, clean energy) scored 81 from Perplexity and 10 from Claude. Both sit at the bottom of the sample on composite, at 62 and 66. The pattern holds across the cohort: the thinner the public record, the less the models agree about it. For 22 of the 25 subjects, every model came back at 70 or above.
Even a strong record is no guarantee. Lex Fridman scored a composite of 90 that week, with Gemini at 94 and Meta AI at 51.
The models agree on how confidently you are described
Accuracy is the dimension they converge on. Our accuracy score does not measure truth. It measures how specific and confidently stated an answer is — dates, titles, named organizations against hedging and vagueness. On that, 22 of the 25 subjects had all seven models land within 10 points of each other, and 162 of the 175 readings scored 90 or above.
That is worth reading correctly. It does not mean the models are right about these people. It means that when a model has something to say about a well-documented person, it says it in roughly the same confident register as every other model. Confidence is cheap to produce and nearly uniform across providers, which is exactly why a fluent AI answer is not evidence of a correct one.
Sentiment was the second-tightest dimension: 12 of the 25 subjects had all seven models within 10 points, and only 7 of the 175 readings came in below 70. Models describe recognized people warmly, and they do it consistently.
The models disagree on whether you show up at all
Presence is where they split, and presence carries the heaviest weight in the composite at 30%. Only 5 of the 25 subjects had all seven models within 10 points on presence. The median spread was 50 points, half the scale.
Presence comes from three answers: a direct question about the person, which carries half the score, and two peer-set questions ("who are the leading people in X?") that split the rest. So it lands on a coarse ladder, and the coarseness is the finding. Four of the 25 subjects earned a perfect presence score from all seven models. Seven of the 25 earned one from none of them. Of the 175 readings, 83 were perfect and another 57 came in at half or below.
Prominence — where in an answer the name lands — was almost as scattered. Only 2 of the 25 subjects had all seven models within 10 points of each other on it.
The practical version: whether a model can describe you is a much more settled question than whether it will bring you up unprompted. The second one is the harder problem, and it is the one most people have not looked at.
Which models scored high and which scored low
Averaged over the 25 subjects, Perplexity scored highest at 87.1, then Grok at 86.7, Gemini at 86.4, DeepSeek at 85.6, ChatGPT at 84.3, Meta AI at 80.5, and Claude at 77.9. A nine-point gap between the top and bottom model, on the same people, in the same week.
Being the lowest of the seven is more concentrated than being the highest. Claude was the single lowest scorer for 9 of the 25 subjects and Meta AI for 7 — between them, two-thirds of the sample. Perplexity was the highest for 7, which is what you would expect from the one model in the set with a live retrieval layer. Neither ChatGPT nor Claude was the top scorer for a single subject.
None of this ranks the models by quality. A low score means the model said less about that person, not that it was wrong. A model that answers "I don't have reliable information about this person" scores low here, and that answer is often the correct one. The score reads the record, not the model's judgment. What we don't claim sets out the rest of the limits.
In two weeks, almost nothing moved
Between the 29 June and 13 July snapshots, 16 of the 25 composites were identical, and 24 of the 25 moved by 2 points or less. The mean absolute change was 0.5 points. The largest single move in the cohort was 3 points, and it was downward: 7 subjects drifted down over the fortnight, 2 drifted up.
Keep that in view if you are working on your own record. Two weeks of AI training data is close to nothing, and drift runs in both directions without anyone doing anything. This is why the measurement exists to report whether the record moved rather than to promise that it will — the reasoning behind that is on the methodology page. Over a fortnight, the honest answer for almost everyone in this cohort was that it didn't.
How to read your own numbers
Read more than one model, and read the lowest one first. A composite in the eighties can contain a model that returned almost nothing about you, and that model is the one carrying information you can act on: some part of your record is invisible to some part of the training data.
Then read presence separately from accuracy, because they fail differently. High accuracy with low presence means the models can describe you when asked directly but won't raise your name on their own, which is a distribution problem — the moves for it are in How to change what AI models say about you. Uneven accuracy means the sources disagree about your facts, which is why ChatGPT gets bios wrong in the first place.
Finally, expect your field to be part of the answer. Across the 84 subjects on the board at the time of this study (it has since grown to 112), sector averages run from 79 in the creator economy to 90 in AI and machine learning. That range is about as wide as the gap between two models reading the same person. Some of what your number reflects is how well documented your whole sector is, which is context worth having before you read any single figure as a verdict on you.
The fastest way to see your own spread is a free reading — it shows all seven answers verbatim. If you want the disagreement tracked over time, Index Monitor rescans weekly and keeps the history.
FAQ
Do different AI models say the same thing about you?+
Not closely. Across 25 public figures read by seven models in the same week, the median gap between a person's highest-scoring and lowest-scoring model was 11 points, and the widest was 72. The thinner the public record, the less the models agree about it.
Which AI model knows the most about people?+
No model wins outright — it varies by person. Averaged over the 25 subjects, Perplexity scored highest and Claude lowest, a nine-point gap. But a low score means the model said less about that person, not that it was wrong. A candid answer admitting no reliable information also scores low.
How quickly does an AI presence score change?+
Slowly, and not always upward. Between two snapshots two weeks apart, 16 of 25 composites were identical and 24 of 25 moved by 2 points or less. The mean absolute change was 0.5 points, and more subjects drifted down than up. The score reports movement rather than promising it.