Skip to content
IndexMeAI
Sign in
Field manual · Rev 2026.06

Methodology.

What we measure, why these four signals, and how they roll up into one number. The framework is open; every reading shows you the answers it was scored from.

PROBESMODELS · weightedDIMENSIONS · per modelINDEXper_model = 0.30·P + 0.25·Pr + 0.25·A + 0.20·S  ·  composite = Σ(weightᵢ · scoreᵢ) renormalized over live models01 · Probe02 · Probe03 · ProbeChatGPT22%Perplexity20%Claude18%Gemini16%DeepSeek12%Grok7%Meta5%Presence30% weightProminence25% weightAccuracy25% weightSentiment20% weightINDEX0–100weighted average
Fig. A · Pipeline. Accent path traces one reading end-to-end — hover any node to re-thread it.
01

How a reading is taken.

For every reading we probe each model the same way, then read four signals off how it responds. The probes themselves, and how we tune them, are part of our calibration and aren't published. What is always open: the four signals below, the weights we apply, the composite math, and, on every reading, the actual answers each model gave, so you can see the evidence your score is built from.

02

Four dimensions, scored end-to-end.

Each model's reading is the weighted sum of four field measurements. The weights sum to 100%.

01
30%

Presence

Did the subject appear at all, across the probes?

Highest when the subject appears across all three answers. The direct answer carries the most weight; the two peer-set answers make up the rest. When a role/field is provided, each direct answer is identity-checked against it — an answer that clearly describes a different person with the same name does not count as the subject's presence.

02
25%

Prominence

When the model mentioned the subject, how early in the response did the name appear?

Higher when the name lands early in an answer rather than buried near the end, averaged across the answers that mention it. Zero if the subject was never mentioned.

03
25%

Accuracy

Did the model give concrete specifics or hedge with uncertainty?

Read from the direct answer by a calibrated LLM judge: it scores how specific and confidently stated the answer is — dates, titles, named organizations versus vague or hedged language. It judges the answer's qualities, not its truth in the world.

04
20%

Sentiment

Was the framing around the mention positive, neutral, or negative?

A calibrated LLM judge reads how the whole direct answer frames the subject — positive, neutral, or negative. 50 is neutral.

03

Every model. Different weight.

Per-model scores roll up into the composite using these product weights. They express our current view of relative discovery importance; they are not an independently audited measure of consumer traffic.

ChatGPT
22%
Perplexity
20%
Claude
18%
Gemini
16%
DeepSeek
12%
Grok
7%
Meta AI
5%

The weights are a published product assumption, not market-share telemetry. We renormalize the composite over whichever model slots are live at scan time, so an unavailable provider never flatters or punishes your reading.

Current API model basket

Family names are interface labels. These are the exact underlying models queried through OpenRouter for this methodology revision.

Family labelAPI modelRoute
ChatGPTopenai/gpt-5.6-lunaopenrouter
Perplexityperplexity/sonaropenrouter
Claudeanthropic/claude-haiku-4.5openrouter
Geminigoogle/gemini-3.5-flash-liteopenrouter
DeepSeekdeepseek/deepseek-v4-flashopenrouter
Grokx-ai/grok-4.3openrouter
Meta AImeta-llama/llama-4-scoutopenrouter
04

The composite, computed.

For each model:

model_score = presence × 0.30
            + prominence × 0.25
            + accuracy   × 0.25
            + sentiment  × 0.20

Then the composite reading:

composite = Σ (model_score_i × model_weight_i)
          / Σ model_weight_i

(renormalized over models live today)
05

What we don’t claim.

  • NoteAccuracy is scored by a calibrated LLM judge, not a fact-check. It reads how concretely and confidently a model describes you — which correlates with the model actually knowing you — and the judge is validated against a hand-labeled rubric set before any methodology change ships.
  • NoteSentiment is scored by the same judge, reading how the whole answer frames you. It cleanly separates positive from neutral from critical framing; sarcasm can still slip past it.
  • NoteIdentity is verified only when you provide a role/field: a direct answer that clearly describes a different person with the same name is excluded from your presence and flagged in the reading. A name-only scan can't be verified — it's marked unverified rather than silently trusted.
  • NoteCanonical facts you save on a subject are checked against each model's answer, and clear contradictions are flagged in the reading's transcript. Findings are receipts, not inputs — they never change the score.
  • NotePer-model scores are stable within a week but not bit-identical. Treat the index as a calibrated reading, not a deterministic score.
  • NoteThe familiar labels name model families, not the consumer products themselves. Consumer apps may add web search, citations, private system prompts, routing, personalization, geography, or conversation history. This benchmark controls the prompts and API basket; it does not claim product parity.
  • NoteA reading is one timestamped sample in English, using the model basket shown above. It is not a confidence interval, a geography-wide survey, or proof of what every user will see. Weekly comparisons are useful for direction; repeated samples are required before claiming a small movement is durable.
  • NoteAll seven slots are attempted on a full reading. If one is temporarily unavailable, its weight is renormalized out of the composite. The transcript and scan time are the receipt for which answers actually contributed.
Cohort-01 · Open

Test weekly monitoring with Cohort-01.

Current test price: $4.99/mo. It is not our validated standard price. No card now.

Calibrate your own reading

Run the procedure on yourself.

One free reading. Each model's verbatim answer about you is yours to inspect.

Run a free scan ↗