Methodology
How we measure your brand's AI visibility
Lumidian queries each AI model directly and measures how often your brand appears in their responses. No scraping, no proxies: the same APIs that power ChatGPT, Claude, Perplexity, and Gemini.
Transparency matters. If you're going to act on a visibility score, you should know exactly how it's calculated, what it represents, and what can move it. This page explains every part of the process, and the key design decisions are backed by published research, cited inline and listed in full in the references below.
Models
How we measure visibility
Every model we query answers against the live web. Each one reaches for sources differently, so we treat them as four independent readings of what the web says about you today. That independence is measurable. In our own study, 900 runs of the same 30 questions on ChatGPT, Gemini, and Perplexity, any two surfaces agreed on the top answer only 20 to 23 percent of the time (Jaccard 0.20 to 0.23).[9] An independent 2026 study across many categories found a similar split: models agree on the top-recommended brand 41.6% of the time.[1] A brand's standing on one surface says little about the others, so tracking a single model misses most of the picture.
ChatGPT
OpenAI
OpenAI's native web search retrieves live results before answering, matching what users see in ChatGPT today.
Claude
Anthropic
Claude Haiku 4.5 with Anthropic's web-search tool. Issues up to three targeted searches per query before responding.
Perplexity
Perplexity
Search-grounded by design. Every answer is built from sources pulled at request time, with inline citations.
Gemini
Google Search grounding is on for every query, so answers reflect current web context rather than static knowledge.
Scoring
How your visibility score is calculated
One number that tells you how visible your brand is across AI. Here's exactly what goes into it.
Prompts sent to each model
Your tracked prompts are sent to each AI model's API: ChatGPT, Claude, Perplexity, and Gemini. Each prompt is run multiple times per model because LLMs give different answers to the same question. Peer-reviewed studies show outputs vary even at "deterministic" settings, so a single response is not a reliable measurement.[2],[3],[4]
Responses checked for mentions
Every response is analyzed for your brand name using both an exact case-insensitive match and a fuzzy normalized check that strips non-alphanumeric characters. A mention is detected if either method finds your brand.
Score calculated
Your visibility score is the percentage of queries where your brand was mentioned. Mention rate over repeated prompts is the measure independent research converges on: exact AI answers almost never repeat, but a brand's mention rate is stable across runs, which is also why we don't sell an "AI ranking position" metric.[6],[7]
Score = (mentions / total queries) × 100
Per-model score is calculated this way for each AI model individually. Your overall visibility score is the average across the models tracked for your brand. Errors excluded
If a query fails due to an API error or timeout, it's excluded from the denominator entirely. Your score only reflects responses that were actually received and analyzed.
Approach
Why we query models directly
We query each AI model's API directly, the same models that power ChatGPT, Claude, Gemini, and Perplexity.
Controlled, comparable conditions
Direct API queries eliminate variation from account state, location, cookies, and session history, so every run measures the model, not your browser. The variation that remains is the model's own response randomness, which no caller can switch off.[3],[4] That's what the repeated runs are for.
Every answer comes from today's web
Every model we query runs against the live web: ChatGPT's native search, Claude's web-search tool, Perplexity's search grounding, and Gemini's Google Search integration. Scores reflect the web as it exists today, not a frozen snapshot from a model's training run.
Stable over time
Some tools scrape consumer chat interfaces, but those results vary by session and break when UIs change. Direct API queries give you stable, comparable scores that you can trend with confidence.
Improving
What moves your score
Every model we query is doing the same thing under the hood: searching the web for sources that answer the prompt, then composing an answer from what it finds. This is measurable and moveable: the foundational peer-reviewed study on generative engine optimization found that targeted content changes boosted a source's visibility in AI answers by up to 40%.[8] Four things consistently show up in the sources that get cited.
Fresh web content that answers the prompt
Reddit threads, Quora answers, and recent articles that mention your brand in the context of what the prompt is actually asking. Relevance to the question beats generic brand mentions every time.
Authority on sources AI search weights heavily
Wikipedia, major publications, and industry-specific subreddits that reliably surface in grounded searches. A mention on a domain the models already trust moves the needle.
Repeat mentions across independent sources
One mention on one site is easy to pass over. Three independent sources corroborating the same claim is much harder to ignore. That’s when models start treating it as the default answer.
Prompt-term adjacency
Your brand name appearing near the prompt’s core keywords on the source page. Proximity is how search-grounded models decide which mentions are relevant to the question being asked.
References
The research behind this methodology
The sources cited above, in full. We label each one honestly: peer-reviewed papers passed independent academic review; preprints and industry studies haven't, but publish their data and methods openly. Our own measurement is labeled as such: one vertical, descriptive, and reviewed by nobody but us. No study validates our exact run count. The research supports measuring over repeated runs as a practice, and three runs is where we balance statistical stability against querying cost.
- [1]Preprint
Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models (opens in a new tab)
Żatuchin, D. · arXiv:2606.23057, 2026
Found only 41.6% agreement between models on the top-recommended brand. A top spot on one model doesn’t carry to another, so tracking a single model gives an incomplete picture.
- [2]Peer-reviewed
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism (opens in a new tab)
Song, Y., et al. · NAACL 2025
Shows that judging an LLM from a single response per prompt is unreliable; the paper itself samples each prompt many times and reports averages across runs.
- [3]Peer-reviewed
Non-Determinism of "Deterministic" LLM Settings (opens in a new tab)
Atil, B., et al. · Eval4NLP @ ACL 2025
Across 5 LLMs, 8 tasks, and 10 runs each, accuracy varied up to 15% between identical runs. Even at temperature 0 with fixed seeds, no model produced repeatable outputs.
- [4]Peer-reviewed
Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference (opens in a new tab)
Yuan, J., et al. · NeurIPS 2025 (oral)
Traces run-to-run variation to the inference infrastructure itself (floating-point non-associativity, GPU batching), which callers of commercial LLM APIs cannot switch off.
- [5]Peer-reviewed
On the Reproducibility of LLM-Centric Empirical Studies (opens in a new tab)
Angermeir, F., et al. · ICSE 2026
A replication of 85 LLM studies ran up to 30 repetitions per study specifically because of non-determinism, observing up to 30% metric differences between repetitions.
- [6]Industry study
AIs Are Highly Inconsistent When Recommending Brands: Marketers Should Take Care When Tracking AI Visibility (opens in a new tab)
Fishkin, R. (SparkToro) & O’Donnell, G. (Gumshoe.ai) · SparkToro Research, 2026
Across 2,961 runs of 12 prompts, exact brand lists almost never repeated, yet per-brand mention rates stayed stable over many runs, concluding that visibility % across many prompts run multiple times is a sound metric, while "AI ranking position" metrics are not.
- [7]Preprint
Don’t Measure Once: Measuring Visibility in AI Search (opens in a new tab)
Schulte, B., Bleeker, F. & Kaufmann, E. · arXiv:2604.07585, 2026
Repeated identical AI-search queries shared only 32–43% of cited sources; brand visibility should be measured as a distribution over repeated queries, not a one-off observation.
- [8]Peer-reviewed
GEO: Generative Engine Optimization (opens in a new tab)
Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. & Deshpande, A. · KDD 2024
The foundational generative-engine-optimization study (Princeton/IIT Delhi): targeted content changes (adding citations, quotations, and statistics) boosted source visibility in generative engine responses by up to 40% on a large multi-domain benchmark.
- [9]Our measurement
Repeatability of AI shopping-agent answers: 30 queries, 5 repeats, 3 surfaces, two collection days
Turner, K. (Lumidian) · Lumidian working paper, September 2026 (Study 0; write-up on request)
900 API runs of 30 headphone queries on ChatGPT (GPT-5.6 with web search), Gemini 2.5 Pro (Search grounding), and Perplexity (Sonar Pro), collected on two days nine days apart. Within-day repeatability of the top-3 answer: Perplexity 0.80, ChatGPT 0.62, Gemini 0.45 (median Jaccard). Agreement between surfaces on the same query: 0.20 to 0.23. Constrained questions (a budget, a use case) were not stable on any surface. One vertical, descriptive, not peer-reviewed.