AI Visibility Index · August 2026

LLM observability (tracing, evaluation and monitoring for LLM/agent apps): who the AI engines actually name

As of August 21, 2026, across 47 AI answers (ChatGPT · Google AI Overviews · Google AI Mode · Gemini) to 4 buyer prompts, Langfuse was named most often (46 mentions). 13 vendors were named at all; only 4 by every engine that answered. Of the 427 sources the engines cited, 33% were vendor-owned pages.

Category
LLM observability (tracing, evaluation and monitoring for LLM/agent apps)
Measured
2026-08-21
Engines
ChatGPT · Google AI Overviews · Google AI Mode · Gemini
Buyer prompts
4
Answers
47
Source
CrediGeo measured AI visibility for LLM observability (tracing, evaluation and monitoring for LLM/agent apps) in 2026. Across 47 AI answers recorded on 2026-08-21, Langfuse was named most often (46 mentions).

Google AI Overviews appeared in 3 of 8 checks, for the other checks, a normal Google search in this category showed no AI answer at all. Its shares below are of the answers that appeared.

The leaderboard. Each engine on its own

Vendors named in at least one answer, sorted by total mentions. Engines are never combined into one number. Stability counts flips, a prompt naming the vendor in some reruns but not others, across each vendor’s checks:

# Vendor ChatGPTAI OverviewsAI ModeGemini Flips observed
1 Langfuse 20/20 3/3 12/12 11/12 1 of 16, chance predicts 1
2 LangSmith 20/20 3/3 11/12 11/12 2 of 16, above the 1.9 chance predicts
3 Arize AI (Phoenix / AX) 20/20 3/3 11/12 9/12 3 of 16, chance predicts 3.5
4 Braintrust 20/20 0/3 11/12 6/12 3 of 16, chance predicts 7.3
5 Datadog LLM Observability 18/20 0/3 6/12 2/12 3 of 16, chance predicts 10.4
6 Helicone 13/20 0/3 6/12 3/12 1 of 16, chance predicts 10.5
7 Confident AI (DeepEval) 1/20 0/3 7/12 5/12 7 of 16, chance predicts 8.6
8 Comet Opik 3/20 3/3 1/12 4/12 5 of 16, chance predicts 7.8
9 Weights & Biases Weave 9/20 0/3 2/12 0/12 3 of 16, chance predicts 7.8
10 Galileo 0/20 0/3 1/12 4/12 3 of 16, chance predicts 4.3
11 Portkey 2/20 0/3 0/12 0/12 1 of 16, chance predicts 1.9
12 Traceloop (OpenLLMetry) 2/20 0/3 0/12 0/12 1 of 16, chance predicts 1.9
13 Pydantic Logfire 0/20 0/3 0/12 1/12 1 of 16, chance predicts 1

Mentions, not quality. This table answers “who is on the AI shortlist”, never “who is best.” Cells read “named / answers shown” per engine. Vendors not named in any answer are not listed, see the FAQ for what that does and doesn’t mean. Product aliases and sub-brands are matched (the alias set ships in the raw data). Vendor names that link out have a profile of their own, showing every category we measured them in.

Flips are counted against what sampling alone would produce at each vendor’s own hit rate. Across this table, 34 flips were observed where chance predicts about 67.9. Reruns are worth doing because a single ask gives you one draw rather than an answer. That is a different claim from saying this category is unusually unstable, and on this measurement it isn’t.

CrediGeo · measured 2026-08-21 · 4 prompts · ChatGPT ×20 · AI Overviews ×3 · AI Mode ×12 · Gemini ×12 · United States

The sources behind the answers, labeled for ownership

The domains the engines cited while building these answers. Each is labeled by who owns it, because a “ranking” cited from a vendor’s own site is a different kind of evidence than an independent one. 33% of all 427 citations in this category came from vendor-owned pages:

Domain Ownership Cited ChatGPTAI OverviewsAI ModeGemini
braintrust.dev Vendor-owned · Braintrust 36× 50724
google.com Other 30× 00300
langchain.com Vendor-owned · LangSmith 23× 10085
mlflow.org Other 21× 03117
confident-ai.com Vendor-owned · Confident AI (DeepEval) 21× 00417
arize.com Vendor-owned · Arize AI (Phoenix / AX) 18× 9054
langfuse.com Vendor-owned · Langfuse 16× 10321
getmaxim.ai Other 16× 00160
medium.com Other 15× 00123
openobserve.ai Other 2340
laminar.sh Other 2340
latitude.so Other 0090
morphllm.com Other 4004
signoz.io Other 1007

Ownership rule: a domain is “vendor-owned” when it belongs to a vendor in this category’s alias set; review platforms and community/video surfaces are labeled as such; everything else is “other.” We label. We don’t exclude: the raw mention counts above include all sources.

What 47 answers can, and can’t, tell you

Enough to separate consistently-named from absent from flipping on August 21, 2026; not enough to make small share differences a ranking, one run moves a vendor’s number noticeably. AI answers move continuously. This page is a dated snapshot, re-measured monthly, not a permanent verdict. We sell the measurement discipline, not placements, and we make no promise about who appears here.

Method, conflicts, corrections

The full method is public, including prompt design, fresh sessions, alias matching and what we refuse to claim. See how we check. The exact prompts used here are in the published raw data.

Conflicts: CrediGeo sells AI-visibility measurement and earned-authority work. We do not sell, resell, or take referral fees from any vendor in this category, and no vendor paid to appear or influence placement. Corrections: if you believe a product was missed by an alias gap, write to info@credigeo.com. We re-check against the raw runs and correct publicly if we got it wrong.

Your product in this category? Get your own breakdown. Which prompts name you, which don’t, per engine. Free, within 2 business days.

Request your breakdown

Questions about this index

What is AI visibility for LLM observability (tracing, evaluation and monitoring for LLM/agent apps)?
AI visibility is how often an AI assistant names a product when buyers ask which LLM observability (tracing, evaluation and monitoring for LLM/agent apps) to use. When ChatGPT, Google's AI or Gemini answers "best LLM observability (tracing, evaluation and monitoring for LLM/agent apps)", it names a short list of specific vendors. This page measures who actually appears on that list, answer by answer.
How is this measured?
We ran 4 real buyer prompts repeatedly on each engine (ChatGPT · Google AI Overviews · Google AI Mode · Gemini) on August 21, 2026, in fresh sessions, and counted which vendors each answer named, 47 answers in total. Every raw answer is published and the method is public.
Why don’t the engines agree?
They read different sources, studies show different AI engines rarely cite the same pages. That per-engine split is the finding: of the 13 vendors named at all here, only 4 were named by every engine that answered. A single blended score would hide exactly that.
Why does Google AI Overviews have fewer answers than the other engines?
AI Overviews only appears for some queries. In this measurement it appeared in 3 of 8 checks. That trigger rate is itself part of the data: it tells you how often buyers even see an AI answer on a normal Google search in this category.
Is this a quality ranking?
No. This measures MENTIONS, how often an engine names a vendor, not product quality. A great product can be rarely named and a mediocre one often. It answers "who is winning the AI shortlist", not "who should you buy".
A vendor is missing, what does that mean?
Only vendors named in at least one of our 47 answers on August 21, 2026 for these 4 prompts appear here. Not appearing means our runs didn’t name them on that date with these prompts. Nothing more. If it’s your product, request your own run-by-run breakdown (free) and we’ll show you exactly what we asked and what came back.

Get your free breakdown

A run is one question, asked once, in a clean session. We ask your buyers’ questions on ChatGPT and on Google’s AI Mode, more than once on each, then send you who got named in every answer with the exact prompts.

Free, and no call required. In your inbox within 2 business days.

Already got our email? Replying to it reaches the same person, and it’s the fastest route.

What arrives, exactly

Engines
ChatGPT and Google AI Mode
Runs
3 ChatGPT and 2 AI Mode per question, each a fresh session
You get
Every answer, who was named in it, and the exact prompts
Arrives
Within 2 business days, from a person
Cost
None, and no call required

Prefer to talk first? Book a walkthrough (opens in a new tab). Your category on screen, no deck. Or read the real sample report.