A real measurement, shown in full
A real category, measured July 16, 2026: real vendors, 32 dated runs, the exact prompts printed. And every raw answer archived and published.
- Category
- Contract management software
- Measured
- 2026-07-16
- Engines
- ChatGPT · Google AI Mode
- Buyer prompts
- 4
- Runs
- 20 + 12 = 32
- Location
- United States
Why contract management? Recognizable vendors in a crowded field, not picked for drama. Your category may look better or worse; that’s what measuring is for.
Who gets named, the same question, five fresh sessions
The first buyer prompt, run 5 times on ChatGPT. Filled cells are runs that named the vendor; count them:
Each square is one run, left to right.
Icertis appears in runs 2 and 3 only. Gatekeeper (not shown) appears in run 1 only.
measured 2026-07-16 · n=5 · ChatGPT
how we run this →Each square is one run, left to right.
A different question flips the winners: Spellbook leads here and is absent above.
measured 2026-07-16 · n=5 · ChatGPT
how we run this →The leaderboard. Each engine on its own
20 ChatGPT runs and 12 Google AI Mode runs, shown separately. Several vendors live almost entirely on one engine, which is exactly what one combined number would hide. Flips count each vendor’s mixed checks across its 8 checks, one buyer prompt on one engine, rerun identically.
A check reads as mixed unless every rerun in it agrees, which happens by chance at any hit rate between never and always. So flips are counted against that baseline. Across this table 31 were observed where chance predicts about 75.5. Rerunning is worth doing because one ask gives you a single draw, not because this category is unusually unstable.
| Vendor | ChatGPT | Share | AI Mode | Share | Stability |
|---|---|---|---|---|---|
| Ironclad | 16/20 | 80 percent of ChatGPT runs | 8/12 | 67 percent of AI Mode runs | 3 of 8, chance predicts 5.3 |
| DocuSign * | 14/20 | 70 percent of ChatGPT runs | 5/12 | 42 percent of AI Mode runs | 3 of 8, chance predicts 6.6 |
| Juro | 13/20 | 65 percent of ChatGPT runs | 5/12 | 42 percent of AI Mode runs | 3 of 8, chance predicts 6.7 |
| LinkSquares | 12/20 | 60 percent of ChatGPT runs | 6/12 | 50 percent of AI Mode runs | 3 of 8, chance predicts 6.7 |
| Agiloft | 12/20 | 60 percent of ChatGPT runs | 4/12 | 33 percent of AI Mode runs | 3 of 8, chance predicts 6.8 |
| Icertis | 12/20 | 60 percent of ChatGPT runs | 0/12 | 0 percent of AI Mode runs | 3 of 8, chance predicts 6.4 |
| Conga | 11/20 | 55 percent of ChatGPT runs | 3/12 | 25 percent of AI Mode runs | 2 of 8, chance predicts 6.7 |
| SpotDraft | 10/20 | 50 percent of ChatGPT runs | 4/12 | 33 percent of AI Mode runs | 3 of 8, chance predicts 6.7 |
| PandaDoc | 6/20 | 30 percent of ChatGPT runs | 8/12 | 67 percent of AI Mode runs | 3 of 8, chance predicts 6.7 |
| Spellbook | 5/20 | 25 percent of ChatGPT runs | 3/12 | 25 percent of AI Mode runs | stable |
| Gatekeeper | 2/20 | 10 percent of ChatGPT runs | 5/12 | 42 percent of AI Mode runs | 3 of 8, chance predicts 4.9 |
Sorted by ChatGPT share, engines are never combined into one number (20 vs 12 runs would quietly weight one). Listed: vendors appearing in ChatGPT’s structured entity lists in ≥2 runs. Vendors the engines only mention in prose (e.g. Sirion) aren’t ranked. They’re auditable in the raw answers. * DocuSign includes runs naming its CLM product line.
measured 2026-07-16 · 4 prompts × (5 ChatGPT + 3 AI Mode) runs · United States
The engines disagree. And an average would hide it
Icertis
60% ChatGPT · 0% AI Mode
An enterprise CLM leader, invisible on one engine. Averaged: “30%”, describing neither.
PandaDoc
30% ChatGPT · 67% AI Mode
The mirror image, strong exactly where Icertis is missing.
The sources AI reads to build these answers
The 12 most-cited domains (of 71 recorded) across the 32 runs, the map for where authority is actually earned in this category:
| Domain | Cited in | Share |
|---|---|---|
| linksquares.com | 11/32 | |
| juro.com | 9/32 | |
| g2.com | 9/32 | |
| docusign.com | 8/32 | |
| reddit.com | 8/32 | |
| ironcladapp.com | 8/32 | |
| sirion.ai | 8/32 | |
| hyperstart.com | 6/32 | |
| gatekeeperhq.com | 6/32 | |
| gartner.com | 6/32 | |
| pandadoc.com | 6/32 | |
| youtube.com | 6/32 |
Both layers matter: vendors’ own pages are cited heavily. And so are surfaces no vendor controls (YouTube, Reddit, G2, analyst pages). The work targets both.
What five runs can, and can’t, tell you
Enough to separate consistently named from absent from flipping; not enough to make “40% vs 60%” a ranking, one run moves it 20 points. Client baselines run more repetitions, and movement only counts once it clears that noise.
What a client report adds on top
- Your prompt set, built in your buyers’ words and frozen for comparability
- Month-over-month movement, dated both times, the scoreboard for the authority work
- The gap analysis: which sources name your competitors but not you, in priority order
- The prioritized 90-day plan the sprint executes
Check this report right now. Paste this into ChatGPT.
best contract management software for mid-market companies
The list will move between runs. That movement is the finding.
Get your free breakdown
A run is one question, asked once, in a clean session. We ask your buyers’ questions on ChatGPT and on Google’s AI Mode, more than once on each, then send you who got named in every answer with the exact prompts.
That didn’t go through, please check the required fields and send again.
What arrives, exactly
- Engines
- ChatGPT and Google AI Mode
- Runs
- 3 ChatGPT and 2 AI Mode per question, each a fresh session
- You get
- Every answer, who was named in it, and the exact prompts
- Arrives
- Within 2 business days, from a person
- Cost
- None, and no call required
Prefer to talk first? Book a walkthrough (opens in a new tab). Your category on screen, no deck. Or read the real sample report.