Day 110: Ask What Your GEO Baseline Refuses to Promise
A GEO baseline should be useful before it sounds ambitious.
That sounds obvious until procurement begins. The buyer asks for a diagnostic. The supplier wants the work to feel valuable. The proposal starts to stretch. It promises visibility improvement, citation opportunity, ranking movement, revenue upside, competitor displacement, or a clean score that leadership can track. The language becomes easier to approve because it sounds closer to growth.
It also becomes less credible.
For CMOs, Marketing Directors, and founders, the safer buying question is not only, “What will this baseline deliver?” It is also, “What will this baseline explicitly refuse to promise?”
A serious diagnostic can promise observable work: the question window, surfaces tested, retained captures, source labels, access limits, technical inspection, prioritised diagnosis, and a recommended next move. It should not promise deterministic answer-engine outcomes from a snapshot.
The first red flag is a missing refusal
A baseline is a procurement instrument before it is a growth story.
Its job is to help leadership decide whether AI-assisted buyer research is creating a material visibility, positioning, source-quality, competitor, or technical retrievability problem. It can reduce uncertainty around what to fix first. It can show where the public record is strong, thin, stale, inaccessible, or being interpreted in ways the company did not intend.
But it is still a bounded diagnostic.
It observes answer-led surfaces under recorded conditions. It may inspect ChatGPT, Claude, Perplexity, Gemini, Google AI features, search results, citation surfaces, directories, review sites, or public source material. Those observations vary by surface, prompt, date, market, language, account state, visible source availability, freshness, and access path.
That does not make the work weak. It makes the boundary part of the product.
If a proposal says what will be checked but not what the supplier will refuse to infer, the buyer is being asked to approve the most fragile part of the work on faith.
A generalised pre-purchase scenario
Imagine a Marketing Director preparing to buy a first GEO baseline.
Sales has heard that prospects arrive with strange assumptions. Some buyers think the company is a monitoring platform. Some think it is a content agency. Some mention competitors the team does not usually consider direct substitutes. Leadership wants to know whether answer-led research is shaping those expectations before the first call.
Two suppliers respond.
The first promises a visibility score, citation opportunities, technical fixes to improve answer-engine inclusion, and a forecast of likely commercial upside. The scope looks reassuring because it turns uncertainty into a tidy growth story.
The second starts with narrower language. It says the baseline will define the buyer-question window, record direct answer captures where access allows, label search or citation-surface proxies separately, retain prompts and timestamps, inspect visible sources where available, check whether important public material is accessible and well structured, and return a ranked diagnosis with a recommended next decision. It also says what the baseline will not claim: guaranteed citations, platform-wide ranking movement, causal attribution from technical hygiene, a universal AI visibility score, or revenue forecast from one snapshot.
The second proposal may sound less exciting.
It is usually easier to approve, defend, and learn from because the buyer knows which claims are inside the work and which claims are outside it.
Use the promise and refusal checklist
A buyer does not need a technical data room before the first call. They do need enough contract language to separate a diagnostic from speculation.
A compact checklist can do that.
| The baseline should promise | The baseline should refuse | Why it matters | What to ask for |
|---|---|---|---|
| A declared question and surface window: buyer roles, prompts, markets, dates, access conditions, and answer surfaces. | A claim that the output represents all buyer research, all markets, or every answer-engine context. | Leadership needs to know what was actually observed before treating the result as management evidence. | “Which buyer situations, surfaces, access conditions, and dates are in scope?” |
| Retained observations: prompts, captures, timestamps, visible sources where available, and practical limitations. | A polished summary that cannot be traced back to what was captured. | The buyer should be able to review the observation, not only the supplier’s interpretation. | “Will we receive the retained captures and limitation notes?” |
| Clear labels for direct answer captures versus search results, citation-surface proxies, source indexes, or technical inspections. | Proxy evidence presented as direct answer-engine visibility. | Search-side visibility can be useful, but it is not the same as a captured answer from the tested surface. | “Which findings are direct answer captures, and which are proxies?” |
| Technical retrievability review: crawl access, semantic structure, metadata, canonical signals, machine-readable surfaces where present, and answer-ready source quality. | A claim that technical hygiene proves why a surface did or did not cite the brand. | Technical work can make important material easier to find and reuse; it rarely proves causality on its own. | “Which technical findings are contributing conditions rather than proven causes?” |
| Prioritised diagnosis tied to the buyer decision: fix, test, monitor, park, or investigate further. | A universal score that hides different findings under one number. | Absence, weak source quality, competitor displacement, misdescription, and technical access problems deserve different actions. | “What decision will this result make safer, and what should we not fund yet?” |
| Explicit outcome boundaries. | Guaranteed citations, guaranteed rankings, deterministic answer control, forecast revenue, or claimed buyer behaviour from the snapshot. | Overpromised outcomes create false ROI expectations and make later work harder to evaluate. | “Which claims will you put in writing that the baseline will not make?” |
The last question is the strongest one.
A supplier that can name its refusal boundaries is not lowering ambition. It is protecting the buyer from treating early observations as proof of outcomes no one can control.
Boundaries make follow-on work easier to judge
Promise boundaries are not only legal caution. They make the work after the baseline easier to contract, approve, and challenge.
If the report says a brand was absent from a defined set of high-intent questions under recorded conditions, the next action may be to improve the offer route, comparison material, source quality, or category language. If the report says a search or citation-surface proxy suggests a source pattern, the next action may be deeper direct capture rather than a public rewrite. If the technical review finds important proof assets blocked, thin, duplicated, or hard to parse, the next action may be retrievability improvement without pretending the fix guarantees citations.
Each action remains attached to the type of observation that justified it and the claims the baseline refused to make.
That makes procurement cleaner. The buyer can approve a contained diagnostic without being forced to believe a revenue forecast. The supplier can recommend further work without disguising uncertainty. Leadership can decide whether to fix, test, monitor, or hold spend with fewer hidden assumptions.
It also protects the relationship. A buyer who was promised guaranteed answer-engine visibility will judge every variable result as failure. A buyer who was promised retained observations, labelled limits, and a ranked next decision can judge whether the diagnostic did its job.
Keep Google and answer-engine claims ordinary
This is especially important when Google AI features or technical tactics enter the conversation.
Google AI features rely on core Search ranking and quality systems. If a baseline observes a weak Google-visible result, the appropriate response is to improve usefulness, relevance, clarity, accessibility, source quality, and public evidence where the finding supports that work. It is not to claim that llms.txt, special AI markup, arbitrary chunking, or over-focused structured data are required switches for Google AI visibility.
The same restraint applies across answer-led surfaces. ChatGPT, Claude, Perplexity, Gemini, Google AI features, search-linked summaries, directories, review sites, and publisher pages can all expose useful signals. They do not give the supplier deterministic control over what future buyers will see.
A credible baseline should therefore speak in observable verbs:
- captured;
- labelled;
- compared;
- inspected;
- prioritised;
- recommended;
- excluded.
It should avoid outcome verbs it cannot own:
- guaranteed;
- forced;
- proved;
- attributed;
- forecast;
- controlled.
The difference is not stylistic. It is the difference between buying a diagnostic and buying a promise the market cannot honour.
Ask for the refusal before you buy the report
The procurement test is simple.
Before approving a GEO baseline, ask the supplier to write two lists.
First: what the baseline will deliver. The answer should name the question window, tested surfaces, retained captures, source and proxy labels, technical retrievability review, prioritised diagnosis, and next-step recommendation.
Second: what the baseline will not claim. The answer should refuse guaranteed citations, guaranteed rankings, deterministic answer control, causal attribution from technical hygiene alone, universal scores, forecast revenue, and buyer behaviour claims from a snapshot.
If the second list is missing, the first list is less useful than it looks.
For CMOs, Marketing Directors, and founders, the value of the baseline is not that it makes AI visibility sound certain. It is that it gives the company a bounded, retained, decision-useful view of what answer-led surfaces appear to show now, what the public record can support, what technical conditions may be limiting reuse, and what work deserves the next pound of attention.
A credible baseline does not pretend to control the answer layer.
It tells you what was observed, what that observation can support, what it cannot support, and what to do next.
Ask for the promise.
Then ask for the refusal.