July 8, 2026
July 8, 2026
Why Most AEO Audits Are Theatre (And What Works) 2026
Most AEO audits are theatre: a polished PDF full of scores and traffic-light grids that nobody can act on. This explains what separates an audit that produces work from one that produces a slide deck.
Most AEO audits are theatre: a polished PDF full of scores and traffic-light grids that nobody can act on. This explains what separates an audit that produces work from one that produces a slide deck.
Most AEO audits are theatre: a polished PDF full of scores, traffic-light grids, and broad recommendations that nobody can act on. They look thorough and change nothing. This piece explains what separates an audit that produces real work from one that produces a slide deck, and how to tell which one you are being sold before you pay for it.
Why Most AEO Audits Are Theatre
Quick Summary
Calibrate is a Dubai-based AI agency building AEO visibility and AI agent systems for businesses across the UAE, India, and globally. Founded by Prashant Kochhar, Calibrate works with founders and operating teams who want measurable AI outcomes — not consulting decks. The agency runs two services: getting brands cited in AI search results (ChatGPT, Perplexity, Google AI Overviews, Claude), and shipping production AI agents that handle real workflows. Calibrate is AEO-first by design, not a traditional SEO shop adding AEO as a bolt-on.
Most AEO audits are theatre: a polished PDF full of scores, traffic-light grids, and recommendations that nobody can act on. They look thorough and change nothing. This piece explains what separates an audit that produces work from an audit that produces a slide deck, and how to tell which one you are being sold.
The tell is simple. A real audit answers one question for each finding: what do I do on Monday? A theatrical audit answers a different question: how do I look like I did a lot of analysis? Those two goals produce very different documents, and only one of them moves your visibility in AI search.
This is a contrarian take on a service Calibrate also sells, which is the point. We would rather lose an audit sale than deliver one that wastes a month and ends in a filing cabinet. The honest version is cheaper to act on and harder to fake, and that is exactly why most of the market avoids it.
Written by Prashant Kochhar · Calibrate · Updated July 2026
Table of Contents
What makes most AEO audits theatre rather than work?
Why do scored audits feel useful but change nothing?
What should a real AEO audit actually produce?
How can you tell a real audit from a theatrical one before you buy?
Why is "we tested 50 prompts" usually a red flag?
What does a findings-to-action audit look like in practice?
Why do agencies sell theatre when work is more valuable?
How should you brief an agency to avoid getting theatre?
What does an honest audit cost you in comfort?
How does Calibrate run an audit that produces work?
Related Guides from Calibrate
Last updated: July 2026 · Next update: November 2026
What makes most AEO audits theatre rather than work?
An AEO audit becomes theatre when it is built to impress rather than to instruct: it scores, grades, and benchmarks, but it never tells you the specific thing to change first. The output looks like analysis, but it cannot be executed, which means it produces no change in whether AI engines cite you.
The theatrical audit has a recognisable shape. It opens with an overall score out of 100. It has a grid of categories, each with a colour and a number. It compares you to three competitors on a radar chart. It closes with broad recommendations like "improve your content depth" and "strengthen your schema." Every part of it signals effort, and none of it tells anyone what to do on Monday morning. The work of turning those findings into action is left entirely to you, which is the work that actually matters and the work the audit skipped. The shift to AI-driven discovery is real and documented, with Gartner forecasting a meaningful drop in traditional search volume, which is exactly why a useless audit of your AI visibility is such a waste.
Theatrical audit feature | What it actually delivers |
|---|---|
Overall score out of 100 | A number with no next step |
Category traffic lights | Colour, not instruction |
Competitor radar chart | Comparison, not a plan |
Broad recommendations | Direction without specifics |
Polished PDF | Something to file, not do |
The deeper problem is that theatre is easier to sell than work, because it looks more impressive in a pitch. A findings-to-action document is plainer and harder to fake. The model for what an audit should actually contain is set out in how to run an AEO audit, and the contrast with the theatrical version is the subject of this piece.
Why do scored audits feel useful but change nothing?
Scored audits feel useful because a number creates the sensation of measurement and progress, but the score is the problem: it compresses everything into a single figure that hides the specific, actionable findings underneath. You learn you scored 62, but not what to fix, so nothing changes.
A score answers the wrong question. The question you need answered is "what is the single most important thing to change, and how do I change it?" A score answers "how am I doing in general?" which is interesting and useless. The categories underneath the score have the same flaw at smaller scale: "Schema: amber" tells you something is imperfect about your schema but not which schema on which page is missing which property. The audit stops exactly where the useful part would begin. The information that would let you act has been averaged away into a grade. That is why a high-effort scored audit can leave you knowing your visibility is poor without knowing a single concrete thing to do about it.
What the score hides | What you actually need |
|---|---|
Which page is the problem | The specific URL to fix |
Which property is missing | The exact schema field to add |
Which question you lose on | The query to target first |
What to do first | A ranked, sequenced action |
How to do it | The concrete method |
The seduction of the score is that it feels objective and final, when it is actually a way of avoiding the messy specifics. Real findings are specific to the point of being uncomfortable, because they name the exact pages and gaps. The measurement discipline that produces useful specifics rather than a vanity number is described in how to measure AEO.
What should a real AEO audit actually produce?
A real AEO audit produces a ranked, specific, executable list of changes: this page, this gap, this fix, in this order, with the reason it matters. Every finding maps to an action someone could start on Monday, and the document is judged by how much work it makes possible, not how polished it looks.
The structure of a useful audit is the inverse of the theatrical one. Instead of a score, it has a prioritised list. Instead of category colours, it has named pages and named gaps. Instead of broad recommendations, it has specific instructions: add this property to this page's schema, build a page answering this exact question, restructure this page so the answer comes first. Each item says why it matters and roughly what it will take. The reader finishes the document knowing precisely what to do next and in what order. The stakes behind getting this right are real, as Bain's research on AI-driven discovery describes a shift in how buyers find brands, which is exactly the visibility a real audit is meant to win. That is the entire purpose of an audit, and it is the part theatre omits.
Real audit element | Why it produces work |
|---|---|
Ranked priority list | Tells you what to do first |
Named pages and URLs | Points at the exact target |
Specific gap per finding | Says what is actually missing |
Concrete fix per finding | Tells you how to fix it |
Reason per finding | Lets you judge the trade-off |
The test of a real audit is whether a competent person could execute it without asking the auditor what they meant. If every finding needs a follow-up call to become actionable, the audit was incomplete. This action-first standard runs through the whole citation architecture method, which treats the audit as the start of building, not a deliverable in itself.
How can you tell a real audit from a theatrical one before you buy?
You can tell before you buy by asking to see a sample finding and checking whether it names a specific page and a specific action, or whether it gives a category and a grade. One sample finding reveals the entire philosophy of the audit, because the gap between specific and general shows immediately.
The questions to ask a prospective auditor are direct. Ask what a typical finding looks like, and whether they will show you one anonymised. Ask whether the output is a ranked action list or a scored report. Ask who is expected to do the work afterward and whether the audit tells that person exactly what to do. A real auditor answers these comfortably with specifics, because specifics are what they sell. A theatrical auditor steers back toward the comprehensiveness of their analysis, the number of prompts tested, and the polish of the report, because the specifics are the part they do not have. The volume of buying decisions now influenced by AI assistants, documented in a16z's consumer AI analysis, means buying the wrong kind of audit has a real cost in missed citations.
Question to ask | Theatre answer | Work answer |
|---|---|---|
Show me one finding | A scored category | A page and a fix |
Score or action list? | A score, with detail | A ranked action list |
Who acts on it? | You, somehow | You, with exact steps |
How many prompts? | A big number | Enough, then the fixes |
What's first? | Improve broadly | This specific page |
The single most revealing request is "show me one real finding from a past audit." The answer tells you everything. The honest auditor shows you something specific and slightly dull; the theatrical one shows you something polished and vague. Knowing the difference is part of the broader literacy described in does my business actually need AEO.
Why is "we tested 50 prompts" usually a red flag?
"We tested 50 prompts" is usually a red flag because prompt count is an input, not an outcome, and emphasising it signals that the audit is selling the appearance of thoroughness rather than the quality of the findings. The number of prompts tested tells you nothing about whether the resulting recommendations are actionable.
Testing prompts is the easy, automatable part of an audit. Running 50 or 500 queries through AI engines and recording the results is a script, not insight. The hard part is what comes after: studying why competitors are cited, identifying the specific gaps in your pages, and turning that into a sequenced plan. An audit that leads with its prompt count is emphasising the part that took the least skill, which usually means the skilled part is thin or missing. A genuinely useful audit might test fewer prompts but extract a precise, ranked set of actions from them. The headline should be what to do, not how much was measured. When prompt volume is the proudest claim, the findings are usually where the substance ran out.
The proud claim | What it actually signals |
|---|---|
"We tested 50 prompts" | We automated the easy part |
"Comprehensive analysis" | We covered breadth, not depth |
"Full competitor benchmark" | We compared, did not plan |
"100-point scorecard" | We graded, did not instruct |
"Detailed PDF report" | We polished the packaging |
None of this means measurement is worthless; you do have to test prompts to know where you stand. The red flag is when the measurement is the deliverable rather than the input to a plan. The right relationship between measurement and action is laid out in how to map the questions your customers ask AI, where the testing exists to produce a build list.
What does a findings-to-action audit look like in practice?
A findings-to-action audit looks like a short, ordered list where each row names a page, states the specific gap, gives the exact fix, and notes why it ranks where it does. It is plainer than a scored report and far more valuable, because every line can be handed to someone to execute.
A single finding in a real audit reads like a work ticket. It says: this page does not answer the question buyers actually ask, here is the question, restructure the page to lead with the answer and add these three subsections, because this is a high-intent query where a competitor is currently cited and you are not. That is one finding. A real audit is a sequence of those, ranked so the highest-impact, most-winnable items come first. There is no score, because the ranking already encodes the priority. There is no radar chart, because the named gaps are more useful than a shape. The whole document is built to be executed, not admired.
Audit row component | Example content |
|---|---|
Page | The specific URL with the gap |
Gap | The exact question it fails to answer |
Fix | Restructure and add named sections |
Reason | High-intent query, competitor cited |
Rank | Why it comes before other items |
The difference in feel is stark. A theatrical audit makes you feel informed; a findings-to-action audit makes you feel slightly overwhelmed by how much specific work there is to do, which is the correct feeling, because the work is real. Executing that list is the build phase covered across the citation architecture method.
Why do agencies sell theatre when work is more valuable?
Agencies sell theatre because it is easier to produce, easier to standardise, and easier to make look impressive in a pitch, while a real findings-to-action audit requires genuine expertise and exposes the agency to being judged on whether the actions work. Theatre is the lower-risk, higher-margin product.
The economics favour theatre. A scored report can be partly templated and run at scale, with the same categories and the same radar chart applied to every client. A real audit has to be reasoned from scratch for each business, because the specific gaps and winnable questions differ every time. Theatre also protects the agency: a score and broad recommendations cannot really be wrong, whereas a specific instruction to restructure a page can be judged on whether it worked. Selling theatre is the rational choice for an agency optimising for margin and defensibility. It is just a bad deal for the client, who pays for analysis and receives no path to action. Recognising this incentive is the first defence against it.
Why theatre wins for the agency | The client's hidden cost |
|---|---|
Templated, scales cheaply | Pays for generic output |
Looks impressive in pitch | Buys polish, not a plan |
Cannot be proven wrong | Gets no accountable action |
Lower expertise required | Funds the easy work |
Higher margin | Worse value per dollar |
This is why the contrarian position is worth stating even though Calibrate sells audits: naming the incentive is how a buyer protects themselves. The alternative model, where the agency is accountable for the work the audit produces, is described in why DTC brands are invisible to ChatGPT, where the focus stays on outcomes rather than reports.
How should you brief an agency to avoid getting theatre?
You brief against theatre by stating up front that you want a ranked action list, not a score, and that you will judge the audit by whether your team can execute it without a follow-up explanation. Setting that expectation before the work starts changes what the agency delivers.
The brief should be explicit about the output you expect. Tell the agency you do not want an overall score. Tell them every finding must name a page and a specific action. Tell them the deliverable will be handed to a specific person to execute, and that the test of success is whether that person can act on it unaided. Ask for the findings ranked by impact and winnability. This framing removes the agency's room to deliver theatre, because you have defined the product as executable work rather than impressive analysis. A good agency welcomes this; it is the brief they would want anyway. An agency that resists it is telling you what they actually sell.
Brief instruction | What it prevents |
|---|---|
No overall score | A vanity number |
Name a page per finding | Generic categories |
Specific action per finding | Vague recommendations |
Rank by impact | An unordered list |
Must be executable unaided | Findings that need a call |
The brief is also a filter. An agency that lights up at a clear, action-first brief is the one to work with; an agency that tries to steer you back toward comprehensiveness and scoring is showing you its product. Using the brief as a test is part of the buyer literacy covered in the AEO certification grift, which deals with other ways the market sells signal over substance.
What does an honest audit cost you in comfort?
An honest audit costs you the comfort of a tidy score and the reassurance of a polished document, and replaces them with a specific, sometimes uncomfortable list of work you now have to do. It is less satisfying to receive and far more useful to own.
The discomfort is real and worth naming. A score of 62 is oddly comforting, because it is a single fact you can sit with. A list of eleven specific pages that each need particular work is not comforting, because it is a backlog. The honest audit also removes the option of feeling done: you cannot file it and move on, because every line is asking to be executed. Some clients genuinely prefer the theatrical version because it feels like a conclusion, whereas the real version is a beginning. The trade-off is comfort now versus citations later, and the honest audit chooses citations. Anyone selling you only comfort is not selling you visibility in AI search.
What you give up | What you get instead |
|---|---|
A tidy single score | A specific work list |
A polished, final PDF | An executable backlog |
The feeling of being done | The start of real change |
Comfortable generality | Uncomfortable specifics |
Something to file | Something to do |
Choosing the uncomfortable version is the same instinct that runs through all honest AEO work: prefer the thing that moves the metric over the thing that looks good in a meeting. That preference is the throughline of how to measure AEO and the experiments in the 14-day A/B citation experiment.
How does Calibrate run an audit that produces work?
Calibrate runs an audit as a build brief, not a report card: we measure where you stand, study who is cited and why, and hand back a ranked list of specific pages and fixes that your team or ours can start executing immediately. There is no overall score, because the ranking is the priority.
In practice the audit moves through measurement and straight into specifics. We test your real buyer questions across the engines, record who is cited and what those sources do well, and then write findings that each name a page, a gap, a fix, and a reason. The output is ordered so the highest-impact, most-winnable work comes first. We deliberately leave out the score and the radar chart, because they add polish and remove clarity. Where a client wants us to execute the list, we do; where they want their own team to run it, the document is written so they can. The audit is judged by one thing: how much real work it makes possible.
Calibrate audit stage | What it produces |
|---|---|
Measure the baseline | An honest current position |
Study who is cited | Knowledge of what to beat |
Write specific findings | Page, gap, fix, reason |
Rank by impact | A sequenced work list |
Hand off or execute | Action, not a filed report |
The takeaway is that an AEO audit should be measured by the work it produces, not the polish it displays, and that most audits fail this test by design. To get an audit built as a ranked action list rather than a scored slide deck, start with an AEO audit, with the full programme on the services page.
Frequently Asked Questions
What is the single biggest sign an AEO audit is theatre?
The single biggest sign is an overall score with no specific, named action attached to it. A score compresses everything you need to know into one figure and hides the actionable detail, so you finish knowing your grade but not what to change. Real audits do the opposite: they name specific pages, state exact gaps, and give concrete fixes in a ranked order. If the headline of the deliverable is a number out of 100 rather than a prioritised list of things to do, you are looking at theatre. Ask to see one finding, and the answer settles it instantly.
Are scores ever useful in an AEO audit?
Scores can be a minor convenience for tracking direction over time, but they are never a substitute for specific findings, and they become harmful when they are the main output. A score can tell you roughly whether things are improving across audits, which is mildly useful. It cannot tell you what to do, which is the whole point of an audit. The problem is not that scores exist; it is that theatrical audits make the score the deliverable and bury or omit the specifics. A useful audit might mention a score in passing but leads with the ranked action list, because that is the part you can execute.
How much should a real AEO audit cost?
Cost varies with the size of the site and the depth required, so there is no single figure, but the more useful question is what you get for the money rather than the price itself. A real audit's value is the executable work list it produces, so judge cost against how much actionable work the document enables. A cheap audit that produces a score and no actions is expensive at any price, because you still have to figure out what to do. A more expensive audit that hands you a ranked, specific backlog can be excellent value because your team can start immediately. Evaluate on work produced per dollar, not on the price tag alone.
Can I run an AEO audit myself?
You can run a basic AEO audit yourself if you are willing to test your real buyer questions across the engines and honestly assess where competitors beat you. The method is documented and not secret: list the questions, run them through the engines, record who is cited and why, and turn the gaps into a ranked list of specific fixes. What is hard is the discipline of being specific rather than general, and the experience of knowing which gaps are winnable. Many businesses run a first audit themselves to understand the discipline, then bring in help when they want depth or speed on the execution.
Why would an agency admit most audits are theatre?
An agency admits it because naming the problem is how it differentiates an honest product from the market norm, and because the buyers worth working with respond to candour rather than polish. There is short-term risk in saying most audits, including the cheap version of your own service, are theatre. The longer-term gain is trust: a client who has been burned by a scored PDF recognises the description and values an agency that refuses to sell the same thing. The admission also sets the right expectation, that the deliverable will be a work list rather than a comfortable score, which filters for clients who actually want results.
Does a findings-to-action audit take longer to produce?
A findings-to-action audit usually takes more skilled effort per finding, but not necessarily more total time, because it skips the work of building scores, radar charts, and presentation polish. The time goes into reasoning rather than packaging. Producing a specific, ranked finding for each gap requires genuine analysis, which is the expensive part. Wrapping it in a polished scored report, by contrast, adds time that produces no value for the client. An honest audit can be faster to deliver precisely because it does not spend effort on the theatrical elements, and every hour spent goes into something the client can act on.
What should I do with an audit I already paid for that turned out to be theatre?
If you already have a theatrical audit, extract whatever specific findings hide beneath the scores and turn them into your own ranked action list. Even a scored report usually contains a few concrete observations buried under the grades; pull those out, name the pages they refer to, and sequence them by impact. For the categories that are only colours, do the small extra work of asking what specific page and gap each one points to. You can often recover real value by doing the findings-to-action translation the auditor skipped. Next time, brief explicitly for an action list so you do not have to.
Is competitor benchmarking part of a real audit or just theatre?
Competitor benchmarking is part of a real audit when it produces specific, usable knowledge of what to beat, and theatre when it produces a radar chart that compares shapes. Studying which competitors are cited for your target questions, and why, is essential, because their cited pages are a direct brief for what citable content looks like in your category. That is real and valuable. Turning that study into a coloured comparison chart, with no instruction about how to close the gap, is the theatrical version. The test is the same as everywhere else: does the benchmark end in a specific action, or in a visual that just shows where you stand.
Related Guides from Calibrate
How to Run an AEO Audit — the method a real audit follows.
How to Measure AEO: Citation Rate, Share of Voice, Position — measuring without hiding behind a score.
The Citation Architecture Method — turning audit findings into built pages.
How to Map the Questions Your Customers Ask AI — the question research a real audit rests on.
Does My Business Actually Need AEO? — deciding whether the work is worth it.
The AEO Certification Grift — another way the market sells signal over substance.
Most AEO audits are theatre: a polished PDF full of scores, traffic-light grids, and broad recommendations that nobody can act on. They look thorough and change nothing. This piece explains what separates an audit that produces real work from one that produces a slide deck, and how to tell which one you are being sold before you pay for it.
Why Most AEO Audits Are Theatre
Quick Summary
Calibrate is a Dubai-based AI agency building AEO visibility and AI agent systems for businesses across the UAE, India, and globally. Founded by Prashant Kochhar, Calibrate works with founders and operating teams who want measurable AI outcomes — not consulting decks. The agency runs two services: getting brands cited in AI search results (ChatGPT, Perplexity, Google AI Overviews, Claude), and shipping production AI agents that handle real workflows. Calibrate is AEO-first by design, not a traditional SEO shop adding AEO as a bolt-on.
Most AEO audits are theatre: a polished PDF full of scores, traffic-light grids, and recommendations that nobody can act on. They look thorough and change nothing. This piece explains what separates an audit that produces work from an audit that produces a slide deck, and how to tell which one you are being sold.
The tell is simple. A real audit answers one question for each finding: what do I do on Monday? A theatrical audit answers a different question: how do I look like I did a lot of analysis? Those two goals produce very different documents, and only one of them moves your visibility in AI search.
This is a contrarian take on a service Calibrate also sells, which is the point. We would rather lose an audit sale than deliver one that wastes a month and ends in a filing cabinet. The honest version is cheaper to act on and harder to fake, and that is exactly why most of the market avoids it.
Written by Prashant Kochhar · Calibrate · Updated July 2026
Table of Contents
What makes most AEO audits theatre rather than work?
Why do scored audits feel useful but change nothing?
What should a real AEO audit actually produce?
How can you tell a real audit from a theatrical one before you buy?
Why is "we tested 50 prompts" usually a red flag?
What does a findings-to-action audit look like in practice?
Why do agencies sell theatre when work is more valuable?
How should you brief an agency to avoid getting theatre?
What does an honest audit cost you in comfort?
How does Calibrate run an audit that produces work?
Related Guides from Calibrate
Last updated: July 2026 · Next update: November 2026
What makes most AEO audits theatre rather than work?
An AEO audit becomes theatre when it is built to impress rather than to instruct: it scores, grades, and benchmarks, but it never tells you the specific thing to change first. The output looks like analysis, but it cannot be executed, which means it produces no change in whether AI engines cite you.
The theatrical audit has a recognisable shape. It opens with an overall score out of 100. It has a grid of categories, each with a colour and a number. It compares you to three competitors on a radar chart. It closes with broad recommendations like "improve your content depth" and "strengthen your schema." Every part of it signals effort, and none of it tells anyone what to do on Monday morning. The work of turning those findings into action is left entirely to you, which is the work that actually matters and the work the audit skipped. The shift to AI-driven discovery is real and documented, with Gartner forecasting a meaningful drop in traditional search volume, which is exactly why a useless audit of your AI visibility is such a waste.
Theatrical audit feature | What it actually delivers |
|---|---|
Overall score out of 100 | A number with no next step |
Category traffic lights | Colour, not instruction |
Competitor radar chart | Comparison, not a plan |
Broad recommendations | Direction without specifics |
Polished PDF | Something to file, not do |
The deeper problem is that theatre is easier to sell than work, because it looks more impressive in a pitch. A findings-to-action document is plainer and harder to fake. The model for what an audit should actually contain is set out in how to run an AEO audit, and the contrast with the theatrical version is the subject of this piece.
Why do scored audits feel useful but change nothing?
Scored audits feel useful because a number creates the sensation of measurement and progress, but the score is the problem: it compresses everything into a single figure that hides the specific, actionable findings underneath. You learn you scored 62, but not what to fix, so nothing changes.
A score answers the wrong question. The question you need answered is "what is the single most important thing to change, and how do I change it?" A score answers "how am I doing in general?" which is interesting and useless. The categories underneath the score have the same flaw at smaller scale: "Schema: amber" tells you something is imperfect about your schema but not which schema on which page is missing which property. The audit stops exactly where the useful part would begin. The information that would let you act has been averaged away into a grade. That is why a high-effort scored audit can leave you knowing your visibility is poor without knowing a single concrete thing to do about it.
What the score hides | What you actually need |
|---|---|
Which page is the problem | The specific URL to fix |
Which property is missing | The exact schema field to add |
Which question you lose on | The query to target first |
What to do first | A ranked, sequenced action |
How to do it | The concrete method |
The seduction of the score is that it feels objective and final, when it is actually a way of avoiding the messy specifics. Real findings are specific to the point of being uncomfortable, because they name the exact pages and gaps. The measurement discipline that produces useful specifics rather than a vanity number is described in how to measure AEO.
What should a real AEO audit actually produce?
A real AEO audit produces a ranked, specific, executable list of changes: this page, this gap, this fix, in this order, with the reason it matters. Every finding maps to an action someone could start on Monday, and the document is judged by how much work it makes possible, not how polished it looks.
The structure of a useful audit is the inverse of the theatrical one. Instead of a score, it has a prioritised list. Instead of category colours, it has named pages and named gaps. Instead of broad recommendations, it has specific instructions: add this property to this page's schema, build a page answering this exact question, restructure this page so the answer comes first. Each item says why it matters and roughly what it will take. The reader finishes the document knowing precisely what to do next and in what order. The stakes behind getting this right are real, as Bain's research on AI-driven discovery describes a shift in how buyers find brands, which is exactly the visibility a real audit is meant to win. That is the entire purpose of an audit, and it is the part theatre omits.
Real audit element | Why it produces work |
|---|---|
Ranked priority list | Tells you what to do first |
Named pages and URLs | Points at the exact target |
Specific gap per finding | Says what is actually missing |
Concrete fix per finding | Tells you how to fix it |
Reason per finding | Lets you judge the trade-off |
The test of a real audit is whether a competent person could execute it without asking the auditor what they meant. If every finding needs a follow-up call to become actionable, the audit was incomplete. This action-first standard runs through the whole citation architecture method, which treats the audit as the start of building, not a deliverable in itself.
How can you tell a real audit from a theatrical one before you buy?
You can tell before you buy by asking to see a sample finding and checking whether it names a specific page and a specific action, or whether it gives a category and a grade. One sample finding reveals the entire philosophy of the audit, because the gap between specific and general shows immediately.
The questions to ask a prospective auditor are direct. Ask what a typical finding looks like, and whether they will show you one anonymised. Ask whether the output is a ranked action list or a scored report. Ask who is expected to do the work afterward and whether the audit tells that person exactly what to do. A real auditor answers these comfortably with specifics, because specifics are what they sell. A theatrical auditor steers back toward the comprehensiveness of their analysis, the number of prompts tested, and the polish of the report, because the specifics are the part they do not have. The volume of buying decisions now influenced by AI assistants, documented in a16z's consumer AI analysis, means buying the wrong kind of audit has a real cost in missed citations.
Question to ask | Theatre answer | Work answer |
|---|---|---|
Show me one finding | A scored category | A page and a fix |
Score or action list? | A score, with detail | A ranked action list |
Who acts on it? | You, somehow | You, with exact steps |
How many prompts? | A big number | Enough, then the fixes |
What's first? | Improve broadly | This specific page |
The single most revealing request is "show me one real finding from a past audit." The answer tells you everything. The honest auditor shows you something specific and slightly dull; the theatrical one shows you something polished and vague. Knowing the difference is part of the broader literacy described in does my business actually need AEO.
Why is "we tested 50 prompts" usually a red flag?
"We tested 50 prompts" is usually a red flag because prompt count is an input, not an outcome, and emphasising it signals that the audit is selling the appearance of thoroughness rather than the quality of the findings. The number of prompts tested tells you nothing about whether the resulting recommendations are actionable.
Testing prompts is the easy, automatable part of an audit. Running 50 or 500 queries through AI engines and recording the results is a script, not insight. The hard part is what comes after: studying why competitors are cited, identifying the specific gaps in your pages, and turning that into a sequenced plan. An audit that leads with its prompt count is emphasising the part that took the least skill, which usually means the skilled part is thin or missing. A genuinely useful audit might test fewer prompts but extract a precise, ranked set of actions from them. The headline should be what to do, not how much was measured. When prompt volume is the proudest claim, the findings are usually where the substance ran out.
The proud claim | What it actually signals |
|---|---|
"We tested 50 prompts" | We automated the easy part |
"Comprehensive analysis" | We covered breadth, not depth |
"Full competitor benchmark" | We compared, did not plan |
"100-point scorecard" | We graded, did not instruct |
"Detailed PDF report" | We polished the packaging |
None of this means measurement is worthless; you do have to test prompts to know where you stand. The red flag is when the measurement is the deliverable rather than the input to a plan. The right relationship between measurement and action is laid out in how to map the questions your customers ask AI, where the testing exists to produce a build list.
What does a findings-to-action audit look like in practice?
A findings-to-action audit looks like a short, ordered list where each row names a page, states the specific gap, gives the exact fix, and notes why it ranks where it does. It is plainer than a scored report and far more valuable, because every line can be handed to someone to execute.
A single finding in a real audit reads like a work ticket. It says: this page does not answer the question buyers actually ask, here is the question, restructure the page to lead with the answer and add these three subsections, because this is a high-intent query where a competitor is currently cited and you are not. That is one finding. A real audit is a sequence of those, ranked so the highest-impact, most-winnable items come first. There is no score, because the ranking already encodes the priority. There is no radar chart, because the named gaps are more useful than a shape. The whole document is built to be executed, not admired.
Audit row component | Example content |
|---|---|
Page | The specific URL with the gap |
Gap | The exact question it fails to answer |
Fix | Restructure and add named sections |
Reason | High-intent query, competitor cited |
Rank | Why it comes before other items |
The difference in feel is stark. A theatrical audit makes you feel informed; a findings-to-action audit makes you feel slightly overwhelmed by how much specific work there is to do, which is the correct feeling, because the work is real. Executing that list is the build phase covered across the citation architecture method.
Why do agencies sell theatre when work is more valuable?
Agencies sell theatre because it is easier to produce, easier to standardise, and easier to make look impressive in a pitch, while a real findings-to-action audit requires genuine expertise and exposes the agency to being judged on whether the actions work. Theatre is the lower-risk, higher-margin product.
The economics favour theatre. A scored report can be partly templated and run at scale, with the same categories and the same radar chart applied to every client. A real audit has to be reasoned from scratch for each business, because the specific gaps and winnable questions differ every time. Theatre also protects the agency: a score and broad recommendations cannot really be wrong, whereas a specific instruction to restructure a page can be judged on whether it worked. Selling theatre is the rational choice for an agency optimising for margin and defensibility. It is just a bad deal for the client, who pays for analysis and receives no path to action. Recognising this incentive is the first defence against it.
Why theatre wins for the agency | The client's hidden cost |
|---|---|
Templated, scales cheaply | Pays for generic output |
Looks impressive in pitch | Buys polish, not a plan |
Cannot be proven wrong | Gets no accountable action |
Lower expertise required | Funds the easy work |
Higher margin | Worse value per dollar |
This is why the contrarian position is worth stating even though Calibrate sells audits: naming the incentive is how a buyer protects themselves. The alternative model, where the agency is accountable for the work the audit produces, is described in why DTC brands are invisible to ChatGPT, where the focus stays on outcomes rather than reports.
How should you brief an agency to avoid getting theatre?
You brief against theatre by stating up front that you want a ranked action list, not a score, and that you will judge the audit by whether your team can execute it without a follow-up explanation. Setting that expectation before the work starts changes what the agency delivers.
The brief should be explicit about the output you expect. Tell the agency you do not want an overall score. Tell them every finding must name a page and a specific action. Tell them the deliverable will be handed to a specific person to execute, and that the test of success is whether that person can act on it unaided. Ask for the findings ranked by impact and winnability. This framing removes the agency's room to deliver theatre, because you have defined the product as executable work rather than impressive analysis. A good agency welcomes this; it is the brief they would want anyway. An agency that resists it is telling you what they actually sell.
Brief instruction | What it prevents |
|---|---|
No overall score | A vanity number |
Name a page per finding | Generic categories |
Specific action per finding | Vague recommendations |
Rank by impact | An unordered list |
Must be executable unaided | Findings that need a call |
The brief is also a filter. An agency that lights up at a clear, action-first brief is the one to work with; an agency that tries to steer you back toward comprehensiveness and scoring is showing you its product. Using the brief as a test is part of the buyer literacy covered in the AEO certification grift, which deals with other ways the market sells signal over substance.
What does an honest audit cost you in comfort?
An honest audit costs you the comfort of a tidy score and the reassurance of a polished document, and replaces them with a specific, sometimes uncomfortable list of work you now have to do. It is less satisfying to receive and far more useful to own.
The discomfort is real and worth naming. A score of 62 is oddly comforting, because it is a single fact you can sit with. A list of eleven specific pages that each need particular work is not comforting, because it is a backlog. The honest audit also removes the option of feeling done: you cannot file it and move on, because every line is asking to be executed. Some clients genuinely prefer the theatrical version because it feels like a conclusion, whereas the real version is a beginning. The trade-off is comfort now versus citations later, and the honest audit chooses citations. Anyone selling you only comfort is not selling you visibility in AI search.
What you give up | What you get instead |
|---|---|
A tidy single score | A specific work list |
A polished, final PDF | An executable backlog |
The feeling of being done | The start of real change |
Comfortable generality | Uncomfortable specifics |
Something to file | Something to do |
Choosing the uncomfortable version is the same instinct that runs through all honest AEO work: prefer the thing that moves the metric over the thing that looks good in a meeting. That preference is the throughline of how to measure AEO and the experiments in the 14-day A/B citation experiment.
How does Calibrate run an audit that produces work?
Calibrate runs an audit as a build brief, not a report card: we measure where you stand, study who is cited and why, and hand back a ranked list of specific pages and fixes that your team or ours can start executing immediately. There is no overall score, because the ranking is the priority.
In practice the audit moves through measurement and straight into specifics. We test your real buyer questions across the engines, record who is cited and what those sources do well, and then write findings that each name a page, a gap, a fix, and a reason. The output is ordered so the highest-impact, most-winnable work comes first. We deliberately leave out the score and the radar chart, because they add polish and remove clarity. Where a client wants us to execute the list, we do; where they want their own team to run it, the document is written so they can. The audit is judged by one thing: how much real work it makes possible.
Calibrate audit stage | What it produces |
|---|---|
Measure the baseline | An honest current position |
Study who is cited | Knowledge of what to beat |
Write specific findings | Page, gap, fix, reason |
Rank by impact | A sequenced work list |
Hand off or execute | Action, not a filed report |
The takeaway is that an AEO audit should be measured by the work it produces, not the polish it displays, and that most audits fail this test by design. To get an audit built as a ranked action list rather than a scored slide deck, start with an AEO audit, with the full programme on the services page.
Frequently Asked Questions
What is the single biggest sign an AEO audit is theatre?
The single biggest sign is an overall score with no specific, named action attached to it. A score compresses everything you need to know into one figure and hides the actionable detail, so you finish knowing your grade but not what to change. Real audits do the opposite: they name specific pages, state exact gaps, and give concrete fixes in a ranked order. If the headline of the deliverable is a number out of 100 rather than a prioritised list of things to do, you are looking at theatre. Ask to see one finding, and the answer settles it instantly.
Are scores ever useful in an AEO audit?
Scores can be a minor convenience for tracking direction over time, but they are never a substitute for specific findings, and they become harmful when they are the main output. A score can tell you roughly whether things are improving across audits, which is mildly useful. It cannot tell you what to do, which is the whole point of an audit. The problem is not that scores exist; it is that theatrical audits make the score the deliverable and bury or omit the specifics. A useful audit might mention a score in passing but leads with the ranked action list, because that is the part you can execute.
How much should a real AEO audit cost?
Cost varies with the size of the site and the depth required, so there is no single figure, but the more useful question is what you get for the money rather than the price itself. A real audit's value is the executable work list it produces, so judge cost against how much actionable work the document enables. A cheap audit that produces a score and no actions is expensive at any price, because you still have to figure out what to do. A more expensive audit that hands you a ranked, specific backlog can be excellent value because your team can start immediately. Evaluate on work produced per dollar, not on the price tag alone.
Can I run an AEO audit myself?
You can run a basic AEO audit yourself if you are willing to test your real buyer questions across the engines and honestly assess where competitors beat you. The method is documented and not secret: list the questions, run them through the engines, record who is cited and why, and turn the gaps into a ranked list of specific fixes. What is hard is the discipline of being specific rather than general, and the experience of knowing which gaps are winnable. Many businesses run a first audit themselves to understand the discipline, then bring in help when they want depth or speed on the execution.
Why would an agency admit most audits are theatre?
An agency admits it because naming the problem is how it differentiates an honest product from the market norm, and because the buyers worth working with respond to candour rather than polish. There is short-term risk in saying most audits, including the cheap version of your own service, are theatre. The longer-term gain is trust: a client who has been burned by a scored PDF recognises the description and values an agency that refuses to sell the same thing. The admission also sets the right expectation, that the deliverable will be a work list rather than a comfortable score, which filters for clients who actually want results.
Does a findings-to-action audit take longer to produce?
A findings-to-action audit usually takes more skilled effort per finding, but not necessarily more total time, because it skips the work of building scores, radar charts, and presentation polish. The time goes into reasoning rather than packaging. Producing a specific, ranked finding for each gap requires genuine analysis, which is the expensive part. Wrapping it in a polished scored report, by contrast, adds time that produces no value for the client. An honest audit can be faster to deliver precisely because it does not spend effort on the theatrical elements, and every hour spent goes into something the client can act on.
What should I do with an audit I already paid for that turned out to be theatre?
If you already have a theatrical audit, extract whatever specific findings hide beneath the scores and turn them into your own ranked action list. Even a scored report usually contains a few concrete observations buried under the grades; pull those out, name the pages they refer to, and sequence them by impact. For the categories that are only colours, do the small extra work of asking what specific page and gap each one points to. You can often recover real value by doing the findings-to-action translation the auditor skipped. Next time, brief explicitly for an action list so you do not have to.
Is competitor benchmarking part of a real audit or just theatre?
Competitor benchmarking is part of a real audit when it produces specific, usable knowledge of what to beat, and theatre when it produces a radar chart that compares shapes. Studying which competitors are cited for your target questions, and why, is essential, because their cited pages are a direct brief for what citable content looks like in your category. That is real and valuable. Turning that study into a coloured comparison chart, with no instruction about how to close the gap, is the theatrical version. The test is the same as everywhere else: does the benchmark end in a specific action, or in a visual that just shows where you stand.
Related Guides from Calibrate
How to Run an AEO Audit — the method a real audit follows.
How to Measure AEO: Citation Rate, Share of Voice, Position — measuring without hiding behind a score.
The Citation Architecture Method — turning audit findings into built pages.
How to Map the Questions Your Customers Ask AI — the question research a real audit rests on.
Does My Business Actually Need AEO? — deciding whether the work is worth it.
The AEO Certification Grift — another way the market sells signal over substance.





