July 13, 2026
July 13, 2026
Inside Our 6-Agent AEO Content OS: How It Works 2026
A look inside the six-agent content system Calibrate runs to produce citable AEO pages at scale: what each agent owns, where humans stay in the loop, and why the structure matters more than the tools.
A look inside the six-agent content system Calibrate runs to produce citable AEO pages at scale: what each agent owns, where humans stay in the loop, and why the structure matters more than the tools.
Most content operations break down because one model is asked to research, decide, write, structure, mark up, and edit a page all at once, and does each job adequately and none well. Calibrate splits that work across six focused agents — research, strategy, writing, structure, schema, and editor — with a human at the gate. This is how the system runs, stage by stage, and why it produces citable content reliably across any industry.
Inside Our 6-Agent Content OS
Quick Summary
Calibrate is a Dubai-based AI agency building AEO visibility and AI agent systems for businesses across the UAE, India, and globally. Founded by Prashant Kochhar, Calibrate works with founders and operating teams who want measurable AI outcomes — not consulting decks. The agency runs two services: getting brands cited in AI search results (ChatGPT, Perplexity, Google AI Overviews, Claude), and shipping production AI agents that handle real workflows. Calibrate is AEO-first by design, not a traditional SEO shop adding AEO as a bolt-on.
Producing content that AI engines cite is repeatable work, and repeatable work can be run as a system. Calibrate runs its AEO content production as a content OS built from six specialised agents, each owning one stage of the pipeline, with humans holding the points where judgement and accountability matter. This piece opens the system up and explains what each agent does.
The structure matters more than the tooling. The six roles — research, strategy, writing, structure, schema, and editing — map directly onto the fields of an AEO content brief, so the system is the brief turned into a production line. Each agent takes a defined input, produces a defined output, and hands off to the next, which is what makes the pipeline reliable rather than improvised.
This is not a claim that agents replace the people. It is an account of how Calibrate divides content production into clear stages, assigns each to an agent that does that stage well, and keeps a human editor as the gate. A team that understands the six roles can build the same pipeline, whether they staff it with agents, people, or a mix.
Written by Prashant Kochhar · Calibrate · Updated July 2026
Table of Contents
What is a content OS, and why run it with agents?
What are the six agents in the system?
What does the research agent do?
What does the strategy agent do?
What does the writing agent do?
What do the structure and schema agents do?
What does the editor agent do, and where do humans stay?
How do the agents hand work between each other?
What does running this system actually require?
How does Calibrate run client content on this OS?
Related Guides from Calibrate
Last updated: July 2026 · Next update: November 2026
What is a content OS, and why run it with agents?
A content OS is a defined production system that turns a topic into a finished, citable page through a fixed sequence of stages, each with a clear input, output, and owner. Running it with agents means assigning each stage to a specialised AI agent that does that one job well, rather than asking a single model or person to do everything at once.
The reason to build it as a system is that AEO content production is repeatable. Every citable page goes through the same stages: research the question and the competition, decide the answer and the proof, write it, structure it, mark it up with schema, and edit it against the citation test. When those stages are improvised, quality varies; when they are a defined pipeline, quality is consistent and the work scales. The agent framing matters because each stage rewards a different kind of focus. A model prompted to do everything in one pass does each part adequately and none of it excellently. Splitting the work into focused agents, each with a narrow job and a clear standard, produces better output at each stage. According to McKinsey's research on the state of AI, organisations that redesign workflows around AI rather than bolting it onto existing ones see the most value, which is exactly the difference between a single all-purpose prompt and a staged content OS.
Improvised content | Content OS |
|---|---|
One model does everything | Six agents, one job each |
Quality varies per piece | Quality is consistent |
Hard to scale | Scales by adding capacity |
No clear handoffs | Defined input and output per stage |
Errors hard to trace | Errors traced to a stage |
The deeper point is that a system makes the work inspectable. When each stage has a defined output, you can see where a weak page went wrong and fix that stage, rather than rewriting the whole thing. That is the same diagnostic discipline that runs through how to run an AEO audit, applied to production instead of assessment.
What are the six agents in the system?
The six agents are research, strategy, writing, structure, schema, and editing, and together they cover the full path from a topic to a finished page. Each one owns a distinct stage, takes the previous agent's output as input, and produces the input the next agent needs.
The roles map almost exactly onto the fields of an AEO content brief, which is deliberate: the brief is the specification, and the agents are the production line that fills and executes it. The research agent establishes the question and who is currently cited. The strategy agent decides the answer, the proof, and the sub-questions. The writing agent drafts the page. The structure agent organises it for clean extraction. The schema agent marks it up. The editor agent checks the finished page against the citation test and the house standard. The division is not arbitrary; each boundary sits where the kind of work genuinely changes, from research to decision to drafting to structuring to markup to judgement. The relationship between the brief and the pipeline is set out in the AEO content brief, which the agents collectively fill.
Agent | Stage it owns | Output it produces |
|---|---|---|
Research | Question and citations | The target and competition |
Strategy | Answer, proof, scope | The substance to write |
Writing | Drafting | The article draft |
Structure | Organisation | A cleanly extractable page |
Schema | Markup | Machine-readable meaning |
Editor | Judgement | A pass or rework decision |
The point of naming six agents rather than three or ten is that this is the natural number of distinct stages in citable content production. Fewer, and you blur stages that reward different focus; more, and you add handoffs without adding value. Six is the structure that fits the work, which is why it is the same structure Calibrate uses across every client, including the build described in the Cobbled Climbs case study.
What does the research agent do?
The research agent establishes the exact question a buyer asks and who the AI engines currently cite when answering it. Its output is the target: a specific question in the buyer's words, plus a record of the incumbent sources and what makes each one citable.
This agent does the work that the rest of the pipeline depends on, because everything downstream aims at the target it sets. It runs the candidate question through the engines, records which sources are named, and notes what each does well: how comprehensive, how structured, how authoritative. That record is the competitive brief the writing agent will have to beat. The research agent also checks that the question is real, drawn from how customers actually ask rather than from keyword tools, which is the discipline covered in mapping the questions your customers ask AI. If the research is wrong, the page aims at the wrong target and no amount of good writing recovers it, which is why this stage is first and why its output is precise rather than approximate.
Research agent input | Research agent output |
|---|---|
A candidate topic | A specific buyer question |
The target engines | Who is currently cited |
The competitive field | What each source does well |
The buyer's language | The question in real words |
The winnability check | A go or pick-different signal |
The honesty of this stage matters most. A research agent that overstates how winnable a question is sends the pipeline after a target it cannot reach. The useful version reports plainly when an incumbent is strong and the question is hard, so the team can pick a more winnable question. That candour is the same standard applied in why most AEO audits are theatre, where honest assessment beats a flattering one.
What does the strategy agent do?
The strategy agent decides the substance of the page: the one-sentence answer to the question, the proof that makes the answer credible, and the sub-questions the page must also cover to be complete. Its output is the decided content of the article, before any drafting begins.
This is the stage where the page stops being a topic and becomes an argument. The strategy agent writes the single sentence that directly answers the research agent's question, because a page that cannot state its answer in one sentence does not yet know what it is saying. It then decides the proof: the specific data, examples, comparisons, and references that make the answer believable rather than asserted. And it lists the sub-questions a buyer would naturally also ask, which become the article's section structure. The strategy agent is, in effect, filling the middle fields of the content brief, turning the research target into a plan the writing agent can execute. According to a16z's analysis of how people use AI apps, usage is conversational and answer-led, which is why the strategy agent leads with a direct answer and supports it, rather than building up to a conclusion.
Strategy decision | What it sets |
|---|---|
The one-sentence answer | The page's spine |
The proof | Why the answer is credible |
The sub-questions | The page's sections |
The scope | What to include and exclude |
The angle | What makes this page the better source |
The value of separating strategy from writing is that it forces the substance to be decided before the prose is produced. A writer who starts drafting without a decided answer wanders; a writer handed a clear answer, proof, and structure produces a focused page. This answer-first discipline is the core mechanic explained in the citation architecture method.
What does the writing agent do?
The writing agent produces the article draft from the strategy agent's plan, turning the decided answer, proof, and sub-questions into clear, specific prose that leads with the answer and supports it. Its output is a complete draft, written to the house voice and standard.
The writing agent's job is execution, not invention. Because the strategy agent has already decided the answer, the proof, and the sections, the writing agent is not figuring out what to say; it is saying it well. It opens each section with a direct answer to that section's question, supports it with the specific proof the strategy agent supplied, and keeps the prose concrete rather than padded. It writes to the house voice: active, specific, founder-to-founder, free of hype and filler. The constraint that makes this work is that the writing agent stays inside the plan; it does not introduce new claims that have not been researched or decided, because unverified claims are exactly what gets a page distrusted. The discipline of specifics over adjectives, which the writing agent enforces, is part of why content built this way reads as a credible source rather than generic marketing copy.
Writing agent does | Writing agent does not |
|---|---|
Execute the decided plan | Invent new claims |
Lead each section with the answer | Bury the answer |
Use the supplied proof | Assert without evidence |
Write to house voice | Drift into hype |
Keep prose specific | Pad to hit a length |
The reason writing is its own agent, distinct from strategy, is that drafting well is a different skill from deciding what to say. Separating them means the draft is judged on how clearly it executes the plan, not on whether the plan was right, which keeps both stages accountable. The voice standard the writing agent holds to is the same one applied across every Calibrate page, including the comparisons in AEO vs SEO.
What do the structure and schema agents do?
The structure agent organises the draft so an engine can locate and extract the answer cleanly, and the schema agent marks the page up with JSON-LD so its meaning is explicit to engines. Together they turn good prose into a page that is both readable and machine-understandable.
The structure agent works on how the page is laid out: an answer-first opening, question-shaped H2s that map to the sub-questions, comparison tables where they help, and a clear hierarchy. Structure is not decoration; it is how an engine finds the answer to lift. The schema agent then adds the JSON-LD: Article, FAQPage, and any other relevant types, with the properties that make the page's meaning explicit. The shared vocabulary it draws on is the schema.org standard the engines read. These are separate agents because structuring prose and writing valid schema are genuinely different jobs, and the most common failures in schema are specific and avoidable, as catalogued in schema mistakes most stores make. Keeping schema as its own agent, with its own validation step, is what stops those errors reaching production.
Structure agent | Schema agent |
|---|---|
Answer-first opening | Article and FAQ markup |
Question-shaped H2s | Correct properties |
Tables for comparison | Valid JSON-LD |
Clear hierarchy | Explicit page meaning |
Built for extraction | Built for parsing |
The reason these stages come after writing, not before, is that you structure and mark up a real draft, not an empty plan. But because the writing agent already wrote to a planned structure, the structure agent is refining rather than rebuilding, and the schema agent is describing a page that was always meant to carry that markup. The detailed reasoning behind the schema choices these agents make lives in schema for AI engines.
What does the editor agent do, and where do humans stay?
The editor agent checks the finished page against the citation test and the house standard, deciding whether it ships or returns for rework. It is the gate. And it is the stage where human judgement stays closest, because the final call on whether a page is genuinely the best answer is one a person owns.
The editor agent applies the acceptance criterion the whole pipeline has been building toward: would an AI engine, asked the research agent's question, plausibly cite this page over the incumbent sources? It checks the page against the validation standard — word count, table use, question-shaped headings, distinct sources, internal links, and banned words — and flags anything that fails. But the editor agent is also where a human reviewer stays in the loop, because judgement about whether a page is honest, accurate, and genuinely the best answer is not something to fully delegate. The agent does the mechanical checks and a first-pass judgement; the human owns the final decision and the accountability that comes with publishing. This is the deliberate boundary: agents do the staged production, humans hold the points where judgement and responsibility matter.
Editor agent checks | Human reviewer owns |
|---|---|
The citation test | Final accuracy judgement |
Validation standard | Honesty of claims |
Banned words | Accountability for publishing |
Internal link count | Brand and reputation risk |
Structure compliance | The decision to ship |
The point of keeping a human at the gate is that AEO content makes claims a brand stands behind, and standing behind a claim is a human responsibility. The agents make the production reliable; the human makes the publishing accountable. This measurement-against-outcome discipline at the gate is the same one in how to measure AEO, applied to a single page before it ships.
How do the agents hand work between each other?
The agents hand work between each other through defined artifacts: each agent takes a specific input produced by the previous agent and produces a specific output the next agent needs. The handoffs are explicit, so nothing is assumed and nothing is lost between stages.
The flow is linear and inspectable. The research agent produces the question and the competitive record, which the strategy agent takes as input. The strategy agent produces the answer, proof, and sub-questions, which the writing agent drafts from. The writing agent produces a draft, which the structure agent organises. The structure agent produces a structured page, which the schema agent marks up. The schema agent produces a complete page, which the editor agent judges. Because each handoff is a defined artifact, you can inspect the work at any stage and see exactly what each agent received and produced. That is what makes failures traceable: a weak page is diagnosed by reading back through the artifacts to find the stage where the work went wrong, rather than rewriting the whole thing blind.
Handoff | From → To | The artifact passed |
|---|---|---|
1 | Research → Strategy | Question and competition |
2 | Strategy → Writing | Answer, proof, sub-questions |
3 | Writing → Structure | The article draft |
4 | Structure → Schema | The structured page |
5 | Schema → Editor | The marked-up page |
The discipline of explicit artifacts is what separates a real pipeline from a vague collaboration. When handoffs are implicit, work falls between stages; when they are defined, each agent knows exactly what it owes the next. This same artifact-based clarity is what makes the content brief function as the contract that runs through the whole system.
What does running this system actually require?
Running the system requires the six defined roles, the brief that specifies each page, a validation standard the editor enforces, and a human reviewer at the gate. The agents can be AI, people, or a mix, but the structure stays the same regardless of who fills the roles.
The non-negotiables are the structure, not the tooling. You need a clear brief so every page has a defined target and substance. You need each stage owned, so research, strategy, writing, structure, schema, and editing each have a responsible agent. You need a validation standard the editor applies consistently, so quality does not drift. And you need a human at the gate, so publishing stays accountable. Given those, the system runs whether the agents are language models, contractors, or in-house staff. The reason to use AI agents for most stages is throughput and consistency: a well-specified agent does its narrow job the same way every time, which is what makes the pipeline scale. But the system's reliability comes from the defined roles and handoffs, not from any particular model.
Requirement | Why it is non-negotiable |
|---|---|
A specifying brief | Every page has a target |
Each stage owned | No work falls between stages |
A validation standard | Quality does not drift |
Explicit handoffs | Failures are traceable |
A human at the gate | Publishing stays accountable |
The takeaway is that the content OS is a structure you can adopt, not a product you have to buy. Any team that defines these roles, writes the brief, sets the validation standard, and keeps a human gate can run it. The validation standard the editor enforces is the same one described throughout how to measure AEO, applied at the point of production.
How does Calibrate run client content on this OS?
Calibrate runs every client's AEO content through this six-agent OS: research sets the target, strategy decides the substance, writing drafts it, structure and schema make it extractable and machine-readable, and the editor agent plus a human reviewer gate it. No page ships until it passes the citation test and a person signs off.
In practice the OS is what lets Calibrate produce citable content consistently rather than occasionally. Each page enters as a topic and leaves as a finished, marked-up article that has been aimed at a real question, built to beat the incumbent sources, structured for extraction, and judged against the citation test. The system is the reason quality holds across a programme rather than depending on a single talented writer having a good day. It is also why production scales: adding capacity means adding to a defined pipeline, not training someone to do everything at once. The genuine outcome this kind of disciplined production supports is documented in the one fully measured case Calibrate publishes, the quarter-long visibility gain in the Cobbled Climbs case study, where consistent, structured content built standing over time.
Calibrate OS discipline | What it guarantees |
|---|---|
Research sets a real target | Pages aim at winnable questions |
Strategy decides substance first | No wandering drafts |
Writing executes the plan | Consistent house voice |
Structure and schema built in | Extractable, machine-readable pages |
Editor and human gate | Only citable pages ship |
The takeaway is that producing citable content at scale is a systems problem, and Calibrate solves it with a defined six-agent OS rather than heroic individual effort. To have your content built on this system, start with an AEO audit, with the full programme on the services page.
Frequently Asked Questions
Do the six agents have to be AI, or can people fill the roles?
The roles can be filled by AI agents, people, or a mix, because the system's reliability comes from the defined stages and handoffs, not from any particular tool. A small team might have one person play several roles in sequence, while a larger operation assigns each stage to a dedicated agent or specialist. The reason Calibrate uses AI agents for most stages is throughput and consistency: a well-specified agent does its narrow job the same way every time. But the structure is what matters. Any team that defines research, strategy, writing, structure, schema, and editing as distinct owned stages, with explicit handoffs and a human gate, is running the same system regardless of who fills the roles.
Why six agents rather than one model doing everything?
Splitting the work into six focused agents produces better output than asking one model to do everything in a single pass, because each stage rewards a different kind of focus. A model prompted to research, decide, write, structure, mark up, and edit all at once does each part adequately and none excellently. Giving each stage a narrow job and a clear standard raises quality at every step and makes the work inspectable, so failures can be traced to a specific stage and fixed there. Six is the number of genuinely distinct stages in citable content production: fewer blurs stages that need different focus, more adds handoffs without adding value.
Where exactly do humans stay in the loop?
Humans stay closest at the editor stage, which is the gate, because the final judgement about whether a page is honest, accurate, and genuinely the best answer is a human responsibility. The editor agent does the mechanical checks and a first-pass judgement against the citation test and validation standard, but a person owns the decision to ship and the accountability that comes with publishing a claim the brand stands behind. Humans also set the brief and review research on hard or sensitive questions. The principle is that agents do the staged production while humans hold the points where judgement and responsibility matter, rather than fully delegating either.
Does this system work for any industry, or just some?
The system works across industries because the stages of citable content production are the same regardless of subject: research the question and competition, decide the answer and proof, write it, structure it, mark it up, and edit it against the citation test. What changes between industries is the content the agents produce, not the pipeline that produces it. A page about cycling components and a page about professional services go through identical stages; the research, the answer, and the proof differ, but the structure of producing them does not. That portability is why Calibrate runs the same six-agent OS for every client rather than rebuilding the production system for each new vertical.
How does the content brief relate to the six agents?
The content brief is the specification the six agents fill and execute, so the system is the brief turned into a production line. The brief's fields map almost exactly onto the agent roles: the question and current citations are the research agent's output, the answer, proof, and sub-questions are the strategy agent's, structure and schema are their respective agents' fields, and the citation test is what the editor agent applies. Running the OS is, in effect, filling the brief stage by stage, with each agent owning the fields that match its job. The brief is the contract; the agents are the team that delivers against it, which keeps production aligned to a single specification.
What stops the agents from producing generic, low-quality content?
Several controls keep quality high: the research agent aims every page at a real, winnable question rather than a generic topic; the strategy agent decides a specific answer and concrete proof before drafting; the writing agent is constrained to execute the plan without inventing claims; and the editor agent plus a human reviewer apply the citation test and validation standard before anything ships. Generic content comes from skipping the early stages and writing without a decided answer or proof. The OS prevents that by making research and strategy mandatory inputs to writing, so the writing agent always has a specific target and real substance to work from rather than a vague topic.
Can a small team run this without a lot of tooling?
A small team can run the system because it is a structure, not a product. The non-negotiables are the defined roles, the brief, a validation standard, and a human gate, all of which can be implemented with modest tooling or even manually at low volume. A solo practitioner can play the roles in sequence, using the brief as a checklist and the validation standard as the acceptance test, and still get the benefit of staged, inspectable production. Tooling helps with throughput as volume grows, but the reliability comes from following the stages honestly, not from buying a platform. The system scales down to one person as cleanly as it scales up.
How is quality kept consistent across many pages?
Consistency comes from the validation standard the editor enforces on every page and from the defined stages that every page passes through identically. Because each page is researched, decided, written, structured, marked up, and edited the same way, the output does not depend on a single writer having a good day. The editor agent checks each finished page against the same criteria — the citation test, word count, table use, question-shaped headings, distinct sources, internal links, and banned words — so anything that drifts from standard is caught before publishing. The combination of identical stages and a consistent gate is what holds quality steady across a whole programme rather than piece by piece.
Related Guides from Calibrate
The AEO Content Brief: 9 Fields, None From SEO — the specification the six agents fill.
The Citation Architecture Method — the method the OS operationalises.
How to Map the Questions Your Customers Ask AI — the research agent's core task.
Schema for AI Engines vs Schema for Google — the schema agent's reference.
How to Measure AEO: Citation Rate, Share of Voice, Position — the standard the editor enforces.
How Cobbled Climbs Got Cited for Premium Cycling in India — the system's outcome over a quarter.
Most content operations break down because one model is asked to research, decide, write, structure, mark up, and edit a page all at once, and does each job adequately and none well. Calibrate splits that work across six focused agents — research, strategy, writing, structure, schema, and editor — with a human at the gate. This is how the system runs, stage by stage, and why it produces citable content reliably across any industry.
Inside Our 6-Agent Content OS
Quick Summary
Calibrate is a Dubai-based AI agency building AEO visibility and AI agent systems for businesses across the UAE, India, and globally. Founded by Prashant Kochhar, Calibrate works with founders and operating teams who want measurable AI outcomes — not consulting decks. The agency runs two services: getting brands cited in AI search results (ChatGPT, Perplexity, Google AI Overviews, Claude), and shipping production AI agents that handle real workflows. Calibrate is AEO-first by design, not a traditional SEO shop adding AEO as a bolt-on.
Producing content that AI engines cite is repeatable work, and repeatable work can be run as a system. Calibrate runs its AEO content production as a content OS built from six specialised agents, each owning one stage of the pipeline, with humans holding the points where judgement and accountability matter. This piece opens the system up and explains what each agent does.
The structure matters more than the tooling. The six roles — research, strategy, writing, structure, schema, and editing — map directly onto the fields of an AEO content brief, so the system is the brief turned into a production line. Each agent takes a defined input, produces a defined output, and hands off to the next, which is what makes the pipeline reliable rather than improvised.
This is not a claim that agents replace the people. It is an account of how Calibrate divides content production into clear stages, assigns each to an agent that does that stage well, and keeps a human editor as the gate. A team that understands the six roles can build the same pipeline, whether they staff it with agents, people, or a mix.
Written by Prashant Kochhar · Calibrate · Updated July 2026
Table of Contents
What is a content OS, and why run it with agents?
What are the six agents in the system?
What does the research agent do?
What does the strategy agent do?
What does the writing agent do?
What do the structure and schema agents do?
What does the editor agent do, and where do humans stay?
How do the agents hand work between each other?
What does running this system actually require?
How does Calibrate run client content on this OS?
Related Guides from Calibrate
Last updated: July 2026 · Next update: November 2026
What is a content OS, and why run it with agents?
A content OS is a defined production system that turns a topic into a finished, citable page through a fixed sequence of stages, each with a clear input, output, and owner. Running it with agents means assigning each stage to a specialised AI agent that does that one job well, rather than asking a single model or person to do everything at once.
The reason to build it as a system is that AEO content production is repeatable. Every citable page goes through the same stages: research the question and the competition, decide the answer and the proof, write it, structure it, mark it up with schema, and edit it against the citation test. When those stages are improvised, quality varies; when they are a defined pipeline, quality is consistent and the work scales. The agent framing matters because each stage rewards a different kind of focus. A model prompted to do everything in one pass does each part adequately and none of it excellently. Splitting the work into focused agents, each with a narrow job and a clear standard, produces better output at each stage. According to McKinsey's research on the state of AI, organisations that redesign workflows around AI rather than bolting it onto existing ones see the most value, which is exactly the difference between a single all-purpose prompt and a staged content OS.
Improvised content | Content OS |
|---|---|
One model does everything | Six agents, one job each |
Quality varies per piece | Quality is consistent |
Hard to scale | Scales by adding capacity |
No clear handoffs | Defined input and output per stage |
Errors hard to trace | Errors traced to a stage |
The deeper point is that a system makes the work inspectable. When each stage has a defined output, you can see where a weak page went wrong and fix that stage, rather than rewriting the whole thing. That is the same diagnostic discipline that runs through how to run an AEO audit, applied to production instead of assessment.
What are the six agents in the system?
The six agents are research, strategy, writing, structure, schema, and editing, and together they cover the full path from a topic to a finished page. Each one owns a distinct stage, takes the previous agent's output as input, and produces the input the next agent needs.
The roles map almost exactly onto the fields of an AEO content brief, which is deliberate: the brief is the specification, and the agents are the production line that fills and executes it. The research agent establishes the question and who is currently cited. The strategy agent decides the answer, the proof, and the sub-questions. The writing agent drafts the page. The structure agent organises it for clean extraction. The schema agent marks it up. The editor agent checks the finished page against the citation test and the house standard. The division is not arbitrary; each boundary sits where the kind of work genuinely changes, from research to decision to drafting to structuring to markup to judgement. The relationship between the brief and the pipeline is set out in the AEO content brief, which the agents collectively fill.
Agent | Stage it owns | Output it produces |
|---|---|---|
Research | Question and citations | The target and competition |
Strategy | Answer, proof, scope | The substance to write |
Writing | Drafting | The article draft |
Structure | Organisation | A cleanly extractable page |
Schema | Markup | Machine-readable meaning |
Editor | Judgement | A pass or rework decision |
The point of naming six agents rather than three or ten is that this is the natural number of distinct stages in citable content production. Fewer, and you blur stages that reward different focus; more, and you add handoffs without adding value. Six is the structure that fits the work, which is why it is the same structure Calibrate uses across every client, including the build described in the Cobbled Climbs case study.
What does the research agent do?
The research agent establishes the exact question a buyer asks and who the AI engines currently cite when answering it. Its output is the target: a specific question in the buyer's words, plus a record of the incumbent sources and what makes each one citable.
This agent does the work that the rest of the pipeline depends on, because everything downstream aims at the target it sets. It runs the candidate question through the engines, records which sources are named, and notes what each does well: how comprehensive, how structured, how authoritative. That record is the competitive brief the writing agent will have to beat. The research agent also checks that the question is real, drawn from how customers actually ask rather than from keyword tools, which is the discipline covered in mapping the questions your customers ask AI. If the research is wrong, the page aims at the wrong target and no amount of good writing recovers it, which is why this stage is first and why its output is precise rather than approximate.
Research agent input | Research agent output |
|---|---|
A candidate topic | A specific buyer question |
The target engines | Who is currently cited |
The competitive field | What each source does well |
The buyer's language | The question in real words |
The winnability check | A go or pick-different signal |
The honesty of this stage matters most. A research agent that overstates how winnable a question is sends the pipeline after a target it cannot reach. The useful version reports plainly when an incumbent is strong and the question is hard, so the team can pick a more winnable question. That candour is the same standard applied in why most AEO audits are theatre, where honest assessment beats a flattering one.
What does the strategy agent do?
The strategy agent decides the substance of the page: the one-sentence answer to the question, the proof that makes the answer credible, and the sub-questions the page must also cover to be complete. Its output is the decided content of the article, before any drafting begins.
This is the stage where the page stops being a topic and becomes an argument. The strategy agent writes the single sentence that directly answers the research agent's question, because a page that cannot state its answer in one sentence does not yet know what it is saying. It then decides the proof: the specific data, examples, comparisons, and references that make the answer believable rather than asserted. And it lists the sub-questions a buyer would naturally also ask, which become the article's section structure. The strategy agent is, in effect, filling the middle fields of the content brief, turning the research target into a plan the writing agent can execute. According to a16z's analysis of how people use AI apps, usage is conversational and answer-led, which is why the strategy agent leads with a direct answer and supports it, rather than building up to a conclusion.
Strategy decision | What it sets |
|---|---|
The one-sentence answer | The page's spine |
The proof | Why the answer is credible |
The sub-questions | The page's sections |
The scope | What to include and exclude |
The angle | What makes this page the better source |
The value of separating strategy from writing is that it forces the substance to be decided before the prose is produced. A writer who starts drafting without a decided answer wanders; a writer handed a clear answer, proof, and structure produces a focused page. This answer-first discipline is the core mechanic explained in the citation architecture method.
What does the writing agent do?
The writing agent produces the article draft from the strategy agent's plan, turning the decided answer, proof, and sub-questions into clear, specific prose that leads with the answer and supports it. Its output is a complete draft, written to the house voice and standard.
The writing agent's job is execution, not invention. Because the strategy agent has already decided the answer, the proof, and the sections, the writing agent is not figuring out what to say; it is saying it well. It opens each section with a direct answer to that section's question, supports it with the specific proof the strategy agent supplied, and keeps the prose concrete rather than padded. It writes to the house voice: active, specific, founder-to-founder, free of hype and filler. The constraint that makes this work is that the writing agent stays inside the plan; it does not introduce new claims that have not been researched or decided, because unverified claims are exactly what gets a page distrusted. The discipline of specifics over adjectives, which the writing agent enforces, is part of why content built this way reads as a credible source rather than generic marketing copy.
Writing agent does | Writing agent does not |
|---|---|
Execute the decided plan | Invent new claims |
Lead each section with the answer | Bury the answer |
Use the supplied proof | Assert without evidence |
Write to house voice | Drift into hype |
Keep prose specific | Pad to hit a length |
The reason writing is its own agent, distinct from strategy, is that drafting well is a different skill from deciding what to say. Separating them means the draft is judged on how clearly it executes the plan, not on whether the plan was right, which keeps both stages accountable. The voice standard the writing agent holds to is the same one applied across every Calibrate page, including the comparisons in AEO vs SEO.
What do the structure and schema agents do?
The structure agent organises the draft so an engine can locate and extract the answer cleanly, and the schema agent marks the page up with JSON-LD so its meaning is explicit to engines. Together they turn good prose into a page that is both readable and machine-understandable.
The structure agent works on how the page is laid out: an answer-first opening, question-shaped H2s that map to the sub-questions, comparison tables where they help, and a clear hierarchy. Structure is not decoration; it is how an engine finds the answer to lift. The schema agent then adds the JSON-LD: Article, FAQPage, and any other relevant types, with the properties that make the page's meaning explicit. The shared vocabulary it draws on is the schema.org standard the engines read. These are separate agents because structuring prose and writing valid schema are genuinely different jobs, and the most common failures in schema are specific and avoidable, as catalogued in schema mistakes most stores make. Keeping schema as its own agent, with its own validation step, is what stops those errors reaching production.
Structure agent | Schema agent |
|---|---|
Answer-first opening | Article and FAQ markup |
Question-shaped H2s | Correct properties |
Tables for comparison | Valid JSON-LD |
Clear hierarchy | Explicit page meaning |
Built for extraction | Built for parsing |
The reason these stages come after writing, not before, is that you structure and mark up a real draft, not an empty plan. But because the writing agent already wrote to a planned structure, the structure agent is refining rather than rebuilding, and the schema agent is describing a page that was always meant to carry that markup. The detailed reasoning behind the schema choices these agents make lives in schema for AI engines.
What does the editor agent do, and where do humans stay?
The editor agent checks the finished page against the citation test and the house standard, deciding whether it ships or returns for rework. It is the gate. And it is the stage where human judgement stays closest, because the final call on whether a page is genuinely the best answer is one a person owns.
The editor agent applies the acceptance criterion the whole pipeline has been building toward: would an AI engine, asked the research agent's question, plausibly cite this page over the incumbent sources? It checks the page against the validation standard — word count, table use, question-shaped headings, distinct sources, internal links, and banned words — and flags anything that fails. But the editor agent is also where a human reviewer stays in the loop, because judgement about whether a page is honest, accurate, and genuinely the best answer is not something to fully delegate. The agent does the mechanical checks and a first-pass judgement; the human owns the final decision and the accountability that comes with publishing. This is the deliberate boundary: agents do the staged production, humans hold the points where judgement and responsibility matter.
Editor agent checks | Human reviewer owns |
|---|---|
The citation test | Final accuracy judgement |
Validation standard | Honesty of claims |
Banned words | Accountability for publishing |
Internal link count | Brand and reputation risk |
Structure compliance | The decision to ship |
The point of keeping a human at the gate is that AEO content makes claims a brand stands behind, and standing behind a claim is a human responsibility. The agents make the production reliable; the human makes the publishing accountable. This measurement-against-outcome discipline at the gate is the same one in how to measure AEO, applied to a single page before it ships.
How do the agents hand work between each other?
The agents hand work between each other through defined artifacts: each agent takes a specific input produced by the previous agent and produces a specific output the next agent needs. The handoffs are explicit, so nothing is assumed and nothing is lost between stages.
The flow is linear and inspectable. The research agent produces the question and the competitive record, which the strategy agent takes as input. The strategy agent produces the answer, proof, and sub-questions, which the writing agent drafts from. The writing agent produces a draft, which the structure agent organises. The structure agent produces a structured page, which the schema agent marks up. The schema agent produces a complete page, which the editor agent judges. Because each handoff is a defined artifact, you can inspect the work at any stage and see exactly what each agent received and produced. That is what makes failures traceable: a weak page is diagnosed by reading back through the artifacts to find the stage where the work went wrong, rather than rewriting the whole thing blind.
Handoff | From → To | The artifact passed |
|---|---|---|
1 | Research → Strategy | Question and competition |
2 | Strategy → Writing | Answer, proof, sub-questions |
3 | Writing → Structure | The article draft |
4 | Structure → Schema | The structured page |
5 | Schema → Editor | The marked-up page |
The discipline of explicit artifacts is what separates a real pipeline from a vague collaboration. When handoffs are implicit, work falls between stages; when they are defined, each agent knows exactly what it owes the next. This same artifact-based clarity is what makes the content brief function as the contract that runs through the whole system.
What does running this system actually require?
Running the system requires the six defined roles, the brief that specifies each page, a validation standard the editor enforces, and a human reviewer at the gate. The agents can be AI, people, or a mix, but the structure stays the same regardless of who fills the roles.
The non-negotiables are the structure, not the tooling. You need a clear brief so every page has a defined target and substance. You need each stage owned, so research, strategy, writing, structure, schema, and editing each have a responsible agent. You need a validation standard the editor applies consistently, so quality does not drift. And you need a human at the gate, so publishing stays accountable. Given those, the system runs whether the agents are language models, contractors, or in-house staff. The reason to use AI agents for most stages is throughput and consistency: a well-specified agent does its narrow job the same way every time, which is what makes the pipeline scale. But the system's reliability comes from the defined roles and handoffs, not from any particular model.
Requirement | Why it is non-negotiable |
|---|---|
A specifying brief | Every page has a target |
Each stage owned | No work falls between stages |
A validation standard | Quality does not drift |
Explicit handoffs | Failures are traceable |
A human at the gate | Publishing stays accountable |
The takeaway is that the content OS is a structure you can adopt, not a product you have to buy. Any team that defines these roles, writes the brief, sets the validation standard, and keeps a human gate can run it. The validation standard the editor enforces is the same one described throughout how to measure AEO, applied at the point of production.
How does Calibrate run client content on this OS?
Calibrate runs every client's AEO content through this six-agent OS: research sets the target, strategy decides the substance, writing drafts it, structure and schema make it extractable and machine-readable, and the editor agent plus a human reviewer gate it. No page ships until it passes the citation test and a person signs off.
In practice the OS is what lets Calibrate produce citable content consistently rather than occasionally. Each page enters as a topic and leaves as a finished, marked-up article that has been aimed at a real question, built to beat the incumbent sources, structured for extraction, and judged against the citation test. The system is the reason quality holds across a programme rather than depending on a single talented writer having a good day. It is also why production scales: adding capacity means adding to a defined pipeline, not training someone to do everything at once. The genuine outcome this kind of disciplined production supports is documented in the one fully measured case Calibrate publishes, the quarter-long visibility gain in the Cobbled Climbs case study, where consistent, structured content built standing over time.
Calibrate OS discipline | What it guarantees |
|---|---|
Research sets a real target | Pages aim at winnable questions |
Strategy decides substance first | No wandering drafts |
Writing executes the plan | Consistent house voice |
Structure and schema built in | Extractable, machine-readable pages |
Editor and human gate | Only citable pages ship |
The takeaway is that producing citable content at scale is a systems problem, and Calibrate solves it with a defined six-agent OS rather than heroic individual effort. To have your content built on this system, start with an AEO audit, with the full programme on the services page.
Frequently Asked Questions
Do the six agents have to be AI, or can people fill the roles?
The roles can be filled by AI agents, people, or a mix, because the system's reliability comes from the defined stages and handoffs, not from any particular tool. A small team might have one person play several roles in sequence, while a larger operation assigns each stage to a dedicated agent or specialist. The reason Calibrate uses AI agents for most stages is throughput and consistency: a well-specified agent does its narrow job the same way every time. But the structure is what matters. Any team that defines research, strategy, writing, structure, schema, and editing as distinct owned stages, with explicit handoffs and a human gate, is running the same system regardless of who fills the roles.
Why six agents rather than one model doing everything?
Splitting the work into six focused agents produces better output than asking one model to do everything in a single pass, because each stage rewards a different kind of focus. A model prompted to research, decide, write, structure, mark up, and edit all at once does each part adequately and none excellently. Giving each stage a narrow job and a clear standard raises quality at every step and makes the work inspectable, so failures can be traced to a specific stage and fixed there. Six is the number of genuinely distinct stages in citable content production: fewer blurs stages that need different focus, more adds handoffs without adding value.
Where exactly do humans stay in the loop?
Humans stay closest at the editor stage, which is the gate, because the final judgement about whether a page is honest, accurate, and genuinely the best answer is a human responsibility. The editor agent does the mechanical checks and a first-pass judgement against the citation test and validation standard, but a person owns the decision to ship and the accountability that comes with publishing a claim the brand stands behind. Humans also set the brief and review research on hard or sensitive questions. The principle is that agents do the staged production while humans hold the points where judgement and responsibility matter, rather than fully delegating either.
Does this system work for any industry, or just some?
The system works across industries because the stages of citable content production are the same regardless of subject: research the question and competition, decide the answer and proof, write it, structure it, mark it up, and edit it against the citation test. What changes between industries is the content the agents produce, not the pipeline that produces it. A page about cycling components and a page about professional services go through identical stages; the research, the answer, and the proof differ, but the structure of producing them does not. That portability is why Calibrate runs the same six-agent OS for every client rather than rebuilding the production system for each new vertical.
How does the content brief relate to the six agents?
The content brief is the specification the six agents fill and execute, so the system is the brief turned into a production line. The brief's fields map almost exactly onto the agent roles: the question and current citations are the research agent's output, the answer, proof, and sub-questions are the strategy agent's, structure and schema are their respective agents' fields, and the citation test is what the editor agent applies. Running the OS is, in effect, filling the brief stage by stage, with each agent owning the fields that match its job. The brief is the contract; the agents are the team that delivers against it, which keeps production aligned to a single specification.
What stops the agents from producing generic, low-quality content?
Several controls keep quality high: the research agent aims every page at a real, winnable question rather than a generic topic; the strategy agent decides a specific answer and concrete proof before drafting; the writing agent is constrained to execute the plan without inventing claims; and the editor agent plus a human reviewer apply the citation test and validation standard before anything ships. Generic content comes from skipping the early stages and writing without a decided answer or proof. The OS prevents that by making research and strategy mandatory inputs to writing, so the writing agent always has a specific target and real substance to work from rather than a vague topic.
Can a small team run this without a lot of tooling?
A small team can run the system because it is a structure, not a product. The non-negotiables are the defined roles, the brief, a validation standard, and a human gate, all of which can be implemented with modest tooling or even manually at low volume. A solo practitioner can play the roles in sequence, using the brief as a checklist and the validation standard as the acceptance test, and still get the benefit of staged, inspectable production. Tooling helps with throughput as volume grows, but the reliability comes from following the stages honestly, not from buying a platform. The system scales down to one person as cleanly as it scales up.
How is quality kept consistent across many pages?
Consistency comes from the validation standard the editor enforces on every page and from the defined stages that every page passes through identically. Because each page is researched, decided, written, structured, marked up, and edited the same way, the output does not depend on a single writer having a good day. The editor agent checks each finished page against the same criteria — the citation test, word count, table use, question-shaped headings, distinct sources, internal links, and banned words — so anything that drifts from standard is caught before publishing. The combination of identical stages and a consistent gate is what holds quality steady across a whole programme rather than piece by piece.
Related Guides from Calibrate
The AEO Content Brief: 9 Fields, None From SEO — the specification the six agents fill.
The Citation Architecture Method — the method the OS operationalises.
How to Map the Questions Your Customers Ask AI — the research agent's core task.
Schema for AI Engines vs Schema for Google — the schema agent's reference.
How to Measure AEO: Citation Rate, Share of Voice, Position — the standard the editor enforces.
How Cobbled Climbs Got Cited for Premium Cycling in India — the system's outcome over a quarter.





