The short answer
To be cited, a page has to be fetchable by the engine's crawler, present in the raw HTML, and written so the relevant section answers one question on its own with concrete facts in it. Beyond your own site, engines lean heavily on how third parties describe you. There is no markup, file, or protocol that grants AI visibility on its own, and several of the ones being sold have been tested and found to do nothing.
In this part
- First, the boring prerequisites
- Write sections that survive being read alone
- The page types that actually earn citations
- Where else you need to exist
- Things that are widely sold and do not work
- What to do first, if you only have a month
- Five failure modes worth naming
- Where PromptScout fits in this part
- What is next
First, the boring prerequisites
Nothing else here matters if an engine cannot fetch and read the page. Most of the gap between brands that get cited and brands that do not sits in this layer, and it is the cheapest and fastest thing to fix.
Work through it in order. Each item is worthless if the one above it is broken.
The retrievability checklist
- The page is indexable. No stray
noindex, no accidentalrobots.txtblock, no canonical tag pointing somewhere else by mistake. - The right crawlers are allowed by name. If your site blocked AI user agents at any point, confirm you blocked the training crawler and not the search crawler.
GPTBotandOAI-SearchBotare different decisions. - The content is in the raw HTML response, not injected by client-side JavaScript after load. Check with
curl -s https://yoursite.com/pricing | grep -i "per month", not with the browser inspector, which shows you the page after JavaScript has run. - Nothing stands between the crawler and the text. No consent wall, no interstitial, no soft paywall, no login.
- Snippet controls are deliberate.
nosnippet,data-nosnippet, andmax-snippetall restrict what Google's AI features may quote from you. If someone added them years ago to protect featured snippet real estate, that decision now has a larger cost. - The page is reachable by an internal link, not only through search. Orphan pages get crawled unreliably.
- Titles, headings, visible text, and structured data agree with each other. Contradictions here reduce confidence in all of them.
Check your robots file first. The most common cause of a brand missing from ChatGPT search is not weak content. It is a
robots.txtrule written during the 2024 wave of AI opt-outs that blocked every AI-looking user agent, including the one that builds the search index. Two minutes to check, free to fix.
Write sections that survive being read alone
The unit of work is a section with a heading. It has to meet one standard: somebody reading only that section, with no page title, no navigation, and no surrounding paragraphs, still gets a correct and usable answer.
The same information, written both ways:
| Reads poorly in isolation | Survives retrieval |
|---|---|
| As mentioned above, it starts at a very affordable price point. | PromptScout plans start at $15 per month. Provider-backed monitoring requires choosing a plan before the first run. |
| Our integration options are extensive and flexible. | PromptScout connects to Google Search Console, Bing Webmaster Tools, and Cloudflare, and exposes an MCP server for agent access. |
| Many teams find it works well for their reporting needs. | Agencies use the weekly report to show one client which AI answers changed, which competitor gained ground, and which sources were cited. |
| It's great for growing companies of all sizes. | It suits teams tracking 20 to 200 prompts across up to five answer engines. Single-brand teams below that volume are usually better served by the free brand checker. |
The rewritten column is not better prose. It is more useful raw material, and the academic evidence points the same way: adding statistics, direct quotations, and citations to primary sources measurably improved a source's visibility inside generated answers in the KDD 2024 benchmark. Specifics get reused because they can be reused.
Seven habits that make the difference:
- Lead with the answer in one sentence, then justify it. Not the other way around.
- Name the entity in the sentence. Do not let a pronoun do work that points back at the H1 three screens up.
- Include the number, the date, the currency, and the unit. Vague ranges get skipped in favor of a competitor who committed to a figure.
- Attribute the facts you borrow, with a link. Pages that cite well tend to get cited well.
- Answer the awkward questions too. Who it is not for, what it does not do, what it costs at the expensive end. Fan-out generates disqualifying questions constantly, and the brand that answers them honestly is the one that gets recommended with a caveat rather than omitted.
- Keep one idea per heading. If a section needs two H3s, it is two sections.
- Write the summary sentence you would want quoted. If you cannot find one sentence in a section worth lifting, the engine will not find one either.
The page types that actually earn citations
Fan-out does not generate an even spread of sub-questions, so some page types pull far more weight than others. Roughly in order of return on effort:
Honest comparison pages
The highest-value page type, and the one most companies refuse to write well. Fan-out generates comparison questions constantly, and somebody will answer them. If you will not publish a fair comparison against your obvious alternative, a review site, a competitor, or an affiliate blog will, and theirs is what gets retrieved.
A page saying you win at everything reads as marketing and gets treated as such. A page that concedes something specific is quotable:
Linear is the better choice if your team is engineering-only and lives in keyboard shortcuts. Jira is the better choice if you report across twelve teams to a PMO. We are built for the case in between: one product team that has to show progress to people outside it.
That paragraph survives being read by a model that has also read three other accounts of the same comparison, because it is checkable rather than promotional.
Pricing pages with actual prices
"Contact us for pricing" is not an answer, so an engine will reach for whatever third party has guessed. That guess is usually wrong and rarely flattering. If you cannot publish a number, publish the shape: what the tiers are, what drives cost up, what a typical range looks like at a given size.
Integration and compatibility pages
"Does it work with X" is one of the most common constraint questions, and it is trivially answerable. One page per significant integration, each stating plainly what connects and what does not, is unglamorous work that shows up in citations.
Documentation
Specific, factual, structured, one question per page: exactly the shape retrieval prefers. Vendor documentation was a substantial share of what several engines cited in our monitoring. Docs behind a login are your most retrievable asset, removed from the web.
Use-case and job pages
Content organized around the problem rather than the product, for the many questions where the person has not yet named a category. These win the questions where nobody is currently mentioned.
What performs worse than teams expect
- Thought-leadership essays. Rarely contain retrievable specifics.
- Press releases. Written in a register models tend not to quote.
- Ungated whitepapers with the substance in a PDF. Often crawled inconsistently, and the summary that is on the page usually says nothing.
- Anything gated. A form in front of the content removes it from the web.
Where else you need to exist
Your own site is necessary and not sufficient. Engines describe you using whatever they can find, and much of it was written by other people.
The usual advice, "be everywhere," is expensive and mostly wrong. Where you need to be present depends on where the engines your buyers use actually look. In our window, Reddit carried 27.5% of ChatGPT's citations and between 0.2% and 4.3% of every other engine's. If ChatGPT is where your buyers ask, community presence is close to the highest-leverage work available to you. If they ask Gemini, the same effort buys very little.
Before commissioning any outreach, know that there is no fixed list of sites to get on. This is the citation distribution across every source our panel recorded.
PromptScout monitoring data
There is no short list of sites to be on
Cumulative share of 69,729 citations covered by the most-cited domains.
- Top 10 domains28.9%
20,141 of 69,729 citations
- Top 25 domains35.6%
24,790 of 69,729 citations
- Top 50 domains42.3%
29,481 of 69,729 citations
- Top 100 domains49.8%
34,716 of 69,729 citations
- All 7,944 cited domains100.0%
69,729 of 69,729 citations
share of all citations covered
The ten most-cited domains covered 28.9% of citations, and the top hundred still only reached 49.8%. The other half came from a tail of thousands, 4,181 of which were cited exactly once.
Source: PromptScout monitoring, May 28 – August 25, 2026. 5,436 completed answers across 135 tracked prompts. Aggregated across all monitored brands.
What the long tail means for you. Half of all citations came from outside the hundred most-cited domains, and more than half of the domains cited at all were cited exactly once. You cannot buy your way into this with a placement list. Being the most specific, most quotable answer to a narrow question wins citations that general authority does not.
The off-site work that holds up
In rough order of durability:
- Correct the basics wherever they are already published. Your category, what you do, who it is for, what it costs. Start with your G2 and Capterra profiles, your Crunchbase entry, your LinkedIn company description, and any "best X tools" roundup that already lists you. A stale description is worse than no description, because it is confidently wrong.
- Earn coverage that describes you in buyer language. A write-up that uses the words your customers use is more retrievable than one that uses the words your brand guidelines prefer.
- Take review platforms seriously. They are structured, comparative, and constantly updated, which is a retrieval-friendly combination.
- Answer questions in public, as yourself, where your buyers actually ask. Community answers are read as evidence by at least one major engine. Do this transparently; astroturfing is both dishonest and, given how these systems weight consistency across sources, unreliable.
- Keep your entity clear. Consistent name, consistent description, consistent category across everywhere you appear. Ambiguity about what you are is a bigger obstacle than obscurity.
Things that are widely sold and do not work
AEO attracted a lot of confident advice very quickly. Some of it has since been tested. Here is what held up and what did not, so you can stop paying for the second column.
| The claim | What the evidence says |
|---|---|
Publish an llms.txt file to guide AI systems |
Ahrefs examined 137,000 sites and found 97% of valid llms.txt files were never requested at all in the month measured. Google's Search team has said plainly it is not used for Search. |
| Add AI-specific schema markup | Google states there is no special schema.org structured data needed for AI Overviews or AI Mode, and no new machine-readable files or markup. |
| There is a separate AI ranking system to optimize for | Google says its AI features are grounded in the same core ranking and quality systems as Search, with no additional technical requirements. |
| Chunk your pages into AI-friendly fragments with special delimiters | Retrieval systems do their own chunking. Writing self-contained sections helps. Inventing a fragment format for machines that never asked for one does not. |
| Repeat your brand name so the model associates it with the category | Nothing in the published benchmarks rewards this, and it measurably degrades the page for the humans who actually read it. |
| Buy placements on "the sites AI cites" | The top hundred domains covered under half of citations in our panel, and the list is unstable quarter to quarter. |
One clarification on structured data, because that row gets over-corrected. Schema markup is still worth publishing: for rich results, for making your entities unambiguous, and for keeping machine-readable facts in sync with visible ones. It is not an AI visibility lever on its own. Publish it because it describes your page accurately, not because you expect a citation in return.
A fair test for any new tactic. Before adopting something, ask three questions. Who published the evidence, and do they sell the solution? Was the effect measured against a control, or observed after the fact and narrated backwards? Would it still be worth doing if no AI engine existed? A tactic that fails all three is a hypothesis wearing a best-practice costume.
What to do first, if you only have a month
Ordering matters more than completeness. Cheap, high-certainty work first, which is also the order that gives you something to show before the budget conversation.
Week 1: Make yourself retrievable
Audit robots.txt for AI user agents and fix anything that blocks a search crawler. Confirm your top ten commercial pages render server-side. Check snippet directives are intentional. Remove any consent wall or interstitial standing in front of key content.
Done when: every page you would want cited returns its key sentences in a raw curl.
Week 2: Establish a baseline
Pick 20 to 40 real buyer questions in the language buyers use. Run them across the engines your buyers actually use. Record, for every answer, whether you appeared, who appeared instead, and what was cited. Do not change anything yet.
Done when: you can state your mention rate per engine and name the five questions you most want to win.
Week 3: Repair what the answers already reach for
Look at which of your pages were cited and which competitor pages were cited instead. Rewrite the sections that read badly in isolation. Add the missing specifics: prices, versions, limits, integration names. This is editing, not commissioning.
Done when: every page in your cited set leads with a quotable, specific sentence.
Week 4: Fill the single largest gap
Publish the one page for the sub-question nobody in your category answers well. Usually that is a comparison, a constraint page, or an honest "who this is not for." Start only the off-site work your own citation data justifies.
Done when: the gap is covered and the next run is scheduled.
Resist the urge to change ten things at once. Answers move on their own between runs, which is the subject of Part 4. If you change your site, your comparison page, and your community presence in the same week, you will not be able to tell which one did anything, or whether anything did.
Five failure modes worth naming
These show up repeatedly in teams that put in real effort and get little back.
- Optimizing for the engine your buyers do not use. All the effort, none of the audience.
- Publishing volume instead of specificity. Twenty generic posts give an engine twenty ways to find the same non-answer.
- Fixing content while the page stays technically unreachable. The most expensive way to change nothing.
- Declaring victory on one run. Covered in Part 4, and it is how internal credibility gets spent.
- Refusing to write the honest comparison. The most reliable way to let a competitor define you.
Where PromptScout fits in this part
The website audit runs the retrievability checklist above against your actual pages, so you are not doing it by hand. Weekly Tasks turn monitoring evidence into one scoped piece of work at a time, keeping the reason, the sources behind it, and the window to watch afterwards together, which makes the one-change-at-a-time discipline the default rather than an act of willpower.
The website audit documentation covers what it checks and how findings are prioritized.
What is next
Part 4 covers knowing whether any of it worked: which metrics mean something, how to design a question panel that produces a trend rather than a series of snapshots, what run-to-run volatility does to your conclusions, and how to write a report that holds up when someone pushes back.
Common questions
- Does llms.txt help AI visibility?
- There is no evidence that it does. Ahrefs found that 97% of valid llms.txt files across 137,000 sites were never requested in the month they measured, and Google has stated that llms.txt is not used for Search. It is cheap to publish and reasonable as a documentation map for coding tools, but it should not be counted as an AI visibility tactic.
- Do I need special schema markup for AI search?
- No. Google states explicitly that no special schema.org structured data and no new machine-readable files are needed to appear in AI Overviews or AI Mode. Schema remains worth publishing for accuracy, entity clarity, and rich results, but it is not an AI visibility lever on its own.
- How do I get my brand cited by ChatGPT?
- Make sure OAI-SearchBot can fetch your pages, make sure the content is in the raw HTML, and write each section so it answers one question on its own with concrete facts. Then look at which sources ChatGPT actually cites for your questions, because in our window more than a quarter of its citations came from community discussion rather than from brand sites.
- Is there a list of sites I should get mentioned on?
- Not a useful one. In our window the hundred most-cited domains covered under half of all citations, and more than half of the domains cited at all appeared exactly once. Coverage of specific questions consistently outperforms placement on a general list.
- Should I write more content or fix existing pages first?
- Fix existing pages first, almost always. Most brands already have pages covering the sub-questions engines ask; those pages simply read badly out of context or are technically unreachable. Rewriting a section costs hours, and commissioning a new page costs weeks.
- Does publishing more content improve AI visibility?
- Only if the new content answers a question nothing else answers well. Volume on its own gives an engine more ways to find the same generic statement. In a market where more than half of cited domains appeared exactly once, specificity beats output.
Sources cited in this part
Primary sources are published by the party that runs the system. Third-party studies are labeled as such, because vendor research in this field disagrees more than the headlines suggest.
Google's statement that no AI-specific markup, files, or optimizations are required.
Which OpenAI crawler governs search visibility, and which one governs training.
Evidence that major AI crawlers fetch JavaScript without executing it.
- GEO: Generative Engine OptimizationAggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, KDD 2024Peer-reviewed research
Controlled evidence that statistics, quotations, and citations increase a source's visibility in generated answers.
Server-log evidence on whether AI systems request llms.txt at all.