GEO & AI Visibility · 2026-09-28 · 13 min read
Getting found in ChatGPT: the seven levers by which AI systems read, cite and recommend mid-market websites

Michael Kaiser
Co-Founder & Head of Systems, Vincency
The short answer first: whether ChatGPT, Perplexity or Google's AI Overviews recommend your company is decided by seven things, and none of them is magic. Generative engines synthesize answers from several sources and cite the text that is easiest to quote: the one that answers directly, defines terms, states numbers and lets its crawlers in at all. The Princeton study that named this field measured up to 40 percent more visibility through targeted optimization. For a mid-sized company that compresses into a checklist you can run in an afternoon, and run it yourself.
This guide is written for the company that owns a website, not for the one that sells SEO. Every step names what to check, what good looks like and where the free self-test sits if you want the answer graded.
How the machines actually read you
A generative engine does not rank ten blue links; it synthesizes one answer out of several sources. Two separate moments decide your presence in it: retrieval, when the engine's crawler or search index is allowed to see your page, and grounding, when the model assembles the answer and picks the text it can reuse. You can lose at either gate. Blocking the crawler is the silent failure: nothing downstream can recover it.
The upstream evidence is worth knowing. The GEO study by Aggarwal et al. (arXiv 2311.09735, accepted at KDD 2024) formalized generative engines and measured the levers: visibility gains of up to 40 percent, with effectiveness varying by domain, which is why the checklist below is ordered by leverage, not by fashion. And the reason it matters commercially is the Pew finding on AI summaries: users who saw an AI summary clicked a traditional result in 8 percent of visits, against 15 percent without one, and clicked a source cited inside the summary in just 1 percent of visits. The click is migrating into the answer; your text has to be inside it.
The seven levers, in order of leverage
| Lever | What it fixes | Effort |
|---|---|---|
| 1. Let the crawlers in | robots.txt and meta tags decide whether GPTBot, ClaudeBot, PerplexityBot and the search-grounded crawlers may read you at all. A block is absolute. | Minutes |
| 2. Answer first | The first paragraph of every money page must be a complete, self-contained answer to the question your customer actually asks. | Hours of rewriting |
| 3. Define your terms | Models quote definitions. A clean one-sentence definition of what you do, who for and where is the most citable text you own. | Hours |
| 4. Publish llms.txt | A curated markdown map of your key pages at the site root, versioned standard since August 2026. Agents use it as the entry door. | An hour, or minutes with the generator |
| 5. Structure the facts | Schema.org markup (Organization, FAQPage, defined terms) turns your statements into machine-readable facts instead of prose to parse. | Hours, once |
| 6. Stay current | Visible update dates and current numbers; models prefer citing sources whose claims are timestamped. | Ongoing habit |
| 7. Be consistent everywhere | Same company name, location, service wording across site, profiles and directories; engines reconcile entities across sources. | A cleanup pass |
Three of these deserve the second look. Lever 1 is binary and the most common failure: plenty of companies block AI crawlers out of an abundance of caution that was never a decision, just a default. Lever 2 is where the writing changes: „We are your partner for digital solutions" tells a model nothing it can repeat, while „Vincency is a digital agency in Frankfurt that builds websites and AI automation for mid-sized companies" is a sentence it can cite whole. Lever 4 is the one with the strongest signal-to-effort ratio: the llms.txt specification moved to version 2 on 10 August 2026, thousands of sites publish one, Lighthouse checks for it, and the AI labs themselves publish their own.
What does not work
Three habits spend effort without return. Keyword density written for 2010 is noise to a model that paraphrases; it wants the fact, not the phrase. Mass-generated thin content produces pages nobody cites, because a model has no reason to quote interchangeable text. And treating AI visibility as a one-time project misses that grounding is live: the engine answers from what it can fetch today, so a stale page is a stale answer about you.
How to measure whether it works
Since 31 August 2026 the instrumentation exists without paid dashboards: Google Search Console carries a dedicated report for generative AI, and Bing Webmaster Tools exposes citations and grounding queries, albeit on a tenth of the volume. We laid out what each report actually measures, including the clicks Google withholds, in the piece on measuring AI visibility. The practical loop for a mid-market company is quarterly: run the self-test, fix the lowest lever, re-check after the engines have re-crawled.
For the first pass you do not need tooling at all beyond your own site and one honest read. Our AI visibility check runs through these levers as twelve weighted questions and scores you from 0 to 100, including which gap to close first. If you want the machine-readable map itself, the llms.txt generator produces a standards-conformant file in minutes; what the format is and why agents read it sits in the glossary entry on llms.txt and the one on GEO itself.
The honest remainder
Two qualifications keep this honest. First, the 40-percent figure is a benchmark ceiling, not a promise: the study measured across domains, effectiveness varied, and a mid-market website in a thin B2B niche has a different baseline than a consumer brand. Second, this field moves monthly, not annually; the crawler list and the llms.txt convention will keep changing, which is an argument for the habit, not for waiting.
Whoever takes one sentence from this piece: AI visibility is won at the boring end, not the clever one. Let the crawlers in, answer first, define yourself in one sentence, publish the map. The companies that get recommended are rarely the ones with the cleverest copy; they are the ones whose websites give a machine something it can repeat.
Related service
Digital projects in the Rhine-Main region
This article compares providers and approaches. If you want to get specific: we support mid-sized companies in Frankfurt and the surrounding area with websites, system integration and AI automation, with one fixed point of contact instead of rotating project teams.
Our Frankfurt servicesFrequently asked questions about visibility in AI answers
What decides which companies ChatGPT recommends?
Two things: whether the crawlers are allowed to read your content, and whether what they read is citable. Generative engines synthesize answers from multiple sources and prefer text that answers directly, defines terms cleanly, names numbers and is structured. A homepage of marketing filler gives the model nothing graspable; a page that answers a question in its first paragraph gives it plenty.
Does llms.txt actually do anything?
It is a standard, not a ranking: the file alone produces no citations, but it hands agents a curated map of your content before they have to guess. The specification was updated to version 2 on 10 August 2026, thousands of sites publish it, Chrome Lighthouse checks it as part of its agentic audits, and the AI labs run their own. It is the digital equivalent of the floor plan at reception.
Is ranking on Google enough to appear in AI answers?
No. Ranking and citation are related but separate games. The GEO study showed targeted optimization can lift visibility in generative answers by up to 40 percent, independent of classic position. Meanwhile, pages optimized only for clicks lose when the AI delivers the answer itself: per the Pew study, users click less once an AI summary appears.
What is the fastest first step for a mid-market company?
The inventory, not the rebuild. Check whether your robots.txt blocks the AI crawlers, whether your main service page answers in its first paragraph the question your customer actually asks, and whether name, location and service are stated identically everywhere. Our free AI visibility check walks through exactly these questions and produces a score from 0 to 100.
Is GEO just SEO under a new name?
The overlap is large, but the target mechanism differs. SEO optimizes a position in a list that gets clicked. GEO optimizes for your text being quoted as part of a generated answer, meaning a machine must be able to read, understand and rephrase it. Whoever ranks well but writes unreadably for models loses the second channel.
Sources, status and note: The study and the specification were read at their primary sources, retrieved on 28 September 2026: Aggarwal et al., „GEO: Generative Engine Optimization" (arXiv:2311.09735, KDD 2024) for the visibility framework and the up-to-40-percent figure; the llms.txt specification v2 (Jeremy Howard, modified 10 August 2026) for the format, its adoption and the Lighthouse audit mention; and the Pew Research finding on AI summaries for the click-through decline. This is a practical field guide, not a guaranteed outcome; generative engines change their sourcing continuously.
Related insights