HomeCategory › Article

AI Search Optimization in 2026: The Founder’s Guide to Being Found by ChatGPT

I spent the second week of July 2026 doing something stupid. I went looking for the instruction manual.

Image credit: Startups World News

TL;DR

AI search optimization is real: Pew found an AI summary on 18 percent of Google searches in March 2025, rising to 53 percent once a query ran ten words or longer, and readers click a traditional link only 8 percent of the time when one is present, against 15 percent when it is not. But the tactics being sold to founders are mostly unverified vendor marketing, and the load-bearing statistics contradict each other by up to four times. The only rigorous study in the field, the Princeton-led GEO paper, found that what actually gets you cited is citing sources, adding quotations and publishing statistics, which is to say being quotable rather than being marked up.

Experts say

Nearly every number in the AEO industry comes from a company selling an AEO dashboard, and when you line those numbers up they disagree by four times on the most basic question in the category. The one non-vendor study, out of Princeton, found the winning move is to publish a fact worth quoting. That is not a tactic anyone can sell you, which is exactly why it is missing from every definitive guide. Your reputation now lives on Reddit, Wikipedia and Forbes, not on your homepage, and no file in your root directory is going to fix that.
What is AI search optimization, in plain terms?
It is getting your company named and cited inside an AI-generated answer instead of ranking as a link below one. The unit of success changed from a position on a results page to a mention in a paragraph. AEO and GEO are the two labels worth knowing; the rest are vocabulary someone wanted to own.
Do I need an llms.txt file to get cited by ChatGPT?
No. Among the 50 most-cited domains in AI search, only 6 percent carry the file, and crawler logs show AI bots request it in roughly one tenth of one percent of visits. Google has said publicly it does not support it. It is genuinely useful if you sell developer tools, because coding agents like Cursor and Claude Code do fetch it.
If I rank number one on Google, will AI Overviews cite me?
Not reliably, and be careful with anyone who quotes you a precise figure here. The share of AI Overview citations coming from top-10 organic results is reported as 76 percent, 38 percent, or 17 percent depending on which study you read. The first two are Ahrefs measuring itself eight months apart with rebuilt detection, and the third is BrightEdge counting something slightly different over sixteen months. What is clear is the direction: ranking and citation are coming apart, so ranking well is helpful and no longer sufficient.
Which AI engine should a small startup focus on?
Almost certainly ChatGPT, because Previsible’s July 2026 study of 6.77 million AI-referred sessions put it at 92.4 percent of trackable standalone AI referral traffic. Treat the others as separate channels rather than the same job, since Ahrefs found 86 percent of the top-cited sources are not shared across ChatGPT, Perplexity and AI Overviews. Perplexity is worth a second look if you sell to analysts, journalists or technical buyers.
What is the single most useful thing I can do in the next ninety days?
Publish one statistic nobody else has, with the methodology attached, and go get named on one property you do not own, such as a G2 profile or a real Reddit thread. The Princeton research found statistics and quotations lift visibility by up to 40 percent, with the size of the effect varying by domain, and the engines’ favorite sources are consistently places where other people talk about you.
I spent the second week of July 2026 doing something stupid. I went looking for the instruction manual.

Last Updated on August 14, 2026 by Taya Ziv

I spent the second week of July 2026 doing something stupid. I went looking for the instruction manual.

The premise seemed reasonable. Half your customers now start their research inside a chat window instead of a search box. There is an entire industry that has sprung up to tell you what to do about that, with acronyms and certifications and pricing pages. So I read the guides. All of them say they are definitive. All of them say 2026. All of them have a data section.

And somewhere around the fourth one I noticed the thing that made me want to write this instead: they do not agree with each other. Not on the edges. On the load-bearing numbers. Ask the simple question “do AI answers cite the pages that rank on Google,” and depending on which definitive guide you opened, the answer is 76 percent, or 38 percent, or 17 percent. Four times the spread, and not one of the guides tells you that those numbers were measured months apart. That is the tell. A moving number is being sold to you as a fixed fact.

So here is what this page is. It is not another tactic list. It is me sorting the evidence in this category into what is solid, what is vendor marketing wearing a lab coat, and what is left over that actually tells a founder what to do on Monday. The shift is real. The instruction manual is mostly fiction. Those two things are both true and almost nobody says the second one out loud, because most of the people writing the guides are selling the dashboard at the bottom of them.

What AI search optimization actually is

AI search optimization is the practice of getting your company named and cited inside an AI-generated answer, rather than ranking as a blue link underneath one. The unit of success is not a position. It is a mention: the assistant says your name, or quotes your number, or lists you among three options, and the reader never sees a results page at all.

You will see four acronyms for roughly this. Here is the honest translation.

AEO, answer engine optimization, is the older term. It came from the featured-snippet era and still fits that world of direct answers and People Also Ask boxes.

GEO, generative engine optimization, is the newer and more precise one. It targets the synthesized answer itself, the paragraph ChatGPT or an AI Overview writes from several sources at once.

LLMO and AIO are the same idea with different letters, mostly invented so someone could own a term.

The useful distinction is only AEO versus GEO, and even that is thinner than the conference agendas suggest. If you strip the vocabulary, all four are asking one question: when a machine writes a paragraph about your category, does your name appear in it.

Why founders should care

The click is disappearing. When Pew Research Center tracked the real browsing of 900 US adults across 68,879 Google searches in March 2025, an AI summary turned up on 18 percent of them, and on 53 percent of the searches that ran to ten words or more. On the pages that carried one, users clicked a traditional result 8 percent of the time, against 15 percent on the pages without one. Meanwhile Semrush’s read of more than a billion lines of US clickstream data from October 2024 to February 2026 has ChatGPT referral traffic up 206 percent between January 2025 and January 2026.

Sit with the shape of that rather than the digits. The searches still happen. The answers still get delivered. What died in the middle is the part where the reader visits your site to find out. If you are not in the answer, you are not in the consideration set, and you will not see it happen in your analytics, because a customer who never heard of you generates no data. That is the whole reason this is uncomfortable. Traditional SEO decline shows up as a falling number. AI invisibility shows up as nothing at all.

This is the same structural story as the one where AI ate 60 percent of a company’s search traffic and the part that survived converted nine times better. Fewer visitors, far higher intent. The traffic did not get worse. It got smaller and more serious. Which means the cost of being left out of the answer went up, not down.

Every engine is a different country

The AI engines mostly do not cite the same websites. This is the finding that should reorganize your thinking before any tactic does. An analysis of 680 million citations, published by the AI visibility firm Leapd in April 2026, found that only 11 percent of domains are cited by both ChatGPT and Perplexity. Notice two things about that sentence. Leapd is a vendor in the exact category I am about to grade, and its post never says who ran the 680 million citation analysis, or when. Hold it at arm’s length. So do I. What keeps it on this page is that a study with a name and a date on it lands in the same place: Ahrefs lined up the 50 most-mentioned sources for AI Overviews, ChatGPT and Perplexity in June 2025, across 76.7 million AI Overviews and roughly 950,000 prompts on each of the other two, and found exactly 7 domains common to all three lists. Which is to say 86 percent of the top sources are not shared.

That is not a rounding difference. That is three separate channels wearing the same costume.

Their sourcing personalities are genuinely different:

ChatGPT leans encyclopedic and communal at the same time. Ahrefs’ Brand Radar tracking of broad US queries in July 2026 puts Reddit first at 16.7 percent of its citations, Wikipedia second at 8.9 percent and Forbes third at 3.3 percent. It runs on two layers: a static training base, plus a Bing-powered retrieval layer that mostly wakes up for commercial-intent queries, the ones containing words like “reviews,” “comparison,” “best,” or a year. It is also stingy about naming anyone, though I have left the popular figure for that out of this piece, because every version of it traces back to one vendor post that never says who ran the study.

Perplexity is the opposite animal. It performs a live web search on every query and carries no knowledge cutoff, which is why Ahrefs describes its citation path as retrieve, rank, quote, cite while ChatGPT’s default path is recall with no citation at all. Perplexity loves Reddit, which is not an accident: Reddit is real people answering the exact question in the exact words a person would use.

Google AI Overviews sit in between, drawing from Google’s own index. In Pew’s audit of the summaries themselves, Wikipedia, YouTube and Reddit were the three most-linked sources and together accounted for 15 percent of everything cited, while .gov domains showed up three times as often as they do in ordinary results. It is reading off a different shelf than the other two.

The practical read: “optimizing for AI search” as a single project is like running one campaign on LinkedIn and TikTok and calling it a strategy. And if your instinct is to chase all of them at once, notice that Previsible’s July 2026 study of 6.77 million AI-referred sessions across 166 sites puts ChatGPT at 92.4 percent of trackable standalone AI referral traffic between November 2024 and May 2026. Previsible sells AI-discovery consulting and the study leaves AI Overviews out entirely, so take the direction and leave the decimal. There is one country worth emigrating to first, and the rest are worth a postcard.

The part nobody grades: how good is this evidence?

Almost every statistic in this category comes from a company selling AI visibility software. That does not make the numbers false. It does mean nobody is checking them, and when I lined them up, several were checking each other and failing. So, grades.

This is the section I actually wanted to write, because it is the one missing from every guide I read.

Solid. The Semrush clickstream work is measurement at scale, over a billion lines of panel data, and it says AI referrals are growing fast off a small base. Semrush also sells an AI visibility product, so by my own rule you should discount it. I am not, and here is why, because it is the whole standard: a clickstream panel can be caught being wrong. A vendor survey cannot. Grade the checkability, not the logo. Pew’s 8 percent click-through finding is independent research from a nonpartisan fact tank with no dashboard to sell, run on what 900 people actually did in their browsers rather than what they told a survey. Trust both of those. Then there is the line I had in this draft until I checked it: that AI Overviews now show up on half of all searches. Pew measured 18 percent in March 2025. I also had a line here putting Semrush’s tracker under 16 percent by late 2025, and I have cut it back to a shrug, because I could not open a Semrush page that states the figure. It reached me through a roundup, which is the exact failure this section is about. The near-half numbers all come from trackers pointed at commercial keyword sets, which is a different question wearing the same words. Retire that one.

Solid and almost never cited. There is exactly one piece of non-vendor science aimed at the question this whole category claims to answer, which is how you get cited, and it predates the hype. In November 2023 a Princeton-led team, with IIT Delhi, published the paper that coined the term generative engine optimization, later accepted to KDD 2024. They built a benchmark of 10,000 queries drawn from a spread of domains and sources, then tested what moves visibility inside a generated answer. Three things won: citing your sources, adding direct quotations, and adding statistics. The paper’s headline is a lift of up to 40 percent, and the same abstract says plainly that how well each tactic works varies by domain. So treat 40 as a ceiling in a benchmark, not a promise on your blog. Even so: not markup. Not a file in your root directory. Being quotable.

It is a small irony worth noticing that the only rigorous study in generative engine optimization concluded that the way to get quoted is to be worth quoting, and the industry built on top of it sells everything except that.

Shaky. The stat that AI Overviews cite the top 10 organic results 76 percent, or 38 percent, or 17 percent of the time. Two of those are the same shop measuring twice. Ahrefs put it at 76.1 percent in July 2025, off 1.9 million citations from a million AI Overviews, then at 38 percent in March 2026, across 863,000 SERPs, after rebuilding how it detects a citation. The 17 is a different measurement entirely: BrightEdge’s sixteen-month tracker, published September 2025, found 16.7 percent of citations coming from the top 10 while overall overlap with organic climbed from 32 to 54 percent. All three are in circulation right now, undated, presented as the current answer, and that is worse than a disagreement. It means the guides quoting them never checked when the ground moved, or what was being counted when it did.

Shakier. The conversion numbers. One widely repeated pair says AI referral traffic converts at 1.66 percent against 0.15 percent for organic, roughly 11 times better. Another says ChatGPT converts at 14 to 16 percent against Google’s 1.76. Both get published as “AI traffic converts better.” They are an order of magnitude apart, and here is the part that should end it: the vendor post where the first pair circulates credits it to “aggregated LLM traffic studies” and never names one of them. The second pair I could not trace to any named study at all. The direction is probably right, because someone arriving after an assistant recommended you has been pre-sold. The magnitude is unknowable, and any founder building a forecast on it is building on sand.

If you take one thing from this page, take the habit rather than the number: in this category, ask who ran the study and what they sell before you let it into a board deck.

The llms.txt question, answered honestly

You have been told to add an llms.txt file. It is the single most common tactical recommendation in the category, and it is the easiest to test, because crawlers leave logs.

The short answer: it is close to useless for AI search visibility today, and genuinely useful for something else.

llms.txt is a proposed standard, a plain markdown file at your root that hands an AI a clean guide to your content. Nice idea. Adoption sits at 10.13 percent of the 300,000 domains SE Ranking crawled for its November 2025 study, which also found no relationship at all between having the file and being cited. And here is the detail that ends the argument, from a different shop entirely: Trakkr scanned 37,894 AI-cited domains in March 2026 and found that only 6 percent of the 50 most-cited domains carry an llms.txt, with adoption climbing the further down the citation rankings you go. The pages winning the thing the file is supposed to win overwhelmingly do not have it. Both SE Ranking and Trakkr sell AI visibility software, and both published a result that makes a popular deliverable look worthless, which is the kind of self-harm that makes me believe a number.

The crawler logs say the same thing louder. OtterlyAI put an llms.txt at the root of a live site and read its server logs for 90 days, publishing in February 2026: 62,100 AI bot visits, 84 of them to the file, about one tenth of one percent. The site’s average content page pulled around 265 in the same window. GPTBot, ClaudeBot, PerplexityBot and the rest walk straight past it and read your HTML like everyone else. No major provider, not OpenAI, Anthropic, Google, Meta or Mistral, has committed to using it as a ranking or retrieval signal. Google’s own guidance on succeeding in AI search does not mention the file once, and Gary Illyes said out loud at Google’s Search Central Deep Dive in July 2025 that Google does not support it and has no plans to. That last one reached me through Search Engine Land’s report from the event, which quotes Kenichi Suzuki writing up the Q&A from the room, rather than a Google transcript, so file it as strong, not primary.

So why does every guide still recommend it? My read: it costs nothing to recommend, it is easy to sell as a deliverable, and nobody checks.

Now the useful part, because I am not telling you to skip it. llms.txt genuinely works for coding agents. Cursor, Claude Code, Copilot, Cline and Aider routinely fetch it when they are pointed at a documentation site. If you sell a developer tool, that file is a real distribution channel into the agent your customer is coding inside right now. That is a good reason to ship one. “It will get me cited by ChatGPT” is not.

What survives every dataset

Strip out everything that only one vendor believes, and four things are left standing. These are the ones every source agrees on, including the academic paper and the crawler logs.

One. Be the source of a fact, not the author of a post. The Princeton work is unambiguous: statistics, quotations and cited sources are what lift you into a generated answer. An AI cannot quote a vibe. It can quote a number that exists nowhere else. Original data is not a content-marketing luxury in 2026, it is the raw material of citation. If your entire blog is a rewrite of what everyone already knows, there is nothing in it to lift.

Two. Your reputation lives on other people’s websites. ChatGPT’s two most-cited sources are Reddit and Wikipedia. Perplexity’s favorite is Reddit as well. Google’s AI summaries lean on Wikipedia, YouTube and Reddit. Every one of those is a place where other people talk about you. G2 profiles, review sites, community threads, podcast transcripts, all of it feeds the model’s picture of your category more than your own homepage does. Founders spend almost their entire content budget on the one property the engines trust least, which is the property that is obviously trying to sell them something. This is the trap behind the finding that 92 percent of brands are invisible to ChatGPT. They are not invisible because their markup is wrong. They are invisible because nobody outside the company has ever written their name down.

Three. Front-load the answer. The mechanism in the Princeton work is that a generated answer is assembled from liftable fragments, which means a paragraph has to survive being pulled out of your page and dropped into a chat window. If it would not make sense on its own out there, it will never be pulled. That is why the sections on this page open with the answer and argue afterwards.

Four. Pick one engine. With 86 percent of the top-cited sources not shared across the three engines, and ChatGPT owning the overwhelming share of standalone referral traffic, “we are doing AEO” is not a plan. Which engine, for which query, is a plan.

What to do in the next ninety days

If you are a founder with no budget and no SEO team, here is the whole thing, in order.

  • Find out where you stand. Open ChatGPT and Perplexity and ask the five questions your buyer would ask before they have heard of you. “Best tool for X.” “X versus Y.” Write down who gets named. That list is your real competitive set now, and it is often not the one in your deck.
  • Go get named somewhere you do not own. One G2 profile, one honest Reddit answer in the community where your buyers actually live, one podcast. This is slower than shipping a file and it is the thing that works.
  • Publish one number nobody else has. You have data. Your funnel, your onboarding times, your churn reasons, your survey of 40 customers. Publish it with the methodology attached. That is a citation magnet, and it is the only kind of content an AI cannot get from your competitor.
  • Restructure, do not rewrite. Put the direct answer in the first 40 words under every heading. Keep the headings as the questions people actually type. This is an afternoon of work.
  • Ship llms.txt only if you sell to developers. Otherwise leave it.
  • Do not block the crawlers. Check your firewall and bot rules. Plenty of companies are paying for AI visibility consulting while a security plugin quietly blocks GPTBot at the edge. This one is free and it is the most common own goal in the category.

And a note on the sixth point, which is the whole ballgame in miniature: the tactics that matter are boring and mostly not for sale. Nobody builds a SaaS product around “get one honest Reddit thread.” So nobody markets it. So it does not end up in the definitive guide.

The bit that should worry you

Here is my actual read, and it is less comfortable than the tactic list.

The discovery layer is being rebuilt, and the rebuild has a bias in it. Wikipedia at the top. Forbes and G2 next. Reddit everywhere. These are institutions and archives and communities, things that accumulate slowly over years. Meanwhile a great deal of what the model cites was published before you existed. The machine has a memory, the memory is old, and you are not in it.

That is a moat with an incumbent standing on it. A startup founded this year has no Wikipedia page, no ten-year archive, no accumulated mentions. The old discovery layer at least let you buy your way in with a credit card and an ad account, and the founders who win the ChatGPT ad era mostly will not do it by buying a single ad either, which is the same lesson arriving from the paid side. Even ChatGPT’s pay-per-sale ads, which look a lot like the early days of AdWords, only buy you the slot next to the answer, not a place inside it.

So the honest strategy is not a hack. It is to become the kind of company other people cite, which is the same advice that has worked for twenty years, now with much higher stakes and a much longer lag. The reason everyone is selling you a technical fix is that the technical fix can be delivered in a sprint and the real answer takes eighteen months.

Start the eighteen months today. And when the next definitive guide lands in your inbox with a number in the headline, check who is selling the dashboard.

Enjoyed this analysis?

Get stories like this in your inbox every Monday morning.

You Might Also Like