What LLM SEO actually is
LLM SEO is one of five names for the same job, and it is the name on this page rather than the other four because it is the one people type.
The vocabulary is the least interesting thing about it. What matters is that an assistant asked to recommend a business writes a paragraph, leans on a handful of sources, and stops, so the habits built over twenty years of climbing a list of ten do not transfer cleanly. Ranking is a ladder and every rung is worth something; citation is a door, and you are either through it or you are standing outside it entirely.
This page is the explanation rather than the sales pitch. How a large language model arrives at the handful of sources it names, why plain writing and outside corroboration matter more than they did, what actually makes a page quotable, and the genuinely unsolved problem of measuring any of it. If you want the service instead – the deliverables, the prices, how to tell a real practitioner from a renamed deck – that is the other page and it is linked further down.
Fair warning before you spend twenty minutes here. My honest read is that this field is about seventy percent old work done more carefully and thirty percent new thinking, and a lot of people are selling the ratio the other way around.
LLM SEO is the practice of getting a language model to name your business, your page or your product when somebody asks it a question you should be the answer to.
The word SEO is doing a lot of work in that phrase and it is slightly wrong, because there is no engine and no ranking in the traditional sense. What there is instead is a selection: from everything the model has absorbed and everything it can fetch in the moment, a few sources get named and the rest do not exist for the purposes of that conversation.
So the object of the work is inclusion, and inclusion is binary in a way that ranking never was. That reframing is most of what a business owner needs to take away, because it changes which efforts are worth making. Being the seventh-best page on a subject was always worth something in search. It is worth close to nothing here.
Why a citation behaves differently from a ranking
In search, the results page is a list and your job is to climb it. Every place you move up is worth measurable traffic, and the relationship between position and clicks is well understood by everyone who has ever looked at a curve of it.
An answer works differently. The model writes a paragraph, and if it attributes anything it attributes to a small number of sources chosen because they supported the specific claims it made. You are not competing to be the best page about plumbing in Kennewick. You are competing to be the source that most cleanly supports the specific sentence the model is about to write.
That is a subtler target and it explains a lot of otherwise confusing behavior. A page that ranks eleventh can get cited because it states one fact clearly that the top ten bury. A page that ranks first can get skipped because it is a long sales page that never plainly says the thing.
The other half of the difference is that a citation does not necessarily send you a visitor. Somebody reading an answer may act on your name without ever loading your website, which means the value arrives as a phone call, a branded search, or somebody typing your name into maps a day later. It is real value with a broken paper trail.
Where an answer gets its facts
There are two sources feeding any answer you read, and telling them apart is the most useful mental model in this whole subject.
The first is what the model learned when it was trained, which is a compressed, lossy impression of an enormous amount of text. Nothing in there can be looked up the way a database is looked up; what survives is a blur of everything it ever read, remembered imperfectly and blended together. Your business is in there only if it was written about enough, consistently enough, for a general impression to survive the compression. For most small businesses, it is not.
The second is retrieval, which is the model going out and fetching current sources during the conversation and writing its answer from what it just read. This is where nearly all small business visibility actually happens – and it is the half you can influence on a timescale that matters. Retrieval-augmented answers are why a business that opened last year can be named at all, and why a change you make to a page can show up in an answer weeks later rather than never.
The practical implication is worth stating plainly. You are not trying to get into the training data, which you cannot do on purpose and which would take years to matter. You are trying to be what gets retrieved, read and quoted at the moment somebody asks.
Retrieval is the part you can influence
When a model retrieves, it is usually running a search of some kind behind the scenes, reading a handful of results, and composing from them.
That means the old search game still sits underneath the new one. If your page cannot be found by a conventional search for the phrase somebody used, it is not in the pool of candidates and nothing else you do matters. This is the first reason I keep telling people that ordinary search work is the foundation rather than the alternative.
But being retrieved is only half of it. The model reads what it fetched and then decides what to actually use, and that second filter is where the new work lives. It will favor a source that answers the question directly over one that circles it, a source whose claim is easy to restate without distortion, and a source that agrees with the other things it just read. A page that contradicts the consensus of the other results can still be the correct one, and it is still unlikely to be the one quoted.
There is a small, useful consequence of this. Making a page easier to quote is a different activity from making it rank, and it is cheaper. It is often a matter of moving one sentence to the top of a section. One sentence.
Why models favor sources that state things plainly
A model summarizing your page is doing extraction, and extraction rewards writing that puts the claim where a reader can find it.
Consider two pages about the same thing. One opens with three paragraphs of positioning and gets to the actual price in the eighth paragraph, hedged, next to an invitation to book a call. The other says a small business site runs six to fifteen thousand dollars and then explains what moves it inside that range. The second page is going to be the one quoted, and it will be quoted correctly, because the claim was sitting in a sentence that survives being lifted out of context.
That last part is the criterion worth internalizing. Write sentences that stay true when removed from the page. A claim that only makes sense with the previous four paragraphs attached is a claim the model will either skip or mangle, and a mangled quote is worse than no quote because you cannot correct it.
The same logic covers specificity. Real numbers, real place names, real dates and real category words are all easier to extract and harder to confuse with somebody else’s page. Vagueness is not just weak writing here; it is unquotable writing – and unquotable is the whole failure.
I find this genuinely encouraging, because the writing that machines prefer turns out to be the writing that impatient humans preferred all along. Nobody ever wanted the eight paragraphs.
Corroboration, or why the model wants to hear it twice
A model reading a single page has one source for a claim. A model that has read the same claim on four unrelated sites has something closer to a fact.
This is the mechanism behind the advice you keep hearing about mentions and digital PR, and it is worth understanding rather than just obeying. When several independent sources agree, the model’s confidence in restating the claim goes up, and confidence is roughly what determines whether it will put your name in an answer or hedge into generality. When sources disagree, the safe move for the model is to say nothing specific, and saying nothing specific means not naming you.
Your own website is the weakest kind of evidence for this purpose, not because it is untrustworthy but because it is one source with an obvious interest. A local news write-up, an industry directory a human maintains, an association listing, a supplier’s partner page, a podcast episode description, a conference program – each one of those is another independent voice saying the same thing about you.
The claims worth corroborating are the boring ones. What you do, where you are, who you serve, what you are known for. Nobody needs four sources agreeing about your brand values.
Entities, and the machine's idea of your business
Underneath all of this is a question the machine has to settle before anything else: is this one business, or several?
An entity is the machine’s internal idea of your business as a distinct thing, separate from every other business with a similar name in a similar category. That idea is assembled from every mention it has ever seen, and it is only as coherent as those mentions are. A business whose name appears three ways, whose address still has an old suite number on half the internet, and whose category is listed as four different things does not resolve into one confident entity. It resolves into a smear. Nothing lands.
The consequence is not that you rank lower. The consequence is that the model is less sure you are a real, specific business, and uncertainty is fatal to being named, because naming a business is a specific act. Generality is always the safe answer for a model, and an ambiguous entity pushes it toward the safe answer.
This is why the least glamorous line on any AI SEO invoice, the listings clean-up, is also the most load-bearing. It is not about the directories themselves. It is about giving the machine one consistent story to compress.
What makes a page citable
Assume the model has retrieved your page and is deciding whether to use it. Six things tip that decision, and none of them are exotic.
The page answers one question completely rather than four questions partially. It states the answer near the top of the relevant section instead of building to it. It uses the vocabulary a customer would use rather than the vocabulary the industry prefers, because the question was asked in customer words. It contains specifics that can be attributed: numbers, ranges, timeframes, place names, conditions. It has a clear heading structure so the section boundaries are obvious, which matters more than people expect, because a model chunking a page uses those boundaries.
And it is current, or at least visibly dated. A page that says what year it is and what it was last reviewed gives the model a reason to prefer it over an undated page making the same claim.
What does not tip it: word count for its own sake, keyword density, a table of contents, or the phrase “in this comprehensive guide”. Long pages get cited constantly and so do short ones. Length was never the variable – not once.
Structured data does less and more than people think
Schema markup is not a ranking signal for language models and marking up your page will not make you get quoted. That is the honest ceiling on it, and it is lower than the pitch you will hear.
What structured data does is remove ambiguity. It says, in a form that requires no interpretation, that this is a business, that this is its address, that these are its hours, that this page is about this service, that these are questions and these are their answers. Every one of those statements is a thing the machine would otherwise have to infer from prose, and inference is where mistakes get made.
The most useful marks are the plain ones. The organization or local business. The service, on the page about that service. The question and answer pairs, because they map a page into exactly the shape a retrieval system wants to chunk it into. The person, where a person is the reason people choose you.
The failure mode is over-marking. Sites with half a dozen schema blocks installed by competing plugins, contradicting each other about the business name, are common and actively harmful, because you have now told the machine two different stories in the language reserved for unambiguous facts.
The citation set is narrower than the search results page
Ten blue links is a generous format. It can afford to include the plausible along with the certain, because the human sorts it out.
An answer cannot afford that, so it names few sources and it tends to reach for ones it has reason to trust: sites it has seen corroborated repeatedly, sites with obvious topical focus, and sites whose claims were easy to verify against the other things it just read. The set is narrower than the top ten – and it is narrower in a way that concentrates. The same handful of sources get named over and over inside a category.
That concentration cuts both ways for a small business. Breaking in is harder than getting to page one, because there are fewer slots and the incumbents are sticky. Once you are in the set for your local category, though, you tend to stay named, because the thing that got you in is the same thing that keeps the model confident.
I am not going to put a number on how much traffic this represents or how many people are asking assistants instead of searching. I do not know, the companies involved do not publish it usefully, and every figure I have seen quoted in a sales deck traces back to somebody’s survey of their own customers. The case for doing this work does not require a number.
LLM optimization, AI SEO, and one job with four names
LLM optimization, LLM SEO, AI SEO, answer engine optimization, generative engine optimization. The vocabulary is a mess – and it is a mess because the category is young and everybody wanted to name it.
They describe the same work. I use them interchangeably and so does almost everyone actually doing it, and when somebody insists on a taxonomy where their term is the advanced version of your term, they are selling the taxonomy.
What is worth distinguishing is not the names but the surfaces. An AI summary sitting on top of a search results page behaves differently from a standalone assistant with no search connection, which behaves differently again from an assistant that browses live before it answers. The first is closest to old search and rewards ranking well. The second depends almost entirely on training-era impressions and is nearly untouchable for a small business. The third is where the work pays, and it is the one growing fastest.
Measuring something that has no rank tracker
The industry is worst at this next part, and it deserves saying before anybody spends money on any of it.
There is no position to track. Two people asking the same question get different answers, the same person asking twice gets different answers, and the answer changes with the phrasing, the account, the location, the date and things nobody outside the company can see. Rank tracking worked because the search results page was close to deterministic for a given query and place – the same search from the same town gave the same list. Answers are not deterministic, and averaging a non-deterministic thing does not make it deterministic; it makes it an average with an unstated variance.
So when a tool sells you an AI visibility score, understand what it is. Somebody picked a set of questions, asked them on a schedule, counted mentions and turned that into an index. That can be a useful directional signal if the question set is yours and you can see the raw answers. It is not a measurement of your visibility, and the confident decimal point on it is decoration.
The second problem is attribution. When a model names you and somebody calls you the next morning, nothing in your analytics knows why. Some assistant traffic arrives with a referrer you can identify, plenty arrives as direct traffic, and the largest share may never arrive as a session at all because the person just picked up the phone.
What a sane measurement setup looks like
You can still do this honestly, and it costs almost nothing. It just refuses to produce a single number. That is the trade. The whole procedure runs to seven steps and about an hour a month.
1. Write down twenty questions a real customer would ask, in customer language, and keep your own business name out of every one of them. 2. Pick one day of the month and keep it, because a Tuesday in March and a Sunday in July are two different tests. 3. Ask all twenty in the same assistants, from the same location, signed out of your accounts wherever the tool allows it. 4. Save the full answers as text, rather than as a summary written afterward or a screenshot nobody will ever open again. 5. Log three things against each answer: whether you were named, who else was named, and which sources the answer leaned on. 6. Track branded search volume, direct traffic and identifiable referrals from assistant domains on the same sheet, month by month. 7. Ask every new customer how they found you and write down the answer they give you.
Over six months that log tells you something true, and the fact that it wobbles is information rather than noise. None of it is as satisfying as a rank chart. It is what is actually available, and I would rather hand a client a log they can read themselves than a score I invented.
Where this actually pays off
The businesses getting value out of this today are the ones where the customer does research before choosing.
Somebody picking a dentist for their family, a web shop for a rebuild, a winery to visit on a Saturday, a contractor for a job that costs real money – those are the questions people put to an assistant, because they are questions where a recommendation is worth more than a list. Emergency work is different. Nobody with a burst pipe opens a chat window.
The other thing that makes this pay is being in a category small enough that the model has a shot at knowing your town. A specific service in a specific small place is a much easier thing to be named for than a general service in a metro, and Walla Walla is exactly the kind of market where this is winnable. That is a local strategy rather than a national one, and local is what most of our clients are.
If you want the service version of all this – what gets built, what it costs, what to ask before you hire anyone – that is our page on the AI SEO agency side of the work. The wider program it belongs to is on small business SEO.
Why we do not sell this on its own
We do not sell LLM SEO as a standalone product, and I would be suspicious of anyone who does, because a business paying for citation work while its site is invisible in ordinary search is paying for the roof before the walls.
What the work is, and what it is not, is on the page for the same reason everything else here is. A page that tells you nothing until you ring costs you a phone call to learn very little.
How long it takes
The foundation is three to four weeks, most of it spent waiting on directories and platforms rather than working.
After that, expect the same shape as ordinary SEO. Three to six months before anything is visibly different and nine to twelve before it is clearly worth what you paid, with one wrinkle: because the measurement is weaker, you will feel less certain at every point along that line than you would with a rank chart in front of you.
Some things move faster. A structured data fix or a listings reconciliation can change what an assistant says about you within weeks, because retrieval reads the current web rather than a year-old snapshot. Mentions are the slowest thing on the list, because they run on other people’s schedules and other people are busy.
Common questions
What is LLM SEO?
It is the work of getting a large language model to name your business or your page when somebody asks it a relevant question. It overlaps heavily with ordinary SEO, and the part that is different is that answers name very few sources and there is no ranking position to occupy.
Is LLM SEO different from regular SEO?
It shares most of its foundation with regular SEO and differs in emphasis. Consistent entity information, corroboration from sites that are not yours, and writing that can be quoted out of context matter more; keyword targeting and position tracking matter less, because there is no position.
How do language models decide which sources to cite?
Mostly by retrieving current pages during the conversation, then preferring sources that answer the question directly, that can be quoted without distortion, and that agree with the other sources it just read. Sites the model has seen corroborated repeatedly get reached for more often.
Can I get into the training data?
Not on purpose, and it is the wrong goal. Training happens in cycles measured in months or years and nothing you publish this week is going to change a model’s general impression of your category. Retrieval is the half you can influence, and it responds far faster.
Does schema markup help with AI search?
It helps by removing ambiguity, not by acting as a ranking signal. Marking up your business, your services and your question-and-answer sections gives the machine unambiguous facts instead of things it has to infer, and cleaning up conflicting markup is usually worth more than adding new markup.
How do I know if it is working?
Ask the same set of customer questions every month, from the same place, and save the full answers. Watch branded search, direct traffic and identifiable assistant referrals underneath that. Any tool offering you a single visibility percentage has averaged away the variance that makes the number meaningful.
Should a small business spend money on this yet?
If your website already ranks and your listings are clean, the incremental cost of doing this properly is small and worth it. If neither of those is true, fix them first, because the same work feeds both and you would be buying the second floor of a building with no first floor.
Does being cited actually send traffic?
Sometimes, and less reliably than a search ranking does. A citation can end with somebody calling you without ever visiting the site, which is a real outcome with no session attached to it, and that gap is the main reason this is harder to justify on a spreadsheet than it deserves to be.