How To Tell If Your AI SEO Agency Is Working
What Padding Looks Like Screenshots of favourable answers with no indication of how many runs produced them. Industry news summaries that could have been written without opening your account. A rising score with no methodology. Traffic charts from unrelated channels included to fill space.
Refusals matter. A report that only contains successes is either describing a suspiciously easy month or omitting the parts that did not work, and the omitted parts are usually where the useful information is.
You cannot control those pages, but you can influence them. Claim and complete your listings. Correct factual errors where the platform allows it. Respond to reviews. Give journalists and analysts accurate material to work from. Where a comparison article about your category exists and gets your details wrong, a polite correction is often accepted.
Marketing copy does not get quoted. A paragraph of adjectives about your commitment to excellence contains nothing a model can attribute, so it is skipped in favour of a competitor who wrote a plain answer. Write the plain answer. chatgpt seo
Distinguish between a supplier who is failing and one who is reporting badly, because the remedies differ entirely. Ask for the raw answers and read them yourself before deciding. It is not unusual to find that sound work has been buried under a dashboard nobody understands, and fixing the reporting is far cheaper and less disruptive than replacing a team that is actually doing the job.
Screenshots of favourable answers with no run count, which say nothing about how many attempts produced them. Impressions or traffic from unrelated channels included to fill a report. And activity described in the language of effort, such as ongoing optimisation, with no countable output attached.
Because there is no independent scoreboard in this channel, an engagement can run for a year on the strength of a number the supplier produces. That is an unusual amount of trust to extend, and it makes knowing what to check more important here than in any other marketing channel.
How to Test Rather Than Trust Everything above is a starting hypothesis. Run twenty prompts in your own category across all three, from signed out sessions, recording the mode and the date, and count the cited domains for each.
Run each one across the assistants your customers use, and write down the answers verbatim. Do this from a signed out session so your own history does not colour the result. What you want at the end is a simple table: which prompts named you, which named competitors, and which sources got cited.
Before leaving, make sure you take the prompt set, the baseline archive and everything published. If those were not yours under the contract, that is a lesson for the next agreement rather than something to negotiate at the exit.
One overlooked cost is your own time. Every engagement in this field needs somebody inside the business to confirm figures, approve crawler changes and answer factual questions, and a plan that assumes this is free will stall. Budget a few hours a month explicitly and name the person, because the alternative is an agency waiting on answers and billing for a month in which little shipped.
It is also worth checking which assistant your customers actually use rather than assuming. The answer varies by profession, age and country far more than industry commentary suggests, and several businesses have built measurement programmes around a system their buyers never open. Adding one question to your enquiry form settles it in a fortnight and can redirect the whole effort.
Build the run into an existing routine rather than creating a new one. Measurement programmes in this field fail through quiet abandonment rather than through a decision, and a modest set attached to an established monthly process survives far longer than an ambitious one that depends on somebody remembering to start it.
If the budget is substantial, add the earned coverage work, which is the slowest and most expensive component and the one you genuinely cannot do quickly on your own. Buying that first, before the cheap fixes are done, is the most common way money gets wasted in this field. chatgpt seo
Pricing in this field is unusually opaque, partly because the work is new and partly because the absence of an independent scoreboard makes it hard for a buyer to tell whether they are getting value. That combination invites vague scoping.
Performance and Score Based Models Both sound aligned and both create problems. Payment tied to mentions creates pressure to shape the prompt set toward questions you already win, which is measurable improvement that means nothing.
Good versions read like this: mention rate on evaluation prompts rose from two in fifteen to six in fifteen, which we attribute to the three directory corrections completed in week two, though a competitor also stopped publishing during the same period.
When to Test More Often Three situations justify a tighter loop. During an active campaign where you need to attribute a specific change, weekly runs on a subset of prompts are reasonable, provided you accept the variance.