Aug 26, 2026

Why Did My AI Visibility Score Drop? We Checked 18 AI Visibility Tools — Most Only Show That It Did

Tracking AI visibility across dated runs stopped being rare in 2026: 17 of the 18 tools we checked in August do it, and at least three of them do it free. Explaining which question and which engine caused the move is still rare. Here is what causal attribution looks like on a real eight-run record — and the statistical guards that decide when a number is even allowed to be called a change.

Why Did My AI Visibility Score Drop? We Checked 18 AI Visibility Tools — Most Only Show That It Did

Most AI visibility trackers tell you that your score moved. Far fewer tell you why. aeo-platform is a free, open-source CLI that names the exact question and engine behind every move in your AI visibility score, on every run. It also refuses to call a change a change until the number beats its own historical noise.

Every AI visibility tool sells the same first screen: a score out of 100, a coloured arrow, and a line that moves. The score is the easy part. The moment someone asks what to do about the arrow, the score stops being useful — an arrow is a symptom, and nobody can act on a symptom.

In August 2026 our team ran a dated check across 18 AI-visibility tools — Otterly, Profound, Peec.ai, Bluefish, AthenaHQ, Goodie, HubSpot's AEO Grader, Evertune, Ahrefs Brand Radar, Semrush's AI toolkit, Discovered Labs, geo-aeo-tracker, Trendos, Keytomic, search-visibility.ai, Minuttia, AEO Engine and Geoptie — reading what each one publishes about its own tracking. The result is not the one we expected to write up.

One thing about the shape of that list before the finding, because it is a fact about the category rather than a technicality. Several of the eighteen are not standalone software vendors. Discovered Labs is an agency that also ships its own AI visibility tracker; Minuttia is a content and SEO agency whose free AEO Report runs your brand through a set of prompts across AI platforms and returns visibility, citations and sentiment organised by topic. Both are products you can actually run, and the line between "tool" and "service" in this market is blurrier than the category's own marketing suggests.

We expected to find that a dated, multi-metric history was the rare thing. It is not. Seventeen of the eighteen track several metrics over time, and at least three of them — geo-aeo-tracker, Trendos and search-visibility.ai — give you a per-brand dated history at no cost. What almost nobody publishes, free or paid, is which specific question, on which specific engine, moved the number.

What Is Driver Attribution in AI Visibility Tracking?

Driver attribution names the specific query, engine and answer that moved your visibility score, instead of only reporting that it moved. A tracker without it gives you a direction; a tracker with it gives you a task. It is the difference between knowing you dropped eight points and knowing which question to go and answer.

Three separate capabilities get bundled together under the single word "tracking", and they are worth pulling apart — because a tool can do the first two well and still leave you with nothing to do on Monday morning.

  1. A time series. The same metric, recorded on dated runs, so today's number sits next to the number from three weeks ago instead of overwriting it. Without this, every measurement is a first measurement.
  2. Two deltas, not one. Movement against the previous run tells you what changed since you last looked. Movement against the very first run ever recorded tells you whether the whole programme is working. A tool that shows only the first can report a bad week on a brand that has tripled since day one.
  3. Attribution. The decomposition that connects a delta to the cells that produced it — which query, which engine, which answer, whether a rival moved into the space you vacated, and how much of the total movement that single cell can actually account for.

aeo-platform — the open-source npm CLI our team builds and uses on client work — carries all three. Version 1.8.0 records six metrics per run, each with a delta against the previous run and a delta against the first run on record, and a separate comparison engine explains the movement rather than merely colouring it.

We Checked 18 AI Visibility Tools. Here Is What Is Actually Rare.

Multi-metric tracking over dated runs is no longer a differentiator: 17 of 18 tools we checked in August 2026 do it, and at least three are free. Causal attribution is the rare capability. Among the free tools in that set, the honest ceiling is a named citation gap; the rest stop at "it changed."

The table below records what we found on 26 August 2026, reading each vendor's own public pages. It is a dated observation of published capability, not a laboratory benchmark. Vendors ship fast, and a row that is accurate today may be wrong next quarter. Where the detail sits behind a paid login we did not open, the table says so rather than guessing — an unopened product is not the same thing as an absent feature.

ToolFree tierDated multi-metric historyNames the query and engine behind a move
aeo-platform (ours, open source)Yes — MIT, runs on your own machineYes — six metrics, delta vs previous run and vs first runYes, on every run, per query and engine
ProfoundNo free tier (a trial is offered on the Growth plan)YesMarkets "Answer Engine Insights"; the diagnostic depth sits behind a paid login we did not open
geo-aeo-tracker (open source)YesYesPartly — a "Citation Opportunities" view names a competitor or URL behind a gap, not a per-query decomposition
AthenaHQA $25 / 300-credit grant on the Essential plan, not an unmetered free tierMarketed with month-over-month framing, but no dated-history or trend feature is itemised on any tier — the paid tiers' "competitor tracking" bullet is benchmarking, not a time seriesNot published
TrendosYes — "Free access. No hidden paywall"Yes — past visibility is viewable before you even log inNot published
GoodieFree tools, yes; per-brand tracking sits on the paid Explorer planPartial — the free AI Visibility Index is an aggregate industry benchmark on 30-day, 90-day and 6-month timeframes, updated monthly, not a confirmed per-brand run-over-run historyNot published
KeytomicTrial only — a "Free for 48 Hours" scan and a "3-Day Trial", not a standing free tierAdvertises "week-over-week trend data"; whether that sits in the trial or behind a paid plan is not statedNot published
search-visibility.aiYesYesNo — its own documentation asks the user to note manually which content shipped
Minuttia (AEO Report)Yes — a free baseline reportMulti-metric (Visibility Score, Share of Voice, Average Position); highlights changes between the latest and previous prompt runsNot published; citations currently tracked for Perplexity only
Otterly, Peec.ai, Bluefish, EvertunePaidYesNot published
Ahrefs Brand Radar, Semrush AI toolkitPaid tier of a larger suiteYesNot published

The table shows the rows most readers will recognise; Discovered Labs, AEO Engine, Geoptie and HubSpot's grader were checked on the same day and none of them changed the picture. "Not published" is also doing deliberate work here. It means we could not find the vendor claiming the capability on a page we could read — not that the product cannot do it. A dated survey of published claims is a weaker instrument than a hands-on comparison, and it is the strongest instrument available without buying eleven subscriptions. Read the rows as a map of what the category currently thinks is worth advertising.

The word "free" is the cell that moved most under checking, and it is worth slowing down on because it is the one buyers screen on. Three of the rows we had initially written down as free did not survive a second look. Goodie's free AI Visibility Index is an aggregate industry benchmark on 30-day, 90-day and 6-month timeframes, refreshed monthly — genuinely useful, but it ranks brands within a category rather than keeping your own dated record; continuous per-brand monitoring sits on the paid Explorer plan. Keytomic advertises a "Free for 48 Hours" scan and a separate three-day trial, which is a trial rather than a standing free tier, and its "week-over-week trend data" is listed among the included features without saying which tier it lands in. AthenaHQ gives new accounts a 25-dollar, 300-credit grant on its Essential plan, which is real usage at no cost — but reading the published per-tier feature lists, no dated-history or trend capability is itemised on any tier, including that one — the paid tiers do list a "competitor tracking" bullet, but that names competitive benchmarking, not a time series. Its month-over-month language lives in case studies rather than in a feature list. None of this is a criticism of any of the three products. It is a warning about how the category uses the word, and about how easily a reader lands on the wrong answer by reading marketing copy instead of the tier a feature actually ships in.

The three we could confirm as free and per-brand were geo-aeo-tracker, Trendos and search-visibility.ai. Trendos was the most unambiguous of them: "Free access. No hidden paywall", with past visibility viewable before you even log in. If you are choosing a tool on price, check whether the free tier tracks you across dated runs, or shows you a category league table on a date filter — those are different products wearing the same word.

Two rows deserve their names spelled out. Profound is the most explicitly diagnostic product in the set and has no free tier: its published plans are 99 US dollars a month for Starter and 399 for Growth on annual billing, with Enterprise on request and a trial on the Growth plan (checked 26 August 2026). Its headline framing for the relevant module is "Answer Engine Insights — see how AI represents your brand in every conversation." If your budget starts at four figures a year, that is a serious tool and we will not pretend otherwise.

geo-aeo-tracker is the closest free relative we found: an open-source, local-first AI visibility dashboard that tracks a brand across six AI models, at roughly 244 GitHub stars when we looked. Its "Citation Opportunities" view names a specific competitor or URL sitting in a gap — genuine partial attribution, and more than most paid dashboards publish. What it does not do is decompose an index movement cell by cell and price each contribution.

The one tool in our set with no dated history at all was a one-shot grader: it scores a domain on request and keeps no record between checks. We wrote about that pattern in some detail when we first measured ourselves — see our Week 1 baseline, where two free graders returned two incompatible answers about the same domain.

What Does a Real "Why It Moved" Look Like?

It looks like one sentence naming the question and the engine, plus an arithmetic account of what the loss cost. On our own TypelessForm record, run 8 did not open with a red arrow. It opened with the line: "Gemini stopped naming you on one question."

aeo-platform report, run 8 of the typelessform.com record: the finding 'Gemini stopped naming you on one question', an eight-run pass-or-fail track dated 2026-04-23 to 2026-08-13 with seven passes and one failure, a breakdown attributing 2.9 index points to the lost answer, a note that no rival was named instead, and an eight-run sparkline rising from 33 to 92 with one run flagged partial.
Run 8 of a real typelessform.com record. The report names the question, shows the seven runs on which the answer was held before it was lost, attributes 2.9 of the 8 lost index points to that single cell, and states plainly that the remainder is not attributable from what this run stored.

The finding is specific enough to act on without opening anything else. Asked "multilingual voice form filling for international websites" on 13 August 2026, Gemini named other tools and did not name TypelessForm. The pass-or-fail track underneath shows the same question answered on seven earlier runs — 23 April, 13 May, 18 May, 25 May, 10 June, 11 June, 11 July — all held. This was the first loss on that cell in 112 days.

Then the arithmetic. One answer out of twelve is 8.3 percentage points of the Presence axis. Presence carries 35 per cent of the index weight, so the lost answer accounts for 2.9 index points. The index as a whole moved eight points down. The report prints both numbers and does not quietly close the gap between them.

The rest is not attributable from this run's record: the run stores one index number, not the four axis values behind it.

That sentence is the one most tools in the category would never ship. It concedes, on the client-facing page, that roughly five of the eight lost points cannot be explained from the data this particular run happened to save. A vendor optimising for a confident dashboard would distribute the remainder across the axes and print a tidy total. The cost of doing that is not aesthetic: whoever reads the report then spends money against a decomposition that was invented rather than measured.

The second panel answers the question that decides what you do next. Who was named instead? In this case, nobody: the engine answered without recommending any tool. That makes it a content gap rather than a competitor problem, and the two call for opposite responses. A competitor problem is a positioning and third-party-mention job. A content gap means the question has no good answer on the open web yet, and the fastest fix is to publish one.

Two disclosures sit in the same view, in smaller type, and both cost the report a cleaner story. The tracked question was reworded during the record, so runs before 10 June asked a different phrasing of the same slot. And run 5 measured only nine of twelve answers — it is marked partial on the sparkline and excluded from comparison rather than being silently averaged into the trend. (The presentation of all this was refreshed in version 1.8.0; the capability itself has been in the record since April.)

Why a Metric Can Report a Win as a Decline

Because two of the four index axes are averages taken only over the answers where your brand actually appears. Win a new mention that is factual rather than glowing, and the average falls — the brand got more visible and the score says it got worse. Attribution is what separates that case from a genuine decline.

The distinction is structural, not a quirk. Presence and Citation are counted across every cell of the run: every query-times-engine combination either names you or does not. Sentiment and Rank are averages computed only over the cells where you appear at all. That makes the second pair conditional — when the population they average over changes, they move for reasons that have nothing to do with your visibility.

Our decomposition caught exactly this on one run pair. Sentiment fell from 100 to 83 and read as the single largest contributor to a 14-point index drop. Split into like-for-like movement and compositional movement, the change across the answers present in both runs was exactly zero. The entire move came from one newly gained mention that happened to be factual rather than glowing. The brand had broken into a third engine — an unambiguous win — and the undecomposed number scored it as its worst decline of the record.

Acting on that number would have meant commissioning reputation work against a tone problem that did not exist, while the actual event — a new engine starting to name the brand — went unnoticed and unreinforced. That is not a rounding error in a report. It is a budget pointed at the wrong quarter.

So every conditional component is now split in two before it is shown: like-for-like, the movement among cells present in both runs, which is a real signal; and compositional, the movement caused purely by which cells entered or left, which is an artefact. Gained cells that scored below the previous average are flagged explicitly, so a gain is never reported as a decline. Where nothing at all was comparable across both runs, the report says that too — a component with no like-for-like population is a finding, not a blank.

When Is a Number Big Enough to Be Called a Change?

A delta passes two tests before the report calls it a move: a fixed floor for its unit, and, from five runs on, its own historical noise. The floor is 3.0 points on a 0-100 axis, one whole unit on a count, and half a position on a rank. Below the floor the number still prints — uncoloured, with no claim attached to it.

A floor keyed to the unit matters because "3" means nothing until you know three of what. Three points on a 0-100 visibility index is a wobble. Three answers gained on a twelve-cell grid is a quarter of the whole surface. One shared threshold across both would either paint every fluctuation red or refuse to colour a real collapse.

The second test separates two questions that most dashboards collapse into one. Clearing the floor earns a coloured chip — the reader can see that a number moved. Being named as the thing to act on is a higher bar, because that line is where money gets spent. On a continuous axis, a delta must also beat that metric's own median run-to-run step across its recorded history before the report will name it.

That noise test only comes online from the fifth measured run. Below five points there is not enough history to call anything noise, and a test that failed silently would suppress every genuine finding on a young account. It also counts measured points rather than array slots: a metric reported on only two of eight runs has exactly one step on record, and that one step is the delta being judged — it could never beat its own median, so an unguarded test would veto every finding on a newly added check.

Count metrics are deliberately exempt from the noise test, and the reason is worth stating because it looks like a loophole. On a count, the floor is already one whole unit — the smallest change the metric can express. A real history from our own record makes the point: answers naming the brand across eight runs ran 4, 5, 5, 7, 9, 10, 12, 11, giving a median step of exactly one. Apply a beat-your-own-median rule there and losing an answer would be permanently unreportable.

The commercial cost of all this is quieter reports. Fewer arrows, fewer movements to celebrate, and several screens per run where the honest output is a number with no story attached. For a consultancy that sells on visibility work, that is an awkward trade — and it is the right one, because a report that names a mover every single week trains its reader to ignore it.

What the Report Refuses to Draw Before It Has Enough History

Every visual in an aeo-platform report implies a claim, so the report unlocks them in stages as runs accumulate. One run gets a baseline and no arrows. Two runs get a delta chip and nothing else. From three runs on, the sparkline, the "since day 1" line and any directional word are allowed.

aeo-platform overview for run 1 on 2026-04-23, headed '33 of 100 on the first run. This is the baseline.' Four KPI cards show visibility index 33 of 100, four of twelve answers naming or citing the brand, zero of four engines naming it everywhere, and zero lift opportunities — with no delta chips anywhere on the screen.
Run 1, 23 April 2026. One run on record, so there is nothing to compare against — and the report says exactly that instead of inventing a direction from a single point.

The first run is labelled a baseline in the headline itself. Nothing on the page carries an arrow, because a first measurement has no delta to report and a zero chip would read as "no change" rather than "no comparison". The phrasing on that screen — "This is the baseline" — is the only claim a single run supports.

aeo-platform overview for run 2 on 2026-05-13, headed 'Up 9 points overall — but Perplexity dropped an answer it had held before.' Visibility index 42 of 100 with a green nine-point chip, five of twelve answers naming or citing the brand with a one-answer chip, one of four engines naming it everywhere.
Run 2, 13 May 2026. A delta chip appears — but still no sparkline and no since-day-one line, because two points cannot describe a shape and the two deltas would say the same thing twice.

Run 2 is also where attribution starts earning its place, and the headline shows why the two capabilities belong together. The index rose nine points, which on its own reads as an unambiguously good week. The same headline adds that Perplexity dropped an answer it had held before. Both are true, and only the second one is actionable.

aeo-platform report at three runs: a 'Score over time' panel reading '3 runs, 25 days', 33 to 42, 'Up 9 points since day 1', with the first sparkline; below it a four-axis panel where Presence (35 per cent weight) and Citation (20 per cent) carry deltas while Sentiment on 4 of 12 answers and Rank on 1 of 12 are both marked 'too few to score'.
Run 3, 18 May 2026. The first sparkline is drawn — and on the same screen two of the four axes decline to produce a delta at all, printing their coverage instead of a number.

The third run unlocks the shape, the since-day-one line and any directional word. It also shows the honesty gating doing something visible. A conditional axis must be reported on at least 60 per cent of the run's answers before its delta may print at all; if coverage between the two runs moved by more than a quarter of the cells, the chip reads "coverage moved" instead of a number. Sentiment was reported on four answers of twelve here, Rank on one of twelve. Both are far under the floor, so both say "too few to score".

This is the part of the design that not one of the eighteen tools in our set advertises. Refusing to draw a line, printing a coverage fraction where a competitor prints a confident percentage, flagging one run as partial rather than folding it into the trend — none of it demos well. It matters the first time someone asks you to defend a number in front of the person who paid for it.

Which Six Numbers Carry a History?

Six metrics are recorded per run, each carrying a delta against the previous run and a delta against the first run ever recorded. They are the visibility index, answers naming or citing you, rival tools named, citations to your own pages, distinct hosts cited, and answer-capsule coverage. One of the six is inverted: more rivals named is worse.

  • Visibility index — the composite score on a 0-100 axis, weighted across Presence, Sentiment, Rank and Citation. Movement below 3.0 points is not coloured.
  • Answers naming or citing you — a count of the cells where the brand appears in the prose or in the cited sources. This is the same population the index's Presence axis counts, which is why the two never disagree about a lost answer.
  • Rival tools named — a count of competing products the engines recommended instead. The one metric where a rising number is a worse result, coloured accordingly.
  • Citations to your own pages — how many times an engine cited your own domain as a source, rather than merely mentioning your name. Naming and citing are different wins and are counted separately.
  • Distinct hosts cited — how many different domains the engines drew on across the run. A widening set tells you the answer space is getting more crowded; a narrowing one tells you a few sources are consolidating.
  • Answer-capsule coverage — the share of tracked questions for which a clean, extractable answer capsule exists, on a 0-100 axis.

Carrying both deltas on every one of them is what makes a bad week legible. A drop of eight points against the previous run and a rise of 59 points since the first run are two facts about the same brand on the same day, and a tool that reports only one of them is telling half a story in whichever direction suits it.

How Is This Different From Asking Whether a Single Score Is Trustworthy?

A single AI visibility score asks whether today's number is honest; a history asks whether the difference between two numbers is honest. The second is the harder question, because a difference can be arithmetically correct and still be meaningless. Both need answering, and they need answering separately.

We took the first question apart in Can You Trust an AI Visibility Score?, which covers how a single measurement is made auditable: raw engine answers saved to disk, competitor mentions verified with two models rather than one, a frozen question set, and an explicit list of what the run did not measure. This article is about what happens once you have two of those measurements and have to explain the space between them.

The measurement-to-plan loop that turns any of this into work is a third question again, covered in aeo-platform: measure, plan, improve. Attribution is the hinge between the two: a measurement tells you where you stand, attribution tells you which cell moved, and only then does a plan have something specific to attack.

How Do You Get a Dated History of Your Own?

Run aeo-platform with npx, point it at your brand and domain, and the record starts on the first run. It is MIT-licensed, has zero dependencies, needs Node 20 or newer, and keeps every raw engine answer on your own disk. The history is a folder of dated runs, so the record belongs to you and survives any decision you later make about the tool.

npx aeo-platform@latest init --yes --brand=YOURBRAND --domain=YOURDOMAIN.COM --auto \
  && npx aeo-platform@latest run \
  && npx aeo-platform@latest report

Three commands, three jobs. init picks your tracking queries and writes them into a config file, which is what freezes the question set so later runs stay comparable. run asks each engine your queries and records the answers — exactly one API call per cell, where a cell is one combination of query, engine, region and pass. report renders the HTML with whatever the run count entitles it to draw. For a weekly rhythm, a global install with npm install -g aeo-platform is less typing.

The first run produces a baseline and nothing else, and that is the correct output. The second gives you a delta. The third unlocks the shape and the since-day-one line. The fifth brings the noise test online, and from there the report starts declining to name movers it would happily have named a month earlier.

You pay your own model providers directly with your own API keys. Nothing routes through Webappski, there is no account to create, and the keys are read from environment variables and never written to disk. Raw responses land in a dated folder under aeo-responses/ on your machine, which is also what makes an old run auditable months later — the numbers can always be rebuilt from the answers that produced them.

Frequently Asked Questions

Why did my AI visibility score drop?

Without attribution you cannot know, and most trackers will not tell you. A drop has at least four distinct causes that look identical on a chart: an engine stopped naming you on one specific question, a rival took the slot, the question set or its wording changed, or a conditional average moved because the set of answers it averages over changed. A report that decomposes the delta per query and engine separates those four; a coloured arrow does not.

Is a free AI visibility tracker with history actually rare in 2026?

No. In our dated check of 18 tools on 26 August 2026, multi-metric tracking across dated runs was near-universal — 17 of the 18 — and at least three tools gave it away free: geo-aeo-tracker, Trendos and search-visibility.ai. Check the fine print, though — some "free" tiers turn out to be a time-limited trial, or an aggregate industry index rather than a record of your own brand. What stayed rare was causal attribution: an explanation of which question and which engine moved the number. That is the capability worth shopping for, not the history itself.

How many runs do I need before the numbers mean anything?

Three before any directional language or shape is drawn, and five before the noise test comes online. One run is a baseline and the report labels it as such. Two give you a single delta with no shape. From three, the sparkline and the since-day-one line appear. From five, a delta on a continuous axis must beat that metric's own median run-to-run step before it is named as the thing to act on.

What stops a tracker from calling random noise a trend?

A fixed floor per unit, plus a per-metric noise test. In aeo-platform the floor is 3.0 points on a 0-100 axis, one whole unit on a count and half a position on a rank; below it the number prints uncoloured with no claim attached. Above it, on continuous axes only, the delta must also exceed the median of that metric's own run-to-run steps before the report names it.

Can a rise in my score be bad news, or a fall be good news?

Yes, and it is the most common misreading in the category. Sentiment and rank are averaged only over the answers where your brand appears, so gaining a new, neutrally worded mention can pull the average down while your real visibility went up. The fix is decomposition: split every conditional axis into like-for-like movement among answers present in both runs, and compositional movement caused purely by which answers entered or left.

Does aeo-platform send my data anywhere?

No. It is local-first: raw responses stay in a dated folder on your own disk, API keys are read from environment variables and never written to disk, and there is no account and no telemetry. The only non-provider network call is a once-a-day version check against the npm registry, which is skipped in CI and can be switched off with an environment variable.

Start a Record You Can Argue With

The useful move this week is to take a first measurement, whatever tool you take it with. A baseline you did not record three months ago is the one number no product can reconstruct for you, and every attribution capability described above is worthless until there are two dated runs to compare.

If you would rather have the record than run it, our team runs the same measurement for clients as part of our AEO services — same engine, same gates, and the dated runs are handed over rather than locked in a dashboard. If you just want to see where you stand first, a free AI visibility audit is the shorter path. Either way, ask whoever measures you the same question we now ask of every tool in the category: when the number moves, will you be told which question moved it?

← Back to all posts