How Do You Know a Change in Your AI Visibility Score Is Real? Our Own Report Called a 42-Point Fall It Could Not Support
On 1 September 2026 the tracker we build reported a 42-point fall for our own product and listed the rivals apparently standing in its place. None of them had displaced anything, and about 17 of the 42 points are demonstrably the instrument — the rest is unattributed. This is the reading discipline that should have caught it, the guards that exist, and the axis that had none.

Webappski's own AI-visibility tracker reported a 42-point fall for our product on 1 September 2026. The tracker had switched to a different model behind the same engine name, then listed the rivals standing where we no longer did — none of whom had displaced anything. About 17 of the 42 points are demonstrably that switch; the rest is not established, and saying so rather than rounding it up is the discipline this article is about. We build the tracker and sell the service, so this failure is ours to publish.
Every number below comes from one measurement record: TypelessForm, a Webappski product, put to three fixed buyer questions across four answer engines, with each engine's untouched response saved to disk. We funded the measurement, we wrote the tool that produced it, and we sell AI-visibility work to other companies. That is a conflict of interest in every direction, and it is the reason this article takes the shape it does — the only version worth publishing is the one where the tool is the thing being caught.
This is not an incident report. It is the discipline the incident should have been caught by: what a visibility number is a measurement of, what separates an artifact from a change, which guards a report needs and what each one is for. The interesting part is not that a guard failed. Every guard held. The number cleared all of them, and was still wrong, because none of them was watching the axis that moved.
What is an AI visibility score actually a measurement of?
An AI visibility score measures one instrument's answers to one fixed set of questions on one date. It is not a measurement of your standing in AI search.
Five terms carry the whole argument below. Collapsing any two of them is the ordinary way a number stops meaning anything, and it is what happened to us.
A question basket is the fixed set of prompts a run puts to every engine. Ours holds three: best voice form filling tools 2026, top one-shot voice form filling services for e-commerce, and multilingual voice form filling for international websites.
An engine is the branded product that answers — ChatGPT, Gemini, Claude, Perplexity. It is a name on a socket.
A model is the specific system that answers on that engine's behalf on a given day. It is what is plugged into the socket, and it is replaced on the vendor's schedule, not yours. Our earlier piece on verifying a free visibility score already made this argument in the abstract. What follows is what it costs when you forget it inside your own pipeline.
A cell is one engine's answer to one question in one run. Our run has twelve of them: three questions across four engines. A score of 83% means ten of those twelve cells named the brand.
A like-for-like comparison is two runs that share a basket, an engine set, and a model per engine. Break any of the three and the difference between the two numbers is partly a difference between two measuring devices, in an unknown proportion.
What did our report get wrong on 1 September?
It reported a fall that had not happened, and set beside it a list of rival brands that were not competing for the question asked.
The dated arc, same basket and same four engines throughout: 13 August scored 92% (11 of 12 cells), 1 September scored 50% (6 of 12), 2 September scored 83% (10 of 12). Between the first two runs the tracker's model discovery quietly preferred a small model over the newest generation, and the OpenAI leg ran on gpt-5.4-mini instead of the search-grounded model that had produced the August figure.
Take the third question first, because it is the cleanest. Asked for multilingual voice form filling for international websites, the 1 September answer was an engineering how-to: set the page lang attribute, map voice input to labels rather than placeholders, localise date and address formats. Its citations were developer.mozilla.org and w3.org. It named no vendor at all.
2026-09-01 · openai · gpt-5.4-mini · q3
"If you're building multilingual voice form filling for international
websites, the main idea is:
1. Detect or let the user choose language early
- Set the page language correctly with `lang`. ..."
cited hosts: developer.mozilla.org, w3.orgA vendor list with no vendors in it cannot contain our brand — and cannot contain anyone else's either. The zero recorded in that cell was not a measurement of our absence. It was a measurement of a question that had been answered as something else. No amount of care downstream can rescue a cell like that, because the thing the number describes was never asked.
The e-commerce question went a different way. It was answered with browser extensions, and the two names the report carried forward, Talk2Forms and Speak2Fill.ai, were cited from chromewebstore.google.com listings — personal autofill add-ons a shopper installs, not services an online store buys. Six brand names appear in the report's own account of where we lost ground — four from this e-commerce answer, two from the Gemini answer further down. None of them had taken a position from us. They were answering a neighbouring question that happened to share our vocabulary.
The first question returned consumer dictation software — Wispr Flow, Superwhisper, Dragon, the built-in Windows and Apple dictation tools — cited from Zapier and G2 round-ups. A real answer, and not to the question our basket had asked. Three questions, three different ways of drifting off target, one cause.
That accounts for three of the five cells the run lost, and an honest reckoning has to say what happened to the other two. One was Gemini, and the control day speaks to it more than we first credited: gemini-3.5-flash and gemini-3.6-flash both answered the same three questions on 13 August, and both named the brand in two of the three. Across that pair the instrument made no difference at all. But the model that answered on 1 September was a third one, gemini-3.7-flash, which the control day never measured — so the pair cannot be stretched to cover it. That cell named us again the following day on an unchanged model, which is what ordinary variance looks like rather than displacement, and it stays unattributed.
The other was Perplexity, and it forces a disclosure this article would be hypocritical to leave out. Half of every run quoted here is read by hand. Six of the twelve cells — all three Claude answers and all three Perplexity answers — carry "model": "manual" and "source": "manual-paste" on 13 August, on 1 September and on 2 September alike. We do not put Claude or Perplexity through their APIs for measurement: each question goes to a fresh, context-free AI session we run ourselves, and the answer is pasted back into the run. That is our standing method and a deliberate choice, not an accident of this run — but it means those six cells record no model identity, so the instrument axis cannot be assessed on any of them in either direction, including the Perplexity cell this run lost. An article telling you to demand the instrument beside the number owes you that about its own numbers first.
Even inside the three ChatGPT cells the control day does not say one thing. On two of them the newer model had already failed to name us on 13 August, so their loss on 1 September is the switch. On the third it named us eight times where the older model managed thirty-two: weakened fourfold, not closed, and closed two and a half weeks later by something the control day cannot identify. Two cells of twelve is about 17 of the 42 points. The remaining 25 are unattributed — not secretly ours, not secretly a rival's, simply not established by this record. A report that rounded them up into the tidy version would be committing the error this whole article is about, one level higher up.
How do you separate the instrument from the brand? Run the control.
You hold everything constant except the instrument, and see whether the number still moves.
We had run exactly that control on 13 August, weeks before it mattered: one day, one brand, the same three questions, two OpenAI models measured side by side. Counting occurrences of the brand name in the answer text only, citation URLs excluded:
- gpt-5-search-api — named the brand in 3 of 3 answers: 32, 7 and 12 mentions.
- gpt-5.4-mini — named the brand in 1 of 3 answers: 8, 0 and 0 mentions.
Nothing about the company changed between those six answers. They were collected on the same day, from the same provider, against the same prompts, by the same pipeline. The only difference was which model answered, and it moved the result from 3 of 3 to 1 of 3.
That control is why the September diagnosis is a finding rather than a guess. The OpenAI leg reads 3 of 3 on 13 August with the search-grounded model, 0 of 3 on 1 September with the mini, and 3 of 3 again on 2 September once the model axis was repaired and the run went out on gpt-5.6-luna. The brand did not oscillate. The instrument did.
One methodological note, because it decides whether those counts mean anything. Every cell above is counted the same way: the model's answer text, case-insensitive, citation URLs and metadata excluded. Counting the whole response object instead — URLs included — turns one of our Gemini cells from 5 mentions into 26. Both counts are reproducible; neither is wrong on its own; a table that mixes the two in adjacent rows refutes itself. If a report shows you mention counts, ask which of the two it counted, and whether it counted every row that way.
Was all of the movement the instrument? No, and that matters more than the part that was.
No. On the same two dates, Gemini ran the identical model and its result still moved.
gemini-3.7-flash answered on 1 September and again on 2 September. On the first date it named the brand in one of three answers; on the second, in two of three. Nothing was swapped, nothing was repaired, nothing about the site changed overnight. The answers simply came out differently.
That one cell is the discipline in miniature. The tidy explanation — the model was swapped, therefore all of the movement was noise — would have been wrong in precisely the way the original report was wrong: a confident story running further than its evidence. Part of the movement was the instrument. Part of it was ordinary run-to-run variance. A report that cannot separate the two has not finished its job, and an author who does not say so is selling the same false certainty in the opposite direction.
It is also why the guards further down exist at all. If answers were deterministic there would be no need for a significance floor or a noise test — you could read every delta literally, and a difference of one would be a difference of one.
What separates run-to-run variance from a signal?
Persistence. A cell that goes and comes back inside one run is variance; a cell that stays shut across three consecutive runs is a signal.
Variance is symmetric and short-lived: it moves in both directions and does not survive re-measurement. A signal is directional and repeats. This is unglamorous, and it is the only discriminator that costs nothing but patience — which is exactly why it is the one most often skipped when a number moves in a direction someone has to explain on a call.
Four questions, in this order, before you accept any delta as real:
- Did the instrument change? Compare the exact model string per engine between the two runs, not the engine name. If either run cannot tell you what answered, the comparison is not available and no further question is worth asking.
- Did the population change? A metric averaged only over the answers that carry it moves when that set of answers changes size, for reasons that have nothing to do with your brand.
- Is the move larger than this metric's usual step? Every metric has a characteristic wobble. A delta inside it is a description of the wobble.
- Has it persisted? One run is a hypothesis. Three consecutive runs in the same direction is a finding you can spend money on.
Our own record carries both cases side by side, on the same engine, which is the clearest illustration we can offer. Gemini has not named the brand on the multilingual question in any run since 13 August: shut on 13 August, shut on 1 September, shut on 2 September, across two different Gemini models. That is a signal, and it is the one cell in this whole episode actually worth acting on.
On the first question the same engine named us on 13 August, went silent on 1 September, and named us again on 2 September — with the last two runs on an identical model. That is variance. Treating it as a loss would have bought a month of work on a page that was never the problem, and the report would have looked just as confident either way.
Which guards does a visibility report need, and what is each one for?
Each guard answers one specific way of fooling yourself, and is worth naming by that failure rather than by its threshold. What the set has in common turns out to matter more than any single one of them.
The ones below run in aeo-platform 1.11.0, the open-source tracker that produced every number in this article, and they live in one readable file, lib/report/trend-model.js. The point is not that our thresholds are the correct thresholds. The point is that a report should be able to tell you what its thresholds are, and where they are written down.
// lib/report/trend-model.js — aeo-platform 1.11.0
export const FLOOR = Object.freeze({ points: 3.0, count: 1, rank: 0.5 });
export const NOISE_TEST_MIN_RUNS = 5;
export const TREND_LANGUAGE_MIN_RUNS = 3;
export const MIN_COVERAGE_RATIO = 0.6;
export const MAX_COVERAGE_SHIFT = 0.25;
export const NOISE_TESTED_UNITS = new Set(['points']);- A significance floor. Answers the habit of reading a rounding wobble as a result. Movement below 3.0 points on a 0-100 axis is not coloured and not called a change. The floor is keyed by unit, because 3 means nothing until you know 3 of what: 3.0 points on a percentage axis, 1 whole unit on a count, half a position on a rank.
- A noise test against the metric's own history. Answers the habit of treating a typical step as an event. A delta may be named as the mover only if it exceeds the median of that metric's own historical steps. It comes online at the fifth measured run, and before then it deliberately fails open — it returns true rather than false, because a young account with almost no history would otherwise have every genuine finding silently suppressed. A guard that quietly vetoes everything is worse than no guard.
- A restriction on which metrics the noise test may veto. Answers the habit of applying a continuous-axis test to a thing that is counted. The test runs on 0-100 axes only. On a count the floor is already one whole unit — the smallest change the metric can express — so a must-beat-the-typical-step rule would permanently veto the smallest meaningful movement. The code carries the real case: a brand whose answers-naming-it ran 4, 5, 5, 7, 9, 10, 12, 11 has a median step of exactly 1, so losing an answer could never have been reported.
- Trend words withheld below three runs. Answers the habit of describing a direction from two points. Falling and climbing are not printed until three runs are on record. A line through two points is a delta, not a trend.
- A refusal to state a delta when the answer population moved. Answers the habit of mistaking a change in who could be measured for a change in what was measured. Metrics averaged only over the answers that carry them — rank, sentiment — must be reported on at least 60% of the run's answers, and that coverage must not have shifted by more than 25 percentage points between the two runs. Below either bar the report prints the coverage instead of a delta.
Read the list as a whole and its shape is obvious: four of the five guard the size of a number, and the fifth guards the population the number was computed over. Not one of them asks what produced the answers.
So why did none of them catch a 42-point artifact?
Because a guard list is a record of the ways its builders already knew how to be fooled, and nobody had yet been fooled this way.
The floor exists because someone once read a wobble as a result. The coverage gate exists because someone once compared two averages taken over different populations. Each threshold is a scar. The failure that gets you is always the axis with no scar on it yet — and it will pass every existing check on the way in, because the checks were written by the same people who did not see it coming.
Trace the 1 September number through the list and every gate opens — not by accident, but exactly as designed. A 42-point delta clears the 3.0-point floor by an order of magnitude. It beats any plausible median historical step, so the noise test waves it through as decisively as it would have blocked a delta of two. The run count was long past three, so directional words were permitted. Presence and citation are measured across every cell rather than averaged over the subset that carries them, so the coverage gate never looks at them at all — it short-circuits for those direct axes and gates only rank and sentiment.
Every guard was satisfied, because every guard was answering a version of the same question: is this move too small to mean anything? A 42-point move is not too small. It was merely wrong. Not one of the five asks the other question — whether the two numbers being subtracted were produced by the same device.
Notice what that says about care. The tool was not careless. It carried a coverage gate, a noise test keyed to each metric's own history, and a run-count ladder governing which words the report was allowed to use — an unusual amount of restraint machinery for a category where most vendors print an integer and move on. All of it was pointed at the question axis. A system can be scrupulous along every axis it thought of and still be silently wrong along the one it did not, and the scrupulousness is what makes the silence convincing.
Our own commit message for the repair states the defect more bluntly than we would have in prose:
The tracker guarded the QUESTION axis (a shifted basket refuses to state a delta) and was blind to the INSTRUMENT.
This is an amendment to something we published ourselves. In our article on driver attribution we argued that the useful thing about a tracker is not that it tells you the number moved, but that it tells you which question and which engine moved it. We still believe that, and this run is the correction to it: driver attribution is only as sound as the axes it holds constant. A report able to name the engine, the question and the direction can still name a driver that does not exist, if it is not also checking what answered. Naming a why is a stronger claim than naming a what, and it fails harder when it is wrong, because a plausible cause stops the reader from looking for the real one.
What actually catches this class of error?
Two things, and neither of them is a threshold: naming the instrument beside the number, and requiring a change to persist.
The first is attribution on the model axis, added to aeo-platform in version 1.11.0. Where a run is compared to a previous one, the report now states the was-and-is model pair for every engine whose model changed, marks the affected rows as not like-for-like, and — the part that does the real work — disqualifies the calm wording. A report is no longer permitted to say that visibility held steady when the ruler moved underneath it. The shipped text says so directly: holding your ground while the engine underneath moved is not measured against the same instrument as last time. Where two runs share no model at all on an engine, it states in plain words that not one answer was measured the same way twice.
The second is persistence, and it needs no code whatsoever. It is the only test in this article that works on a report you did not build, produced by a vendor who will not tell you which model answered. Re-run, and look at whether the cell is still shut. Variance will not survive the third run; a real loss will.
Both are cheaper than any of the five guards above, and neither is statistical. That is the uncomfortable half of the lesson. The machinery we had built was the sophisticated kind. The thing that would have saved us was the boring kind: write down what answered, and wait for it to happen twice more.
A guard printed where nobody looks is not a guard
The first of those two mechanisms existed before the false conclusion was ever read. It printed on a surface nobody had open.
aeo-platform emits the same run in more than one form: a markdown report, an HTML report, and a machine payload for the customer portal. In version 1.11.0, published on 2 September 2026, the markdown report prints a Run Comparison section that marks any engine whose model changed since the previous run and says in plain words that no like-for-like comparison is available on it. In that same version, the HTML renderer computed that identical comparison model and read it nowhere.
The HTML is the artifact that gets opened and shown to a client. The 42-point conclusion was read off it, beside a movement figure with nothing next to it. We consider that the more serious of the two defects on display here. The model was right, the logic was right, the warning existed, and the reader still received a bare number. Correctness that does not reach the surface where the decision is taken has not been delivered. Version 1.12.0, published later the same day, carries the warning into the HTML as well, and moves the wording into a module both surfaces share so they cannot drift apart again; the same release also names an engine that was measured last time and not this time. We are leaving the sequence in rather than folding it into a single tidy repair. The cause was fixed first, then the surface nobody reads, and only last the surface the conclusion was actually read from — and the order of those three is the more useful thing to know about how a team works than the fact that all three eventually happened.
The same class of gap is worth checking in your own config files, not only in a vendor's renderer. Ours declares a basket version of 1 with a healthcare question in its recorded history, while the live basket's third question has been a multilingual one for months. The version was never raised when the question was replaced. A version number that is not incremented on edit is worse than no version at all, because it certifies a comparability that does not exist.
It generalises past our own tool. When you evaluate any visibility report, the question is not whether the vendor's methodology page mentions a caveat somewhere. It is whether the caveat is printed on the page where the number is.
What does measurement practice elsewhere do that AI-visibility reporting does not?
It requires a result to travel together with a statement of how uncertain it is.
NIST's public guidance on measurement uncertainty describes how each component of uncertainty, however evaluated, is represented by an estimated standard deviation, termed standard uncertainty. That guidance sits inside a framework — NIST Technical Note 1297, following the ISO Guide to the Expression of Uncertainty in Measurement — whose entire purpose is that a number and its uncertainty are reported as one object rather than two, so a result cannot be quoted stripped of the conditions that produced it.
AI-visibility reporting has no equivalent convention. Scores are published as bare integers: no stated instrument, no stated basket, no stated repeatability. Our 83% is a bare integer too. What we are arguing for is small and cheap by comparison: name the instrument, name the basket, name the date, and refuse to state a delta when any of the three changed between the runs being compared.
None of this is specific to one market or one language. A score computed in English for a US software company and a score computed in German for a Mittelstand supplier fail in exactly the same way when the model behind the engine name changes between runs, and the checks below are word-for-word the same in both cases. The instrument problem does not respect a border.
What should you demand from any AI visibility report?
Seven answers, every one of which a competent vendor can give you in a single email.
- Which model answered, per engine, per run — the exact model string, not the brand name. A vendor who cannot produce it cannot tell you whether last month's number and this month's were made by the same device.
- Whether the question basket changed, and when — with a version that is raised every time a question is edited, and a recorded history that matches the questions actually being asked.
- Which engines were measured this run and which were not, by name. A footnote reading partial run without naming the missing engine is not a disclosure; it is the shape of one.
- How many runs are on record. Below three, no directional word — falling, climbing, recovering — should appear anywhere in the document.
- The smallest movement the vendor is willing to call a change, and on which axis. If there is no such number, every delta printed is being presented as meaningful, including the ones that are not.
- Whether the raw answers were saved, and whether you may read them. A score you cannot audit down to the sentence that produced it is a claim, not a measurement. Ours are written to disk on every run, which is the only reason the counts in this article could be recomputed by hand — the design behind that is in our note on what makes a visibility number auditable.
- What the report refuses to say. A report with no refusals in it is not being careful. It is being confident, which is a different thing and is cheaper to produce.
How do you apply this to your own numbers this week?
Run the control before you run the analysis, because the control is the cheap part and the analysis is the part you will act on.
- Pin the model. Fix the exact model string per engine in your configuration, and confirm that the string requested is the string that answered. A pin that is silently overridden by an auto-discovery step is not a pin, and that is the precise defect that produced this article.
- Run one control pair. Same day, same brand, same basket, two different models on a single engine. It costs one extra run and it tells you how much of your score belongs to the instrument rather than to you.
- Check three things before accepting a delta, in order: instrument, population, size. Most people check size only, because size is the number printed in bold.
- Treat a single-run move as a hypothesis. Re-run before you rewrite a page, brief an agency, or take a fall to a board. Ask whether the cell is still shut on the third run.
- Read the answers, not only the score. Open the cell that lost you the most points and read what the engine actually said. In our worst cell the engine had answered a different question entirely, and no number could have told us that. The text did, in one paragraph.
Frequently asked questions
Does a drop in my AI visibility score mean I lost visibility?
Not on its own. A drop means the measured cells named you less often than last time, which can be caused by your visibility falling, by a different model answering, by an edited question, or by ordinary variance. Establish which before you act: ours fell 42 points and nothing about the company had changed.
How many runs before a fall is worth acting on?
Three consecutive runs in the same direction before you rewrite a page or brief an agency on it, and five before a noise test has enough history to judge the delta against that metric's own typical step. Two runs give you a delta, not a trend. A single fall is a hypothesis worth re-measuring, not a result worth spending on.
The engine name did not change, so the instrument did not change — correct?
No. The engine name is a product label and the model behind it is replaced on the vendor's schedule. On 13 August two OpenAI models answered our identical three questions on the same day: one named the brand in three answers of three, the other in one of three. Same engine name on both rows.
Should a rise get the same scrutiny as a fall?
Yes, and in practice it gets far less, because nobody interrogates a flattering number. A model swap can lift a score exactly as easily as it can sink one, and a run that lost an engine to an API failure can score higher than a complete one purely because it normalised over fewer answers. Run instrument, population and size against good news too.
How do I tell whether a competitor really took my place?
Read the answer that names them and check what it was answering. Two of the rivals our report surfaced were cited from Chrome Web Store listings for personal autofill extensions, in response to a question about services for e-commerce. They had not displaced anyone; the question had drifted underneath the measurement.
Why publish a failure in your own tool?
Because the alternative is asking you to trust a number on the strength of its vendor never having reported being wrong, which is not evidence of anything. A measurement tool earns credibility by publishing the cases where it misled its own authors, and by naming the repairs that are still open.
Read your next report against these questions
If you have an AI-visibility report in front of you, take the seven demands above to it before you take its conclusions to anyone else. If it cannot name the model that answered, the rest of it is a story about a number rather than a measurement of one.
The tracker used throughout this article is aeo-platform, open source and free to run against your own domain, thresholds and all. If you would rather see where you currently stand before deciding whether any of this applies to you, our free AI visibility check runs the same basic question and shows you the answers it collected, not only the score it computed.


