A real SEO experiment needs four things decided before you touch anything: the exact change, the metric it should move, the measurement window, and the condition under which you’ll call it wrong. We run this pattern on our own site for every on-page fix we ship – six proposals on file, six pre-registered wrong-if conditions – and this article walks through what we changed, what we said would prove us wrong, and what actually happened when we checked. Two of the six are still open. One taught us that a low-traffic site can go a month without enough search data to call a result either way.
Most SEO advice tells you what to change. It rarely tells you how you’d know if the change didn’t work. We run organic-os, the AI SEO agent that proposes and applies fixes on this site, under a rule borrowed from science before marketing: no proposal gets approved without a falsifiability section stating the exact condition that would prove it wrong. Here’s what that looks like run six times, with the results as they actually stand today.
Why most SEO tests fail before they start
A test without a pre-registered wrong-if condition isn’t an experiment – it’s a change you’ll rationalize afterward. If rankings go up, you credit the change. If they don’t, you blame seasonality, a core update, or “these things take time.” Nothing was ever falsifiable, so nothing was ever actually tested. The fix is small and unglamorous: write down, before you ship, the specific number and the specific window that would make you reverse the change or admit it didn’t help. Every proposal on this site does that in a section literally titled “Falsifiability.” It’s the only part of the process that keeps a fix from becoming a permanent, unverified belief.
Five experiments we’ve actually run, with exact setups
1. Title-length trim. Two posts had auto-templated titles running 69 and 73 characters, past the roughly 60-character point where SERPs start truncating. Change: set an explicit title override on each post, dropping the auto-appended site-name suffix. Metric: whether the full title renders untruncated. Window: next snippet inspection, no fixed day count. Wrong-if: a snippet check still shows truncation 7+ days after apply. Result: applied and verified live against both URLs’ rendered title tag – full titles render, wrong-if condition not triggered.
2. Meta-description trim. Two pages had descriptions at 177 and 170 characters, past the roughly 160-character truncation point. Change: cut only the redundant closing clause from each – no new claims added. Wrong-if: either description still over 160 characters after apply, or truncation still visible 7+ days later. Result: applied and verified; one check briefly looked like a mismatch until we accounted for RankMath HTML-entity-encoding the apostrophe in the live output – the underlying text matched exactly once decoded.
3. Homepage H1, schema, and keyword fix. Three changes bundled into one proposal: remove a duplicate H1 (the theme’s masthead was rendering as a second H1 alongside the hero), add the primary keyword to visible body copy (it appeared zero times outside meta fields), and remove an incorrectly-typed schema node. Wrong-if: any of the three still present on re-fetch. Result: partial. The keyword and hero-copy change landed and verified. The duplicate H1 and the schema-type change hit a hard wall – both live only in places the connected WordPress role can’t reach over the REST API – and remain open. Checking the live page again while writing this, more than a month later, the homepage still renders two H1 elements.
4. Homepage title and meta rewrite. The homepage title and description were both over length and contained none of this site’s tracked keyword targets. Change: shorten both and add a focus keyword. Wrong-if: average search position for the target queries showing no improvement, or click-through dropping on any query the page already ranked for, 30 days out. This is the experiment with the most instructive result – see the next section, because the honest answer isn’t pass or fail.
5. Blank metadata on three orphaned posts. A whole-site scan found three published posts with every RankMath field – title, description, focus keyword – completely empty. Change: set all three fields on each post, paraphrased directly from each post’s own opening paragraph so no new claim was introduced. Wrong-if: search impressions for the three URLs showing no change 30 days after apply, or RankMath silently overriding the explicit values with its own fallback. Result: still mid-flight, and worth reading honestly rather than rounding up – see below.
Reading results honestly on a low-traffic site
Experiment 4’s 30-day window closed before this article was written, so we pulled the actual Search Console data for the homepage over that exact span. The result: search impressions on 8 of the 32 days, ranging from 0 to 4 per day, zero clicks across the entire window, and a position that swung from as good as 2.25 to as poor as 6 depending on which of those 8 days you look at. A broader pull for any query containing the target phrase over roughly the same period turned up exactly one matching row in the whole month.
That’s not a failed experiment. It’s a sample size too small to call anything. A site with this little query volume can’t distinguish a real ranking shift from a single day’s noise, and treating either the up-swing or the down-swing as the answer would be exactly the kind of unfalsifiable storytelling we’re trying to avoid. The honest read is: inconclusive, re-check in another 30 days once more query data has accumulated. We also know applied fixes on this site have been silently reverted by unattributed bulk edits before – one incident touched six pages within a four-second window – so a clean 30-day read requires checking that the change is still live, not just that a metric moved.
Our own open experiments and their current state
Two of the six are unresolved right now, and we’re naming both rather than letting them quietly age out. The homepage’s duplicate H1 and incorrect schema type are still live, more than a month past the original proposal, blocked on capability rather than decision – both changes live in places our connected WordPress role can’t write to. And the three-post metadata fix is genuinely between states: checking post 6 directly today shows its description field is now set, but to different wording than the proposal specified, while its title and focus keyword are still blank. That’s neither the “applied” nor the “untouched” outcome the wrong-if condition anticipated – it’s a third state the experiment template didn’t account for, and it’s a more useful thing to log than to smooth over.
The pattern across all six: kill criteria only work if you actually go back and check them, on a fixed schedule, against live data – not against your memory of what you meant to ship. That discipline matters more on a small site than a large one, precisely because there’s less data to lean on and more temptation to read a coincidence as a result. It’s the same evidence-first habit behind how we test whether AI-assisted content actually ranks, the review boundary we draw in what still needs a human in this pipeline, the log we keep on who’s actually reading our pages before a person does, the keyword grouping described in how we structure a B2B SaaS keyword portfolio, and the approval loop documented in what an AI SEO agent actually does.
Frequently asked questions
What four things does a real SEO experiment need before you start?
The exact change, the metric it should move, a fixed measurement window, and a wrong-if condition – the specific result that would prove the change didn’t work. Without that last part, a test can’t fail, which means it was never really a test.
Why can’t a low-traffic site trust a 30-day ranking check?
Because a handful of impressions spread across a month is noise, not signal. One real 30-day check on our own homepage found search impressions on only 8 of 32 days and zero clicks the entire window – too sparse to call the result a pass or a fail either way.
What happens when an experiment only partially lands?
It gets recorded as partial, not rounded up to done. One of our own fixes landed two of three changes and left the third blocked on a capability wall; the record still names the exact remaining step rather than marking the whole thing applied.
What’s the biggest risk to a clean experiment result?
An unattributed revert. Applied fixes on this site have been silently reverted by external bulk edits before, in one case within four seconds of a batch operation. A results check needs to confirm the change is still live, not just that a metric moved.









