A hypothesis connects an action to an outcome
‘Improve SEO’ is not an experiment. State the change, population, metric, and period—for example, unique titles on product detail pages will improve non-brand CTR among indexed target URLs over 28 days.
Write the condition for rejecting the hypothesis. Insufficient impressions or unindexed pages are measurement failures, not evidence of no effect.
Keep the treatment as small as practical
One template and one element are easier to explain. Changing titles, body copy, internal links, and speed together makes attribution impossible.
With low traffic, a before/after comparison may be more realistic than forced A/B splits. Record seasonality, campaigns, brand demand, and algorithm updates as confounders.
The observation window includes crawl delay
Search experiments have crawl and index delay between deployment and exposure change. Do not average from release day; identify a stable window after target URLs are recrawled and indexed.
Aggregate volatile metrics weekly and keep both counts and rates. A CTR increase with a changed query mix is not necessarily the same outcome.
Failures and side effects are results
If clicks rise but conversions fall, or head URLs improve while long-tail URLs disappear, record the side effect. Optimizing one search metric can move away from the business goal.
A failed hypothesis prevents repeated mistakes. Preserve code and reports with a searchable conclusion that the effect did not reproduce under stated conditions.
Pre-release checklist
- ✓ Does the hypothesis include change, population, metric, and period?
- ✓ Did you predefine success, failure, and unmeasurable conditions?
- ✓ Did you record concurrent releases and external variables?
- ✓ Did you compare stable windows after crawl and index delay?
- ✓ Did you check conversion and behavioral side effects?
Frequently asked questions
How many weeks should an SEO experiment run?
There is no universal duration. Run long enough after recrawl to include sufficient exposure and normal cycles. Predefined sample and stable-window criteria matter more than a fixed number of weeks.
Is it invalid without a control group?
Causal inference is weaker, but the observation can still be useful. Document baselines, confounders, and limitations, and avoid definitive claims.