Write a falsifiable hypothesis
A useful hypothesis connects a change, a mechanism and an observable outcome. ‘The new page will perform better’ cannot guide analysis. State which audience encounters the change, why their behaviour may differ and what result would challenge the belief. Keep assumptions visible. If the mechanism is unclear, the first experiment may need to test comprehension or usability before a larger traffic split.
Choose a metric near the decision
Metrics should sit close to the action the experiment is designed to improve. A content comparison may focus on task completion or qualified next steps rather than raw time on page. A navigation change may use successful route completion and backtracking. Select one primary metric to prevent opportunistic interpretation, then add a small number of diagnostic measures that explain why it moved.
Add guardrails
Guardrails capture harm the primary metric can conceal. A faster flow is not an improvement if errors, cancellations or support contacts rise. Common guardrails include accessibility failures, opt-outs, complaint rates, latency, error frequency and downstream quality. Set them before seeing results. A guardrail is meaningful only when the team agrees what level will pause or reverse the change.
Understand denominator and exposure
Always define who had the opportunity to perform the measured action. Using all site visitors as a denominator for a control seen by only one page segment will dilute the effect. Record eligibility, assignment, actual exposure and exclusions. Check whether bots, internal users, repeat visits or tracking prevention alter counts. Consistent definitions matter more than a dashboard with many decimal places.
Plan duration and stopping
Set a minimum observation window that covers normal weekly patterns and gives delayed outcomes time to appear. Avoid stopping when a chart first crosses a convenient line. Repeated peeking increases the chance of acting on noise unless the analysis accounts for it. Operational incidents, campaigns and holidays should be logged. If volume is low, a carefully observed sequential rollout may be more honest than an underpowered significance claim.
Publish a decision note
At the end, write what changed, who was included, the outcome, guardrails, limitations and decision. Link the implementation and source definitions. Distinguish ‘no detected difference’ from proof that two options are identical. Record follow-up questions and an expiry date for the conclusion. Decision notes reduce repeated debates and make the organisation’s learning inspectable instead of trapped in dashboards.
Worked example: a clearer onboarding explanation
Suppose a team believes a short example will help new users choose the right setup path. The hypothesis links the example to reduced uncertainty, measured by correct path completion among eligible new users. Backtracking, support contacts and accessibility errors are guardrails. Assignment, exposure and completion are defined before release. The test runs across two normal weekly cycles and excludes internal traffic. If completion improves beyond the practical threshold with stable guardrails, the example can ship. If clicks rise but wrong-path corrections also rise, the mechanism is contradicted even though the first chart looks positive.
Analysis failure modes
Dashboards encourage teams to inspect every segment until something moves. Prevent this by naming the primary analysis and a limited set of diagnostics in advance. Check sample-ratio mismatches, missing events, duplicate identities, timezone boundaries and delayed outcomes before interpreting lift. Do not remove inconvenient days without a documented operational reason. Segment findings are hypotheses unless the test was designed and powered for them. Report absolute counts and baseline alongside percentages, since a dramatic relative change can represent very few people. Preserve the query or source definition so another analyst can reproduce the number.
Maintenance rhythm
Metric definitions require version control. When an event, denominator or attribution window changes, date the change and avoid splicing incomparable periods without annotation. Audit high-stakes events after releases and compare client events with server or operational records where possible. Retire dashboards that no longer inform a decision. Review guardrails as products and harms evolve. The measurement system should make uncertainty visible, not decorate it with precision. A short decision note with a reproducible source is more durable than a crowded dashboard whose filters nobody remembers.
Focused workshop
Use this short working session to turn the guide into a decision-ready artefact. Keep the scope to one live example and record assumptions beside the output.
- Rewrite one broad goal as change, mechanism and outcome.
- Choose a primary metric and two harm guardrails.
- Draft an end-of-test decision note before data is available.
Definition of ready: another practitioner can inspect the evidence, understand the proposed change and name the result that would stop or reverse it.
Review questions
What evidence supports the problem? Which audience and context does it describe? What is the smallest change that can test the mechanism? Which signal is close enough to the decision to be useful? What harm could a single success metric hide? Record answers before implementation, then return to them when the review window closes.