Implementation worksheet · 6 min read
An end-of-month onboarding experiment review
Review every onboarding experiment that reached its planned end date this month in the same order. First confirm the observed split matches the planned split; a mismatch blocks the readout until it's explained. Then confirm exposure was logged when users could actually see the variant. Read only the metric declared before launch, with its interval, check the guardrails, and look at the effect by signup week. Finally record one decision: ship, stop, or extend under a rule written before launch. Experiments still running get a health check, not a decision. Keep the log, null results included.
Onboarding experiments are easy to start and awkward to finish. Tours and checklists get edited mid-test, results are read early, and a metric that 'looked better' quietly replaces the one that was declared. A monthly review with a fixed agenda is a cheap defense. Actor: the product team plus whoever owns experiment analysis. Boundary: experiments on in-app onboarding (tours, checklists, prompts, empty states) whose planned end date fell in the month.
Put it into practice
1. List what ended, what's running and what was abandoned
Abandoned experiments count. Write down why each was stopped; a pattern of abandoned tests often points at tooling or scope rather than at bad ideas.
2. Check the split before anything else
Compare assigned counts with the planned ratio using a chi-square test. A sample ratio mismatch means something removed users unevenly from one arm, like a variant that fails to load or eligibility applied after assignment, and any result read on top of it is suspect.
3. Check when exposure was logged
Users assigned but never shown the variant dilute the effect. Analyze everyone assigned (intention-to-treat) and report the exposure rate beside it, so a small effect can be read as 'weak variant' or 'variant rarely seen'.
4. Read the declared metric with its interval
Only the primary metric declared before launch, reported as a difference with a 95% interval. Secondary metrics are labelled exploratory and can't decide anything on their own.
5. Check the guardrails
Support contacts, dismissal rate, time to value, page performance. A win on the primary metric with a broken guardrail is a trade to make deliberately, not a win.
6. Look at the effect by signup week
Onboarding tests are partly protected from novelty, because every user meets the flow for the first time. What can change over a month is who signs up: a campaign or a launch shifts the mix. If the effect differs sharply by week, find out why before shipping.
7. Record the decision and what would change it
Ship, stop or extend, with the numbers and the date. Extensions happen only under a rule written before launch. The log is how next quarter's plan avoids re-running last quarter's null result.
Monthly experiment review log
Copy this structure into your review document and record your observed result for each row.
| Check | Rule | Experiment A (illustrative) | Experiment B (illustrative) |
|---|---|---|---|
| Hypothesis | Written before launch | Shorter checklist raises 7-day activation | Tour step on import raises first imports |
| Planned split | Declared before launch | 50/50 | 50/50 |
| Observed split | Chi-square check for mismatch | 10,044 vs 9,956; no mismatch | 6,410 vs 5,590; mismatch |
| Cause of mismatch | Found before any readout | No mismatch to explain | Assignment logged after the tour script loaded; failed loads dropped |
| Exposure rate | Share of assigned users who saw the variant | 96% vs 97% | Not read until the split is fixed |
| Primary metric | Declared metric, 95% interval | 31.4% vs 30.0%; +1.4 points (+0.1 to +2.7) | Not read |
| Guardrails | Support contacts, dismissals, time to value | No change detected | Not read |
| Effect by signup week | Similar across weeks? | Similar in all four weeks | Not read |
| Decision | Ship, stop or extend by the pre-set rule | Ship to all eligible users | Stop; fix assignment logging; rerun |
A failure worth checking
Reading the lucky metric. The declared metric, activation within seven days, is flat. A secondary chart shows checklist items completed up sharply, and the review ships the variant on that. But the variant changed the checklist, so of course checklist items moved. The agenda reads the declared metric first and labels everything else exploratory for exactly this reason.
Common questions
What if an experiment is almost significant?
Then its interval includes effects near zero, and the honest options are to stop, or to extend under a rule declared before launch. Extending because the result is close, without such a rule, raises the chance of a false positive. Record which one you did.
Should we review experiments that are still running?
For health, yes: split, exposure and guardrails. For decisions, no. Reading the primary metric before the planned end and acting on it is the peeking problem that a fixed end date exists to prevent.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.