PLG OS

Implementation worksheet · 6 min read

An end-of-month onboarding experiment review

Review every onboarding experiment that reached its planned end date this month in the same order. First confirm the observed split matches the planned split; a mismatch blocks the readout until it's explained. Then confirm exposure was logged when users could actually see the variant. Read only the metric declared before launch, with its interval, check the guardrails, and look at the effect by signup week. Finally record one decision: ship, stop, or extend under a rule written before launch. Experiments still running get a health check, not a decision. Keep the log, null results included.

Onboarding experiments are easy to start and awkward to finish. Tours and checklists get edited mid-test, results are read early, and a metric that 'looked better' quietly replaces the one that was declared. A monthly review with a fixed agenda is a cheap defense. Actor: the product team plus whoever owns experiment analysis. Boundary: experiments on in-app onboarding (tours, checklists, prompts, empty states) whose planned end date fell in the month.

Put it into practice

1. List what ended, what's running and what was abandoned

Abandoned experiments count. Write down why each was stopped; a pattern of abandoned tests often points at tooling or scope rather than at bad ideas.

2. Check the split before anything else

Compare assigned counts with the planned ratio using a chi-square test. A sample ratio mismatch means something removed users unevenly from one arm, like a variant that fails to load or eligibility applied after assignment, and any result read on top of it is suspect.

3. Check when exposure was logged

Users assigned but never shown the variant dilute the effect. Analyze everyone assigned (intention-to-treat) and report the exposure rate beside it, so a small effect can be read as 'weak variant' or 'variant rarely seen'.

4. Read the declared metric with its interval

Only the primary metric declared before launch, reported as a difference with a 95% interval. Secondary metrics are labelled exploratory and can't decide anything on their own.

5. Check the guardrails

Support contacts, dismissal rate, time to value, page performance. A win on the primary metric with a broken guardrail is a trade to make deliberately, not a win.

6. Look at the effect by signup week

Onboarding tests are partly protected from novelty, because every user meets the flow for the first time. What can change over a month is who signs up: a campaign or a launch shifts the mix. If the effect differs sharply by week, find out why before shipping.

7. Record the decision and what would change it

Ship, stop or extend, with the numbers and the date. Extensions happen only under a rule written before launch. The log is how next quarter's plan avoids re-running last quarter's null result.

Monthly experiment review log

Copy this structure into your review document and record your observed result for each row.

Monthly experiment review log
CheckRuleExperiment A (illustrative)Experiment B (illustrative)
HypothesisWritten before launchShorter checklist raises 7-day activationTour step on import raises first imports
Planned splitDeclared before launch50/5050/50
Observed splitChi-square check for mismatch10,044 vs 9,956; no mismatch6,410 vs 5,590; mismatch
Cause of mismatchFound before any readoutNo mismatch to explainAssignment logged after the tour script loaded; failed loads dropped
Exposure rateShare of assigned users who saw the variant96% vs 97%Not read until the split is fixed
Primary metricDeclared metric, 95% interval31.4% vs 30.0%; +1.4 points (+0.1 to +2.7)Not read
GuardrailsSupport contacts, dismissals, time to valueNo change detectedNot read
Effect by signup weekSimilar across weeks?Similar in all four weeksNot read
DecisionShip, stop or extend by the pre-set ruleShip to all eligible usersStop; fix assignment logging; rerun

A failure worth checking

Reading the lucky metric. The declared metric, activation within seven days, is flat. A secondary chart shows checklist items completed up sharply, and the review ships the variant on that. But the variant changed the checklist, so of course checklist items moved. The agenda reads the declared metric first and labels everything else exploratory for exactly this reason.

Common questions

What if an experiment is almost significant?

Then its interval includes effects near zero, and the honest options are to stop, or to extend under a rule declared before launch. Extending because the result is close, without such a rule, raises the chance of a false positive. Record which one you did.

Should we review experiments that are still running?

For health, yes: split, exposure and guardrails. For decisions, no. Reading the primary metric before the planned end and acting on it is the peeking problem that a fixed end date exists to prevent.

Basis and scope

This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.

Continue with PLG OS

Explore onboarding →