PLG OS

Implementation worksheet · 6 min read

An onboarding case-study evidence collection template

Collect the evidence while the onboarding change is live, not when someone asks for a story three months later. The packet needs the before and after flow (captures with version IDs and dates), an exposure log showing who saw which version, an outcome metric declared before launch, the comparison design, the result with its uncertainty, a list of everything else that shipped in the same period, and written customer approval for every number and quote you plan to publish. If the comparison was before-and-after rather than against a holdout, the case study has to say so and avoid causal words like 'drove' or 'caused'.

Onboarding case studies persuade because they're concrete, and they're hard to check for the same reason. A reader sees 'activation rose after the new checklist' and can't tell whether a pricing change, a seasonal spike or a sales push did the work. This template is for the product or customer team that ran one onboarding change, in its own product or a customer's, and wants a claim a skeptical reader could audit. Boundary: one flow change, from the week before launch to the end of the declared measurement window.

Put it into practice

1. Open the packet on launch day

Create the folder and fill four fields before the change ships: flow ID and version, launch date, the outcome metric and how it's defined, and the comparison you intend to run. Anything decided after the results are visible is a post-hoc story, however honest the intent.

2. Capture both versions of the flow as users saw them

Screenshot or record every step of the old and new flow, with version and date. Flows get edited in place, and without captures nobody can say later what 'the new onboarding' actually was. Include the empty states and the checklist as a brand-new account sees them.

3. Keep the exposure log, not only the outcome

Record who was eligible, who saw each version and who completed it. If you kept a holdout on the old flow, log assignment separately from exposure. The outcome number means little without knowing how many people it describes.

4. List everything else that shipped

Pricing changes, sales campaigns, marketing launches, outages, holidays, a new signup source. Ask the customer's team too, since they know about their own launches and you don't. With a randomized holdout these affect both groups; with before-and-after they can explain the whole result.

5. Match the claim to the design

A randomized holdout with an interval that excludes zero can support 'caused'. Before-and-after supports 'followed' or 'coincided with'. A customer saying 'onboarding feels smoother' is a quote, attributed as one. Mixing the three in a sentence is how case studies overstate.

6. Get written approval for each number and each quote

Name, logo, figures and quotes are separate permissions. Record what the customer approved, what they declined, and the date. A declined figure stays out; it doesn't become 'significantly improved'.

7. Save the query and when it ran

Store the query that produced each number, who ran it and when, so it can be re-run. A figure nobody can reproduce six months later is a figure you should stop quoting.

Case-study evidence packet

Copy this structure into your review document and record your observed result for each row.

Case-study evidence packet
FieldWhat to collectWorked example (illustrative)Status
Flow and versionFlow ID, version, launch datesetup-checklist v3, launched 2 JuneCollected
Before and after capturesEvery step of both versions, datedv2 had 6 steps; v3 has 4Collected
Outcome metricName and definition, declared before launchFirst shared project within 7 days of signupDeclared 28 May
Comparison designHoldout, staggered rollout or before-and-after10% of eligible signups kept on v2 for 6 weeksCollected
Exposure logEligible, exposed and completed, per version4,812 eligible; 4,331 on v3; 481 held outCollected
Result with uncertaintyDifference and 95% interval, not one number31.2% vs 27.9%; interval about -0.9 to +7.5 pointsInconclusive; say so
Other changes in the windowPricing, campaigns, outages, seasonalityAnnual-plan discount launched 15 JuneListed; hit both groups
Customer approvalWritten sign-off per number and quoteName and quote approved; figures declinedPartial
Query and run dateSaved query, who ran it, whenactivation_v3.sql, run 20 JulyCollected
Claim wordingMatches design and interval'Early signal, not yet conclusive'Agreed with customer

A failure worth checking

Publishing a before-and-after as if it were an experiment. Activation rose after the new checklist shipped, the case study says the checklist drove it, and a reader later notices the launch coincided with the customer's biggest campaign of the year. Nothing in the numbers was false; the causal word was. The worked example above shows the subtler version: a real holdout that was too small, where a 3.3-point difference sits inside an interval running from slightly negative to strongly positive. The honest write-up says the signal is early.

Common questions

What if the customer won't approve any figures?

Publish what they approve and say plainly what's missing. A named customer, a described change and an approved quote make a weaker case study, and an honest one. Swapping declined figures for vague adjectives is worse than leaving them out.

How large should the holdout be?

Large enough that the interval around the difference is narrow enough to act on, which depends on the baseline rate and the smallest effect you care about. Run a power calculation before launch. The illustrative example shows what an undersized holdout produces: a positive difference you can't tell apart from zero.

Basis and scope

This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.

Continue with PLG OS

Explore onboarding →