Implementation worksheet · 7 min read
A user support deflection measurement protocol
Define deflection as the reduction in support contacts per active user caused by in-app help, measured against a randomized holdout that doesn't get the help, over a window declared in advance. Don't count 'viewed an article and didn't open a ticket' as a deflected ticket: many of those people would never have contacted support, and some contacted support through another channel. Count contacts across every channel you can attribute, check that problems were actually resolved rather than postponed, and report the difference with an interval. Without a holdout, report a before-and-after association and label it as one.
Support deflection is one of the easiest numbers in software to overclaim, because the thing being counted didn't happen. A common shortcut counts help sessions that ended without a ticket and calls each one a deflection; that measures help usage, not avoided support. This protocol is for a team adding or changing contextual help, a help widget or in-app guidance, and wanting a number a finance reviewer would accept. Boundary: signed-in users of one product, all attributable support channels, one declared window.
Put it into practice
1. Write the definition and the metric before launch
Primary metric: support contacts per 100 weekly active users, all channels combined. Secondary: contacts on the topics the help covers, and a resolution guardrail. Declare the window and any extension rule now. A definition written after seeing the data will drift toward the flattering one.
2. Randomize a holdout at the right level
For team products, randomize by account, because colleagues share answers and a user-level split leaks. The holdout doesn't get the in-app help; the normal help center link stays available to everyone. Support contacts are rare events with high variance, so size the holdout with a power calculation rather than habit.
3. Count contacts across every channel
Tickets, chat, email, calls and community posts where you can tie them to a user. Help can move contacts rather than remove them: fewer tickets and more chats is a smaller win than the ticket chart shows. List channels you can't attribute, like unidentified phone calls, and say so in the report.
4. Check resolution, not just absence
For each help topic, track the product event that means the user succeeded, and recontacts within 14 days. A user who read an article, gave up and wrote in a week later wasn't deflected. A user who gave up and churned quietly wasn't either.
5. Compare everyone assigned, not only those who opened help
People who open help differ from people who don't: newer, more stuck, sometimes more motivated. Comparing help openers with non-openers measures those differences, not the help. The headline comparison is the whole treatment group against the whole holdout (intention-to-treat).
6. Look for novelty and calendar effects
New widgets can draw curiosity clicks in the first weeks, and billing dates and releases move contact volume. The holdout absorbs calendar effects because both groups live through them; novelty only hits the treatment group. Report the effect for early and late weeks separately and treat the late figure as the better guide to steady state.
7. Report the difference with its interval, then the cost
State contacts avoided per 100 users per week with a 95% interval clustered by account, the channel shifts, and the resolution check. Only then convert to cost, carrying the interval through. A single savings figure without the range is the overclaim this protocol exists to prevent.
Deflection measurement worksheet
Copy this structure into your review document and record your observed result for each row.
| Item | Protocol rule | Illustrative value | Check |
|---|---|---|---|
| Definition | Fewer contacts per active user than the holdout, all channels | Contacts per 100 weekly active users | Written before launch |
| Randomization | By account for team products | 90% help on, 10% holdout | Observed split checked against 90/10 |
| Window | Declared in advance, covers a billing cycle | 6 weeks | Not extended after looking |
| Tickets per 100 users a week | Treatment vs holdout | 2.4 vs 3.1 | Web form and email combined |
| Chats per 100 users a week | Treatment vs holdout | 0.7 vs 0.5 | Channel shift counted, not ignored |
| All contacts per 100 users a week | Primary metric | 3.1 vs 3.6 | Difference 0.5 |
| Interval | 95%, clustered by account | About 0.1 to 0.9 fewer | Excludes zero; report the range |
| Resolution guardrail | Task success on help topics; 14-day recontact | No drop against holdout | Checked |
| Novelty check | Weeks 1-2 vs weeks 5-6 | 0.7 early, 0.4 late | Late figure used for planning |
| Claim wording | Matches the design and the interval | 'About 0.5 fewer contacts per 100 users a week (0.1 to 0.9)' | No savings figure without the range |
A failure worth checking
Counting help views as saved tickets. A hypothetical dashboard shows 40,000 article views last quarter and 'estimated tickets deflected: 40,000', multiplied by a cost per ticket. Most of those viewers were never going to write in; some wrote in anyway through chat; a few read three articles, gave up and called. The figure isn't a generous estimate of deflection. It's a different metric wearing deflection's name, and it will be quoted in a budget meeting as if it were real.
Common questions
Can we measure deflection without a holdout?
You can measure an association. Compare contact rates per active user before and after launch, list what else changed, and present it as 'contacts per user fell after launch', not as deflection caused by the help. If the decision matters, a holdout is cheap compared with acting on a wrong number.
What about the 'Was this helpful?' buttons?
They're useful for improving articles and weak evidence of deflection. The people who answer are a self-selected minority, and 'yes' doesn't mean they would otherwise have contacted support. Treat the rate as a content signal, not as the outcome.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.