PLG OS

Implementation worksheet · 6 min read

A Product Tour Broken-Target Incident Runbook

Detect it yourself rather than waiting for a report: a step whose target resolution rate drops sharply is the signal, and it is available before any customer notices. Then contain first — disable the affected step or tour, which must be possible remotely without a deploy — establish scope by counting affected sessions, and only then fix. The fix is usually not yours: the target disappeared because the customer's application changed, so the durable answer is a stable selector attribute in their markup rather than a more clever selector in your code. A tour that survives one refactor by guessing harder will break on the next one.

In-app guidance depends on selectors pointing at elements in an application you do not control and cannot see changing. A customer ships a refactor on a Tuesday and a tour that worked for months now highlights empty space or silently skips. Nothing errors on either side, and the usual first signal is a support ticket days later.

Put it into practice

1. Instrument target resolution as a metric, per step

Every step attempt records whether its target was found. A step resolving 99% then 4% is unmistakable and available within an hour. Without this metric the detection path is a customer noticing, which is days and a damaged impression.

2. Alert on a drop, not on an absolute rate

Some steps legitimately resolve at 60% because the target only exists in certain states. A sudden fall from that step's own baseline is the signal; a fixed threshold generates noise on the normal cases and silence on the abnormal ones.

3. Contain remotely, before diagnosing

Disable the step or the tour without a deploy. A tour pointing at nothing is worse than no tour, and every minute of diagnosis is more sessions seeing it. If remote disable does not exist, build it before the next incident rather than during it.

4. Count affected sessions and tell the customer first

How many users hit the broken step, over what period. Telling the customer before they tell you changes the conversation entirely — from a complaint about your reliability to a joint problem, which is what it is.

5. Find what changed on their side

Almost always a class rename, a component restructure, a conditional render, or a flag that changed what renders. Ask for the deploy time and compare against the resolution drop; the correlation is usually immediate and it locates the change without access to their code.

6. Fix with a stable attribute, not a cleverer selector

The durable fix is a dedicated attribute in their markup that exists for this purpose and is not touched by styling refactors. A more specific CSS selector survives this refactor and not the next one. This is a conversation with the customer rather than a change on your side, and it is worth having once per integration.

7. Add the step to a monitored set and re-enable deliberately

Re-enable after confirming resolution recovers in real sessions, not after deploying the fix. Those are different moments, and the gap between them is where a second incident starts.

8. Write the post-incident note for the customer

What broke, how many were affected, what prevents it recurring, and what you need from them. The ask for a stable attribute lands far better attached to a specific incident than as a general best practice.

The runbook

Copy this structure into your review document and record your observed result for each row.

The runbook
PhaseActionTargetOwner
Detecttarget resolution drop alertwithin 1 hourautomated
Containdisable step or tour remotely5 minon call
Scopecount affected sessions and users15 minon call
Notifytell the customer first30 minaccount owner
Diagnosecorrelate with their deploy time1 hourengineering
Fix (short term)update selectorsame dayengineering
Fix (durable)request a stable attributenext releaseaccount owner
Verifyresolution recovers in real sessionsafter deployautomated
Re-enableonly after verification—on call
Post-incident notesent to the customerwithin 2 daysaccount owner

A failure worth checking

Waiting for the support ticket. Resolution for one step drops on Tuesday afternoon. Nobody is watching that metric, so the first signal is a message on Friday from a customer whose new users have been seeing a highlight over blank space for three days. The technical fix takes ten minutes; the damage is that the customer discovered a fault in your product before you did, in front of their users. The detection metric costs one boolean per step attempt.

Common questions

Should tours fail open or closed?

Closed. A step that cannot find its target should skip or end the tour, not render pointing at nothing. Guidance that is absent is a gap; guidance that is confidently wrong is a defect the user attributes to the product around it.

How do we get customers to add stable attributes?

Ask at integration time, and ask again attached to a real incident. The general request is easy to defer; the same request alongside 'this affected 340 of your users on Tuesday' is usually actioned in the next release.

Basis and scope

This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.

Continue with PLG OS

Explore onboarding →