
Quick answer: SpeedCurve and Calibre are the specialists, strongest at synthetic testing and tracking change over time. Datadog RUM and New Relic fit when performance data belongs beside your existing infrastructure monitoring. Sentry Performance is convenient if Sentry already handles your errors. Vercel Speed Insights is the low-friction choice if you deploy there. Open-source RUM collection works when you want to own the data. Choose on synthetic-versus-real-user first — those are different tools, not different features.
Most performance tooling reports page-level scores, which attribute every millisecond to the page. That is the right model when you own everything on it. It is the wrong model when the question is what one embedded script costs — an onboarding SDK, an analytics tag, a chat widget — because the answer a customer wants is not "the page scored 72" but "how much worse is this page because of your script". Those are different measurements and only one of them is a defensible answer.
| Tool | Synthetic | Real user | Best for |
|---|---|---|---|
| SpeedCurve | Strong | Yes | Tracking change and benchmarking over time |
| Calibre | Strong | Yes | Scheduled testing with clear budgets |
| Datadog RUM | Yes | Strong | Beside existing infrastructure monitoring |
| New Relic | Yes | Strong | Full-stack correlation |
| Sentry Performance | No | Yes | Teams already using Sentry for errors |
| Vercel Speed Insights | No | Yes | Low-friction on Vercel deployments |
| Open-source RUM | No | Yes | Owning the data and the model |
Capabilities and pricing checked September 2026 and both move — verify on each vendor's own pages before committing.
SpeedCurve and Calibre are built for the job rather than extended into it, and it shows in the parts that matter over months: consistent synthetic runs, budgets you can fail a build against, and competitive comparison. If performance is something you manage continuously rather than investigate occasionally, start here.
Datadog RUM and New Relic earn their place through correlation — a slow page next to the slow query that caused it, in one place. The trade is that you are buying into a platform, and their performance offering is one part of a much larger product and bill.
Sentry Performance and Vercel Speed Insights are the pragmatic choices: if you already run one of them, the marginal cost of turning on performance data is low and the data is usually good enough to find a regression. Open-source RUM — collecting the browser's own performance entries into your own store — is more work than it sounds and gives you exactly the model you define, which matters if you need per-script attribution nobody sells.
If you ship an SDK into other people's applications — which is our situation, so read this as an interested party describing its own problem — page scores are not the measurement. Three things are: bytes transferred and parsed for your script, main-thread time attributable to it, and the difference in Core Web Vitals between sessions where it loaded and sessions where it did not.
The third is the only one that answers what a customer actually asks, and it needs a controlled comparison rather than a dashboard reading. Our SDK performance budget worksheet covers setting the numbers and the measurement method, and the release acceptance checklist covers holding to them across releases. Publishing those numbers is uncomfortable and it is the only version of this claim worth making.
Which tool should I start with?
SpeedCurve or Calibre if performance is an ongoing concern, Datadog or New Relic if it belongs beside infrastructure data, Sentry or Vercel if you already run them.
Synthetic or real-user monitoring?
Both, eventually. A regression is easiest to spot in synthetic data and easiest to size in real-user data.
How do I measure one script's cost?
Bytes, main-thread time, and a with-versus-without comparison of Core Web Vitals. Page scores attribute everything to the page and cannot answer it.