Skip to content

Proof · The Experiment Ledger

Experiments we shipped,
and what they moved.

A working index of the A/B tests Tyrell Lab has run for SaaS teams since 2019 — baseline, variant, lift, and the days it took. Numbers first. Story after. No highlight reel of vanity wins; only the experiments a senior partner would defend in a quarterly review.

Updated quarterly. 412 winning experiments and 1,180 total runs as of last refresh.

  • 412 Winning experiments shipped since 2019
  • +32% Average activation lift in the first 90 days
  • 92% Year-over-year client retention rate
  • 47-pt Proprietary Experiment Readiness Audit

Median client lift: +27%. Full methodology and audit scoring rubric available under Method.

Browse by vertical

Find your category before you read a single line.

Experiments behave differently in fintech than in dev tools. The same test that lifts a PLG sign-up funnel can depress a sales-assisted onboarding. Pick the vertical closest to yours — the lift tables are not interpolated from another industry.

  1. 01 Fintech B2B treasury, cards, spend management · 14 engagements
  2. 02 Dev Tools PLG, IDE, infra · 11 engagements
  3. 03 Vertical SaaS Healthcare, construction, legal · 9 engagements
  4. 04 Subscription D2C, media, prosumer · 18 engagements

Case study · 01 · Fintech

Onboarding completion, +44% in 41 days.

A Series B spend-management platform — three-month engagement, two partner-led strategists, our 47-point Experiment Readiness Audit run on day four.

Strategist reviewing a printed A/B test result table beside a laptop.
Audit scorecard from week one — the team was running twelve tests but only three were powered correctly.
Metric Baseline Variant Lift Duration Confidence
KYC completion rate 61.2% 74.8% +22.2% rel. 14 days 99%
Bank-link → first card issued 38.4% 55.3% +44.0% rel. 27 days 97%
Time-to-first-card (median) 6d 4h 2d 19h −57% 41 days 95%
30-day activated accounts 29.1% 39.7% +36.4% rel. 41 days 99%

Sequential testing, Bonferroni-corrected across four co-primary metrics. Source: Tyrell Lab engagement ledger, 2023.

The client's team had been running onboarding tests for nine months but had never power-calculated any of them. Eleven of the thirteen "winning" variants were underpowered and one was a false positive that had been live in production for two quarters. Our first week was unlearning. The KYC reordering test alone, removing a single redundant identity check, moved bank-link completion by +22% relative.

What stayed was the playbook. Within six weeks the client's growth PM was running weekly sequential tests on her own, with our partners reviewing designs on a 30-minute call. By the end of month three the team had shipped 11 validated experiments, four of them winners, without a single engineering dispute over methodology. They have since expanded scope to pricing-page and billing-recovery experiments — the engagement is in its second quarter.

Book a Free Experiment Diagnostic →

They killed two tests of ours on day one because we weren't powered to run them. We were a little annoyed. Six weeks later we'd shipped our first genuine win and we had the calculator to prove it.

Head of Growth Series B spend-management platform · client since 2023

Case study · 02 · Dev Tools

Activation to second project, +28% in 33 days.

A PLG developer-infrastructure product — seven-week engagement, single partner embedded with the founding PM, focused on the second-session activation event that nobody had instrumented.

Handwritten power calculations next to a plotted lift curve.

The activation curve, week 1 vs. week 5. The inflection on day 12 is when the empty-state prompt shipped.

Metric Baseline Variant Lift Days Conf.
First project created 72.0% 74.1% +2.9% rel. 21 90%
Second project within 7 days 18.5% 23.7% +28.1% rel. 33 98%
Invite a teammate 11.2% 16.9% +50.9% rel. 33 99%
14-day retained seats 34.4% 41.6% +20.9% rel. 33 97%

Dev-tool activation is a two-session problem. Almost nobody measured the second session. We instrumented it, found that 81% of new accounts never returned after creating a first project — and that the entire onboarding surface had been optimised for the wrong event.

The empty-state prompt for a second project, designed by our partner and reviewed by the client's product designer, shipped in week three. It is, by any visual measure, the least interesting screen in the product. It moved the second-project rate by 28%. The team keeps it because it is boring and it works, and we left them with the framework to find the next boring screen that probably isn't.

Book a Free Experiment Diagnostic →