The Method · An essay in five chapters
A senior bench, a 47-point audit, and a playbook your team can run after we leave.
Most experimentation agencies sell hours and ship junior work. We do neither. The Tyrell Lab method is a small, evidence-led engagement built around three artifacts: the partners who design every test, the audit that decides which tests to run, and the playbook that lets your team keep shipping without us. This is a long read. It is also, deliberately, the most important page on the site.
- +32%avg. activation lift in first 90 days
- 412winning experiments shipped since 2019
- 47point Experiment Readiness Audit
- 14senior strategists, zero juniors
Chapter 01 · The Bench
Fourteen senior strategists. Every test designed by the partner who sold the engagement.
The single most expensive mistake a VP of Product can make in an experimentation engagement is to confuse the people in the sales deck with the people who will design the tests. In our experience, the gap between those two groups is where experiments die. A clever strategist pitches the work; a junior PM inherits it three weeks later; the test that ships is, at best, a polite echo of the one that was pitched.
Tyrell Lab was founded in 2019 by former experimentation leads at Booking.com and Stripe specifically to close that gap. Our bench is fourteen people — fourteen senior strategists, data scientists, and product designers — and every experiment we ship is designed by a partner with at least eight years in the discipline. We do not employ junior strategists. We do not have a “delivery team.” We do not subcontract the analysis. The person who walked you through the diagnostic is the person who will be in your Slack on day six.
This is not a hiring philosophy. It is a unit economics choice that constrains the rest of the business. A senior-only bench means we cannot scale to eighty retainers, cannot chase RFPs from Fortune 500 procurement teams, and cannot promise same-week kickoffs across six time zones. What we can do is keep the bench small enough that every engagement is a named relationship, and disciplined enough that the median client lifts activation by 27% within a quarter — and the top quartile by considerably more.
“The first thing I check when a CRO agency pitches us is who, by name, is actually going to design the test. Tyrell Lab was the only firm that sent me a partner's résumé, not a delivery-team org chart.”
If you are evaluating an experimentation partner, the question worth asking is not how many tests they have shipped, or what tooling stack they prefer, or even how their pricing scales. The question is: who, by name, will be sitting in our experiment review next Tuesday? Everything else on this page is downstream of that answer.
Chapter 02 · The Instrument
Inside the 47-point Experiment Readiness Audit.
The audit is the operational backbone of every Tyrell Lab engagement. It is a proprietary 47-point diagnostic, refined across 180+ SaaS onboarding engagements since 2019, that turns the vague mandate — “we should test more” — into a ranked backlog of evidence-driven experiments with predicted lift bands, required sample sizes, and an explicit sequencing rationale.
The audit has three layers. The first layer (twelve points) measures infrastructure readiness: event taxonomy hygiene, assignment integrity, sample-ratio-mismatch detection, guardrail coverage, and the boring but catastrophic question of whether your analytics platform can actually distinguish a winning test from a fluke. The second layer (nineteen points) measures funnel instrumentation: drop-off attribution across activation, retention surfaces, and the invisible transitions between them. The third layer (sixteen points) measures experiment literacy: how the team writes hypotheses, debates results, archives learnings, and avoids the seven recurring biases that quietly corrupt most in-house testing programs.
The output is a single document — typically sixteen to twenty-two pages — that becomes the engagement's working contract. Every experiment we ship in the first quarter traces back to a numbered audit finding. Every finding carries a confidence band, a recommended test design, and an honest note about where the data is too thin to make a call. We have shipped over 1,180 experiments against audit-derived backlogs and the pattern is consistent: teams that treat the audit as a living instrument, not a one-time report, compound their win rate quarter over quarter.
A worked example of the audit, anonymised and with the host company's permission, is available to qualified prospects after the first diagnostic call. We do not publish it publicly because, like most rigorous instruments, its value is in the calibration.
Figure 01 · The Audit, Rendered
A sample scorecard beside a representative 90-day lift curve.
Two artifacts from a Q3 2023 engagement with a Series B vertical SaaS company (NDA, names withheld). On the left, an excerpt of the 47-point audit with two under-served domains flagged for the first sprint. On the right, the cumulative activation rate across the 90-day engagement, with the eight shipped experiments annotated at their ship date.
Both figures are real and reproduced with permission. The company's name is withheld under the engagement NDA. The audit is the same instrument we run on every engagement, with the scoring rubric adapted to the host's vertical.
Chapter 03 · The Hand-off
What happens to your team after we leave.
The unspoken question in every experimentation engagement is the one nobody asks in the sales call: what do we actually own when you are gone? It is the right question, and the answer is the only thing that determines whether the engagement was worth the spend. A lift that evaporates twelve weeks after the partner's last standup is not a lift; it is a consulting receipt.
Tyrell Lab engagements are structured around an explicit 90-day internalization milestone. In the first thirty days, our partners design and ship experiments; your team shadows. In days thirty through sixty, your team leads the experiment design with our partners in the review seat. In days sixty through ninety, your team ships, reviews, and archives experiments independently, with our partners available only for the hardest calls. By day ninety, the engagement is a contract, not a dependency. Our 71% scope-expansion rate inside that same window is a downstream effect: teams that have internalized the method usually want more of it, not less, but they want it on their terms.
“The playbook alone was worth the engagement. We have run 34 experiments in the nine months since Tyrell handed off, with no drop in rigor and no lift in our internal headcount.”
The playbook itself is not a slide deck. It is a living document — typically 60 to 90 pages, customized to the host's stack and stage — that captures the audit findings, the shipped experiments, the learnings (including the failed ones, especially the failed ones), the governance cadence, and the explicit anti-patterns the team has agreed to avoid. It is the artifact your successor PM will read on day one. It is the reason your next experimentation hire does not start from zero. And it is the part of the engagement we take the most seriously, because it is the only part that has to outlast us.
Chapter 04 · The Numbers, Quietly
Four figures that summarize what the method actually delivers.
- +32% average activation lift across 410+ shipped experiments, measured at the 90-day mark. Median client: +27%.
- 412 winning experiments shipped since inception, against 1,180 total experiments run for clients.
- 92% client retention rate year-over-year. 71% of clients expand scope within 90 days of kickoff.
- 47 point Experiment Readiness Audit, the same instrument on every engagement, with vertical-specific calibration.
If these numbers describe a problem you are trying to solve, the next step is a thirty-minute diagnostic — not a sales call, a working session with a Tyrell Lab partner.
Book a Free Experiment Diagnostic