Experimentation
Structured experimentation where your traffic supports a conclusion — and an honest answer where it doesn’t. Tests designed around what a change is worth, not which version people prefer.
240+ tests for one brand37 tests across 15 funnels for another+49% conversion from a testing programme
The problem
Not because the variant lost — because the test could never have concluded anything. Too little traffic, too small a change, a metric nobody agreed on beforehand, or a hypothesis that was really just a preference wearing a lab coat.
A test is expensive in the only currency that matters at your stage: time. Two weeks spent proving a button colour doesn’t matter is two weeks not spent on the checkout step losing you real orders. Which is why the hard part isn’t running tests — it’s choosing which ones deserve to exist.
What the work involves
Testing works when it’s a system rather than a series of one-off experiments. That system has three parts.
Supplements & Wellness
A functional foods brand scaling across Europe and the US, on a site originally built for education rather than ecommerce. Testing became the mechanism for improving it without stopping to redesign.
Where this sits
This service isn’t sold on its own — it’s scoped into a program from your roadmap, at a fixed price agreed before anything starts.
The highest-confidence fixes shipped, with before-and-after measurement.
Structured testing on the changes worth proving, alongside the rebuild.
A continuous testing cycle where each month’s results set the next month’s plan.
Related services
Finding and fixing where a store loses revenue.
A structured review of the whole buying path.
Every step between wanting it and paying for it.
Recurring purchase that feels like a choice.
FAQ
Enough for a result to mean something, which depends on your conversion rate and the size of the change you’re testing. Rather than quote a threshold that sounds authoritative and misleads, we model it against your actual numbers and tell you plainly whether testing is the right tool for you yet.
Then we don’t test — we prioritise changes that are well-evidenced from behaviour data, heuristics, and eight years of seeing the same patterns. That’s a slower kind of certainty, but it beats running a test that can’t conclude and calling the result a decision.
Whatever fits your stack and budget, and we’ll tell you honestly when a cheaper option does the same job. The tool matters far less than hypothesis quality and clean implementation — a badly built test on expensive software is still a badly built test.
Badly implemented ones do, and that alone can invalidate the result — you end up measuring the flicker rather than the change. Implementation is treated as part of the test design, not an afterthought.
We report it, explain what we think it means, and use it to narrow the next hypothesis. A programme where every test wins isn’t a programme — it’s a marketing document.