What a split test has to prove before we call it
Sample size, the metric that decides, and the three ways a test lies to you.
Most of the split tests we inherit are not tests. They are two messages that went out on the same evening, to two lists that were never the same list, judged on whichever number came out higher. That is not evidence. That is a coin toss with a report attached.
A test earns the right to change what you send next month only if it survives being argued with. Here is what we make one prove, and what we do when it proves nothing.
Clicks are not what you are buying
The commonest fault in operator messaging tests is that the click decides. Clicks are easy to count and they move a lot, which makes them feel like a signal. A tap tells you the message was interesting. It does not tell you the offer was worth funding an account for.
Our blended figures are an 11.4% click rate and 4.2% of those clicks going on to deposit. The two rates move independently, often against each other. Widen the promise in the copy and the click rate climbs while the deposit rate falls, because the extra taps are curiosity rather than intent. That variant reads as a clear winner for six hours and then costs you a quarter of sends.
So the metric that decides is deposits per 1,000 delivered, or on VIP tiers, attributed deposit value per 1,000 delivered. Click rate stays on the report. It does not get a vote.
Split on a hash of the player id
Both arms come out of one segment, cut on something with nothing to do with player behaviour. We hash the player id and split on the remainder, hash(player_id) % 2. The halves are comparable by construction and stay stable if the test runs again next month.
What we see instead is variant A sent to one brand or market or recency band and variant B sent to another. Those are two similar segments, not two halves of one, and what you measure at the end is mostly the difference you started with.
Before the send we check the halves match on four things:
- Count, inside a percent of each other
- Median days since last deposit
- Median lifetime deposit value
- Market and language mix
If any of those disagree by more than a couple of percent, we re-cut. The check takes two minutes and it is the cheapest thing in this note.
Change one thing
Test the offer, or the send hour, or the first line. Not two of them. A test that changes the bonus and the timing tells you that pair beat the other pair, and you will never know which half did the work. Next time you reuse the wrong half.
Operators push back because it feels slow. Two clean single-variable tests a week apart give you two facts you can keep. A four-way test on a 20,000 player file gives you four arms of 5,000 and an argument.
Write down what decides it, before anything sends
One line in the brief, agreed by both sides, before a single message leaves. It names the deciding metric, the minimum arm size, the attribution window, and the result that would change your mind.
That last one matters more than it sounds. If nobody can say in advance what would make them drop the favourite, the test is theatre and the favourite was always going to win.
It is written rather than understood for one reason. Once the numbers land, everybody can find a metric that agrees with them. The person who wrote variant B will notice B pulled better on the thirty day cohort. He is not lying. He is doing what anyone does with a table of numbers and no rule agreed in advance.
A lucky evening cannot call a test
We will not call a test on less than 9,000 delivered per arm. At our blended rates that arm produces about forty deposits, enough that one player having a good Friday cannot decide it, and not much more than enough.
Below that the noise is bigger than any difference worth acting on. Two arms of 3,000 will show you a forty percent gap that reverses the next time you run the same two messages. If a segment is too small to split, we do not split it. We send the variant the library already supports, whole, and queue the test for a month when the file can answer it.
The three ways a test lies
Peeking early
Deposits are attributed while the campaign is running, so anyone with the report open can watch one arm lead at 9pm and call it there. Early deposits are the fastest players in the file, and the fastest player is not the average one. We read the arms live to stop a variant doing damage, nothing else. No winner is called until both arms have closed the window.
Halves that were never matched
Worth saying twice, because it is the lie that survives review. The arithmetic is right, the table is right, the winner is wrong, and nothing in the report shows it. Only the split does.
A segment that was already moving
The nastiest of the three. You test a win-back message against last month's numbers, it comes back thirty percent up, and the reason is a fixture weekend, a payday cycle or your own email push already pulling that file back. Compare a variant to a historical baseline and you credit the message with work the calendar did. Both arms go out in the same hours on the same day. The control has to be running now, not in March.
Sometimes the answer is nothing, and we say so
Roughly one test in three overturns the offer everyone expected to win. About one in five comes back flat: both arms inside the noise, nothing to roll out. That goes in the report in those words, with no rescue attempt.
What we will not do is hunt for a subgroup where our arm happened to win. Cut a flat test finely enough and you will always find a market, a tier or a weekday where one variant looks good. That is a search, not a finding, and a calendar built on it is how a programme drifts into sending things nobody proved.
What happens after a flat test is short:
- The incumbent stays. It was not beaten.
- The variable goes in the library as one that does not move this list.
- We stop testing it and spend the next send on something that might.
Knowing a lever does nothing on your file is worth money. It ends the same argument every quarter and puts the next test somewhere it can pay.
What to take from this
- Name the deciding metric before the send, in writing
- Split on a hash of the player id, never on two similar segments
- An inconclusive test is a result, not a failure to report
Want this run on your list?
Send us the market, the segment size and the offer you have in mind. You get the plan and the number back, not a discovery call.