What a split test has to prove before we call it

Sample size, the metric that decides, and the three ways a test lies to you.

Field note 02 - 7 min

Most of the split tests we inherit are not tests. They are two messages that went out on the same evening, to two lists that were never the same list, judged on whichever number came out higher. That is not evidence. That is a coin toss with a report attached.

A test earns the right to change what you send next month only if it survives being argued with. Here is what we make one prove, and what we do when it proves nothing.

Clicks are not what you are buying

The commonest fault in operator messaging tests is that the click decides. Clicks are easy to count and they move a lot, which makes them feel like a signal. A tap tells you the message was interesting. It does not tell you the offer was worth funding an account for.

Our blended figures are an 11.4% click rate and 4.2% of those clicks going on to deposit. The two rates move independently, often against each other. Widen the promise in the copy and the click rate climbs while the deposit rate falls, because the extra taps are curiosity rather than intent. That variant reads as a clear winner for six hours and then costs you a quarter of sends.

So the metric that decides is deposits per 1,000 delivered, or on VIP tiers, attributed deposit value per 1,000 delivered. Click rate stays on the report. It does not get a vote.

Split on a hash of the player id

Both arms come out of one segment, cut on something with nothing to do with player behaviour. We hash the player id and split on the remainder, hash(player_id) % 2. The halves are comparable by construction and stay stable if the test runs again next month.

What we see instead is variant A sent to one brand or market or recency band and variant B sent to another. Those are two similar segments, not two halves of one, and what you measure at the end is mostly the difference you started with.

Before the send we check the halves match on four things:

  • Count, inside a percent of each other
  • Median days since last deposit
  • Median lifetime deposit value
  • Market and language mix

If any of those disagree by more than a couple of percent, we re-cut. The check takes two minutes and it is the cheapest thing in this note.

Change one thing

Test the offer, or the send hour, or the first line. Not two of them. A test that changes the bonus and the timing tells you that pair beat the other pair, and you will never know which half did the work. Next time you reuse the wrong half.

Operators push back because it feels slow. Two clean single-variable tests a week apart give you two facts you can keep. A four-way test on a 20,000 player file gives you four arms of 5,000 and an argument.

Write down what decides it, before anything sends

One line in the brief, agreed by both sides, before a single message leaves. It names the deciding metric, the minimum arm size, the attribution window, and the result that would change your mind.

That last one matters more than it sounds. If nobody can say in advance what would make them drop the favourite, the test is theatre and the favourite was always going to win.

It is written rather than understood for one reason. Once the numbers land, everybody can find a metric that agrees with them. The person who wrote variant B will notice B pulled better on the thirty day cohort. He is not lying. He is doing what anyone does with a table of numbers and no rule agreed in advance.

A lucky evening cannot call a test

We will not call a test on less than 9,000 delivered per arm. At our blended rates that arm produces about forty deposits, enough that one player having a good Friday cannot decide it, and not much more than enough.

Below that the noise is bigger than any difference worth acting on. Two arms of 3,000 will show you a forty percent gap that reverses the next time you run the same two messages. If a segment is too small to split, we do not split it. We send the variant the library already supports, whole, and queue the test for a month when the file can answer it.

The three ways a test lies

Peeking early

Deposits are attributed while the campaign is running, so anyone with the report open can watch one arm lead at 9pm and call it there. Early deposits are the fastest players in the file, and the fastest player is not the average one. We read the arms live to stop a variant doing damage, nothing else. No winner is called until both arms have closed the window.

Halves that were never matched

Worth saying twice, because it is the lie that survives review. The arithmetic is right, the table is right, the winner is wrong, and nothing in the report shows it. Only the split does.

A segment that was already moving

The nastiest of the three. You test a win-back message against last month's numbers, it comes back thirty percent up, and the reason is a fixture weekend, a payday cycle or your own email push already pulling that file back. Compare a variant to a historical baseline and you credit the message with work the calendar did. Both arms go out in the same hours on the same day. The control has to be running now, not in March.

Sometimes the answer is nothing, and we say so

Roughly one test in three overturns the offer everyone expected to win. About one in five comes back flat: both arms inside the noise, nothing to roll out. That goes in the report in those words, with no rescue attempt.

What we will not do is hunt for a subgroup where our arm happened to win. Cut a flat test finely enough and you will always find a market, a tier or a weekday where one variant looks good. That is a search, not a finding, and a calendar built on it is how a programme drifts into sending things nobody proved.

What happens after a flat test is short:

  1. The incumbent stays. It was not beaten.
  2. The variable goes in the library as one that does not move this list.
  3. We stop testing it and spend the next send on something that might.

Knowing a lever does nothing on your file is worth money. It ends the same argument every quarter and puts the next test somewhere it can pay.

What to take from this

  • Name the deciding metric before the send, in writing
  • Split on a hash of the player id, never on two similar segments
  • An inconclusive test is a result, not a failure to report
Next: Field note 03Dormant is not dead: reading a lapsed segment

Want this run on your list?

Send us the market, the segment size and the offer you have in mind. You get the plan and the number back, not a discovery call.

Messages sent today

407,891