Why we are re-opening a question we thought was closed
Round 6 taught us to distrust a flat early-funnel metric and a lopsided sample. Turning that same lens backwards on our own record, the creative ranking we have treated as settled since Round 4 does not survive.
Round 4 was a three-way tie
It is the only round where the three finalists got roughly equal budgets. We reported that all three beat their own controls, which was true. We then quietly treated the order between them as meaningful. It never was.
| Comparison | Reached the picker | p | Verdict |
|---|---|---|---|
| honest-a vs instant-a | 53.7% vs 48.9% | 0.148 | tie |
| honest-a vs daily-a | 53.7% vs 50.2% | 0.296 | tie |
| daily-a vs instant-a | 50.2% vs 48.9% | 0.694 | tie |
Round 5's tiebreak was not ours to make
Round 5 put the three finalists head to head and left campaign budget optimisation on. That hands the split to the platform, and the platform optimises for its objective, which is clicks. Round 3 had already proved that clicks do not predict who plays.
| Ad | Players it brought | Share of the round |
|---|---|---|
| honest-a | 2,992 | 76.0% |
| daily-a | 582 | 14.8% |
| instant-a | 364 | 9.2% |
The winner got 8.2 times the sample of the ad that came last, and the two starved arms were plausibly stuck in the platform's opening learning phase for their entire run. We then declared a champion on that.
There is a third gap. Every creative round we have run was measured on the audience Round 6 just retired for being burnt out. The ranking has never once been measured on a fresh pool. Round 7 fixes all three problems at once, and it costs nothing extra to fix them, because the three ads already exist.
The three pitches
Unchanged since Round 4, byte for byte. Each one names a different barrier the player is being freed from.
"No energy. No paywall."
Play one run or ten. Free, in your browser.
"No app store. No account."
Tap the link. You are already in.
"Same run. Whole world. Today."
How far do you get? Free, no download.
The setup
One variable. Everything except the picture and the sentence on it is identical across all three.
Held constant
- The audience: keyword-only, 23 terms, no community list at all. This is the configuration Round 6 measured as the best per returning player, and the one with the most room left to grow.
- Budget optimisation: OFF. The single fix that makes this round mean something Round 5 could not.
- Global, automatic placements, lowest-cost bidding, daily budgets.
- Same landing page, same tracking scheme.
The numbers
- $13.54 per ad per day, three ads, $40.62 a day
- 32 days, July 31 to September 1
- $1,299.84 total, a hard cap
- About 5,500 players per ad if delivery matches Round 6
- Running on a brand new ad account, so early delivery will be cold
The month is not vanity. Two weeks resolves a 3.6 point gap in whether people start a run, which is enough to rank the ads but not enough to see whether they bring back different players. A month gets that second question inside reach for the first time.
The metric, declared before launch
Round 6's lesson was that the metric you choose decides the answer you get, so we are fixing ours in public first.
- Primary: cost per started run. Not click-through, not cost per engaged player. Round 3 showed click-through actively misleads, one creative tied for the best click rate in the round and finished last on everything else. Round 6 showed cost per engaged player is blind to quality, identical to three decimal places across audiences that differed twofold on returning players.
- Secondary: cost per returning player. Newly readable at this volume and the real reason for a month.
- No peeking, no reallocation, no pausing a losing arm for the full 32 days. That discipline is what made Rounds 3, 4 and 6 readable and its absence is what cost us Round 5.
- We stop the round early only if cost per player lands more than 40% above Round 6's, which would mean the auction has turned and the comparison is no longer about the ads.
Predictions, in writing, before launch
Same rule as last time. A test you cannot lose is not a test, so here is what we expect, in public, where next month's readout can mark it.
The scoreboard we will publish
Next journal entry, after September 1: this table filled in, plus a grade on each prediction above. Same as last round, and the round before.
| Ad | Landings | Started a run | $ / started run | Returned in 48h | $ / returner | Verdict |
|---|---|---|---|---|---|---|
| honest-a | — | — | — | — | — | pending |
| instant-a | — | — | — | — | — | pending |
| daily-a | — | — | — | — | — | pending |
One thing this round will not answer: whether a different pitch entirely would beat all three. We are re-running finalists, not searching. That search is a bigger question than a month of this budget can hold, and the honest reason to run this round first is that we have to know whether our current answer was ever real before we go looking for a better one.