Game development journal · Paid acquisition · Round 6 readout

The Audience That Clicked Most, Played Least

Last week we froze our best ad and pointed it at three different audiences. The one that has been seeing that ad for four rounds running produced the best click-through rate in the round and the worst players in it, by a factor of two. Here is the scoreboard we promised, a grade on all five predictions, and the reason "the art is worn out" turned out to be the wrong answer.

$366.99
spent
706,248
impressions
494,715
people reached
4,666
players landed
1.15
frequency, control
2 of 5
predictions wrong

The scoreboard we promised

Last week's entry ended with an empty table and a commitment to fill it in. Engagement rates are settled-days only, exactly as we said we would judge them. Landings and spend are the full round.

AudienceLandingsEngaged %Inert %$ / engagedReturned in 48hVerdict
r6-corethe proven list 1,48066.0%34.0%$0.125 1.80%18 / 1,001 retire
r6-genre13 fresh communities 1,59065.7%34.3%$0.126 3.42%39 / 1,142 keep
r6-broadkeywords only 1,59665.9%34.1%$0.125 2.78%32 / 1,150 keep and scale
Look at those middle four columns. Engagement rate, inert rate and cost per engaged player are identical across all three to within noise. The two metrics we pre-registered to judge this round on could not tell the audiences apart at all.
pairwise engaged%: genre vs core p=0.921 · broad vs core p=0.973 · genre vs broad p=0.948

Then look at the last numeric column. On the thing that actually matters, whether the person comes back, the proven audience returns at half the rate of the fresh ones. A flat metric is not a null result about the world. It is a null result about the metric.

What actually happened

Everything past the title screen separated hard, and it separated the same way every time.

AudienceStarted a runReached floor 2Won a runReturned in 48h
r6-genre41.1%22.1%7.5%3.42%
r6-broad41.9%23.2%7.5%2.78%
r6-core39.7%20.3%4.3%1.80%
fresh pools pooled vs the control  ·  won a run: +3.24pp, p=0.00003  ·  returned in 48h: +1.30pp, p=0.034

It is not a phone problem

The obvious objection is that the proven list simply drew cheaper devices. It did skew that way: 82% Android against 77%, and double the share on slow networks. The gap survives inside each group anyway.

AudienceAndroid: wonAndroid: stayediOS: woniOS: stayed
r6-genre7.4%4.08%6.1%2.86%
r6-broad7.4%3.26%6.4%5.18%
r6-core4.4%1.98%3.5%2.33%

A two-point difference in slow-network share cannot produce a 70% relative gap in win rate.

Grading the five predictions

A test you cannot lose is not a test. We wrote these down before spending a dollar, so here they are marked in public. One right, two half right, two wrong, and one of the wrong ones is the finding of the round.

P1

We said: the control declines. Cost per engaged player comes in worse than round 5's $0.11, with click-through continuing its measured 10 to 15% per-round decay.

half right  Cost per engaged did worsen, $0.11 to $0.125, up 14%. But click-through did not decay. It came in at 0.885%, the highest in the round, and 16.4% above the two audiences seeing the ad for the first time. Depletion turned out to be real and worse than we predicted, and it does not show up in clicks at all.

P2

We said: the fresh genre pool posts the best engagement rate, highest engaged share and lowest inert share of the three.

wrong  Genre posted the lowest engaged share of the three. More to the point, the metric was flat across every pair (p above 0.92 everywhere), so it was never going to rank anything. We picked a rung too early in the funnel to see audience quality.

P3

We said: broad buys the cheapest clicks and the most bounces. Open question we refused to predict: whether its cost per engaged player still lands near the control.

wrong twice, right where it counted  Cost per click was identical everywhere ($0.065 / $0.066 / $0.065), and broad did not have the highest inert share, genre did. But the open question resolved decisively: broad's cost per engaged player came in at $0.125, the same as the control to three decimals, while reaching 205,297 people on keywords alone at a frequency of 1.21. Our addressable pool is not 1.6 million people. It is much larger, and we still do not know by how much.

P4

We said: retention will not care which audience we buy. Next-day return lands between 3 and 4% for all three. Ads move the top of the funnel enormously and the bottom not at all.

wrong, and this is the round  Retention cared a great deal. The fresh pools returned at 3.42% and 2.78%; the proven list at 1.80%, roughly half. Fresh against control is +1.30pp at p=0.034. Five rounds had taught us that creative does not move retention, which is still true. We over-generalised it to audience, and audience moves it a lot.

P5

We said: the first three days will look worse than the truth, because fresh ad groups spend their opening days in the platform's learning phase.

right  Engagement on settled days ran 2 to 4 points above the full-round figure on every arm: core 63.8% to 66.0%, genre 62.0% to 65.7%, broad 62.3% to 65.9%. Worth keeping as a standing rule rather than a per-round note.

Was it just that the art is worn out?

This was the first thing everyone asked, including us. It is the right instinct and it is the wrong answer, and this round happens to be a clean test of it.

Creative fatigue makes two predictions. The audience that has seen the art most should show the highest frequency and the lowest click-through. Both came back inverted.

AudiencePrior exposureFrequencySaw it onceCTR
r6-core4 rounds1.1585.3%0.885%
r6-genrenever1.2079.7%0.753%
r6-broadnever1.2178.9%0.768%
The honest ad: a runed stone door with its chains snapped and a padlock falling away
The one creative, byte-identical across all three audiences for the whole round.
The veterans saw the art less often than the newcomers did, and clicked it 16.4% more. z = +5.35, p = 8.6e-8, across half a million impressions.

At a frequency of 1.15, 85% of that audience saw the ad exactly once all week. Repetition fatigue needs repetition, and there was none to be had. Round 4 measured the same thing independently: per-ad frequency between 1.10 and 1.26, and only about 3.4% of an audience had ever seen that exact creative before.

Round-level click-through is not falling either. R3 blended 0.775%, R4 0.746%, and this round 0.797%, the highest we have run, on the most impressions we have bought.

The version that does survive

Call it recognition without novelty. Someone who has scrolled past this ad across four rounds still recognises it and still taps, but they formed a verdict long ago and the tap is reflex rather than intent. That predicts high click-through and worthless clicks, which is exactly the control's fingerprint: best click-through, worst click-to-landing survival (81.0% against 84.1% and 83.6%), worst everything below. It is still an audience problem, not an art problem, and the fix is still a different audience.

Which is why this distinction was worth a round of budget. If it were fatigue, new art fixes it cheaply. If it is depletion, new art fixes nothing and you have to go find new people.

What each audience actually cost

AudienceSpend$ / 1k reached$ / engaged$ / win$ / returner
r6-genre$124.07$0.595$0.126$1.04$2.00
r6-broad$124.62$0.607$0.125$1.04$2.15
r6-core$118.30$0.657$0.125$1.88$3.94

Identical at the top of the funnel. Nearly double at the bottom. The control also costs 8 to 10% more per human reached, because the auction charges a premium for a pool that is running out. Per 100,000 impressions it buys 15% more engaged players and 40% fewer returning ones.

One honest caveat on all of this: the retention verdict rests on 18, 39 and 32 returning players. The direction is solid and the significance holds, but these are small numbers and we would rather say so than dress them up.

What Round 7 does about it

Six rounds have now closed every acquisition question we can afford to ask, with one exception we found while writing this up.

  • The proven list is retired. It still clicks, which is exactly the trap, and it costs roughly double per returning player.
  • We judge on cost per returner from now on, not cost per engaged. This round is a clean demonstration that the early metric is blind to audience quality.
  • Read the reach and frequency columns before diagnosing anything. Two numbers separated fatigue from depletion here, and they are two clicks away in the ad manager. We ran five rounds without them.
  • One question is still open, and it is our own fault. Going back through the record with fresh eyes, the creative ranking we have treated as settled since Round 4 was never significant: the three finalists finished in a statistical tie (p = 0.148, 0.296, 0.694), and Round 5's tiebreak was decided by the platform's budget optimiser, which handed one ad 76% of the round on the basis of click-through. So Round 7 re-runs those three creatives on one fresh audience, with equal budgets we control, for a month.
  • None of this is the real constraint. Our best audience still returns 3.4% at 48 hours. No audience fixes that number. The next work is the home screen and the reason to come back, and that happens in the game, not in the ad manager.
Play Atlas of Doors free in your browser No download, no account, no energy bars. The ads are honest because the game can afford them to be.

← Back to the forum