Paid Acquisition · Reddit · Round 4 “refinement” · 3×3 hook test

Round 4 Readout & the R3 Retention Verdict

Two questions, one report. First the one we promised to answer on a fixed date: now that Round 3 has aged a full week, which ads actually brought players back? Then Round 4: which of the nine hooks won, which three we carry, and — the part that matters more — what this round says about whether Reddit is a channel worth scaling at all.

$350.01
R4 spend
9
Ads (3×3)
599,463
Impressions
4,473
Clicks
0.746%
Blended CTR
3,435
Players landed
1,420
Reached picker
Jul 10–17
~7 days

TL;DR

  • The R3 retention verdict is in, and it is a flat no. Matured D7 across all of Round 3 is 1.14% (48 of 4,206 players), statistically identical to the old benchmark campaign's 0.99% (p=0.663). The hooks did lift D1 for real (3.90% vs 2.33%, p=0.012) — and that lift completely washed out by day 7.
  • No R3 ad's D7 is distinguishable from any other's. χ²=5.4, df=9, p=0.80. The per-ad D7 column on the dashboard is noise — and it proves it by re-ranking itself when you change the definition of "came back" (codex goes from 10th to 1st). Do not pick ads on it.
  • The one retention question the money can answer: "stayed ≥2 days." There the ads genuinely differ (p=0.030), and the best-retaining hooks are daily (12.5%) and instant (12.2%) — the same two the R3 report carried forward. cards is last at 5.0%.
  • Round 4's three to carry, on the pre-declared metric ($ per player who reaches the path-picker): honest-a · daily-a · instant-a. All three "-a" variants top their lanes. honest-a is the round's best ad at $0.134/picker — 46% cheaper than the round's blended $0.246.
  • The obvious objection — “the controls had already been seen, so of course they lost” — was tested against Reddit's reach data, and it does not hold. Only 27% of R4's audience had ever seen an R3 ad (73% were brand new), per-ad frequency is ~1.2 (almost nobody saw anything twice), and only ~3.4% of a control's viewers had seen that exact creative. Worst-case, prior exposure drags a control by ≤2.2pp — against declines of up to 23pp. Subtract it and honest-a (+16.8pp) and daily-a (+21.5pp) still win; instant-a stays a tie (-0.9pp).
  • ⚠ The real reason to stop: we have burned 46% of the addressable audience. The overlap arithmetic puts the effective pool at ~1,572,564, and two rounds have already reached 719,614 of it. R3 creamed off the most responsive 423,125; R4 had to work the remainder — which is why the identical creative gets worse without anyone being tired of it. This, not creative, is the ceiling on Reddit.
  • The clean result — the only comparison no confound touches — is a vs b. Both siblings are new, launched the same day, never seen. The "-a" pattern (name what the player is freed to do, right now) beats the "-b" pattern in 2 of 3 independent lanes. That replication, not any "+20pp vs control" number, is what Round 4 actually taught us.
  • R4's best ad is still worse than R3's best ad was — and note this compares two different ads: r3-instant (last round's champion) bought pickers at $0.088; our new champion r4-honest-a costs $0.134 — 52% more. Better creative, worse economics, because the auction got dearer (eCPM +12%).
  • Round 4 missed every efficiency target the R3 report set. Cost per player went up 22% ($0.083 → $0.102), volume fell 18%, blended CTR did not move. The media got more expensive; the creative got better. Those two cancelled out.
  • The number that should drive the next decision: $7.29 per day-7 player. $350 of Round 3 bought 4,209 players and 48 who came back a week later. No hook changes that. Round 4 should be the last wide creative round — exactly as the R3 report predicted.
Part ① · the promised fixed-date read

Round 3's retention, now that it has actually aged

The R3 report shipped on Jul 10 with a dated promise: “re-read R3's real D7 on/after Jul 16 — this is the single number that says whether the winners actually retain.” That date has passed. R3 landed players Jul 3–10, so every cohort has now had its full 7 days. Here is the answer.

The headline: better ads bought better first days, and zero extra second weeks.

Round 3's blended D1 is 3.90% vs the benchmark's 2.33% — a real, significant improvement (z=2.50, p=0.012). Round 3's blended D7 is 1.14% vs the benchmark's 0.99% — no difference whatsoever (z=0.44, p=0.663). The R3 report guessed that a 3× D1 lift might carry through to a D7 of 1.5–3%. It did not. Everything the hook bought was gone within a week.

Which ads gave the best 7-day retention?

Every R3 ad, ranked by matured exact-day D7. Fractions are the real numerator/denominator — the dashboard rounds these to whole percents, which is exactly what makes them look more meaningful than they are. D7 = active on exactly day 7. D7+ = active on any day from 7 onward (same cohort, looser rule). Stayed = seen on ≥2 distinct days.

AdPlayersPickerD7 exact95% CID7+ aliveStayed ≥2dD1
r3-daily 408 52.0% 2.0%8/406 1.0–3.8 3.0%12/406 12.5%51/408 6.4%26/408
r3-combat 343 25.4% 1.8%6/342 0.8–3.8 3.5%12/342 8.7%30/343 3.5%12/343
r3-doors 402 31.3% 1.2%5/402 0.5–2.9 1.7%7/402 10.0%40/402 4.2%17/402
r3-instant 598 61.4% 1.2%7/598 0.6–2.4 3.3%20/598 12.2%73/598 4.5%27/598
r3-boss 385 16.6% 1.0%4/385 0.4–2.6 2.3%9/385 9.6%37/385 2.3%9/385
r3-dev 485 46.6% 1.0%5/485 0.4–2.4 2.5%12/485 10.5%51/485 3.1%15/485
r3-honest 392 41.6% 1.0%4/392 0.4–2.6 1.0%4/392 9.7%38/392 5.6%22/392
r3-cards 377 16.7% 0.8%3/377 0.3–2.3 1.6%6/377 5.0%19/377 1.6%6/377
r3-time 407 50.9% 0.7%3/407 0.3–2.1 2.5%10/407 11.1%45/407 3.9%16/407
r3-codex 412 37.4% 0.7%3/412 0.2–2.1 3.9%16/412 11.9%49/412 3.4%14/412
climax_curiositymatured benchmark1,11424.1%1.0%11/11140.6–1.84.6%51/111411.0%123/11142.3%26/1114
R3 pooled4,2091.14%48/4,2060.9–1.52.57%108/4,20610.29%433/4,2093.90%
Read the fractions, not the percents

The whole D7 column is built on 3 to 8 people per ad. One player either side moves an ad several rank positions. A χ² homogeneity test across all ten ads returns χ²=5.41, df=9, p=0.799 — that is as close to "these ten ads are the same ad" as data gets. Even best-vs-worst (daily 8/406 vs codex 3/412) is z=1.54, p=0.123 — not significant before you even account for the 45 pairwise comparisons on the table.

The proof that it's noise: the ranking won't hold still

Same players, same week, three reasonable definitions of “came back.” If the D7 column carried signal, an ad's rank would barely move between them. Instead:

AdRank · D7 exactRank · D7+ aliveRank · stayed ≥2dSwing
r3-daily141±3
r3-combat229±7
r3-doors386±5
r3-instant432±2
r3-boss578±3
r3-dev655±1
r3-honest7107±3
r3-cards8910±2
r3-time964±5
r3-codex1013±9

codex is dead last on exact-day D7 and first on D7+. combat is 2nd on D7 and 9th on stayed. time is 9th on D7 and 4th on stayed. A metric that reshuffles its own leaderboard when you nudge the definition by one day is not measuring the ad.

The one retention read this budget can actually support

“Stayed ≥2 days” uses every player as its denominator instead of the handful who returned on one specific calendar day — so it has ~10× the events behind it, and it separates: χ²=18.51, p=0.0296. Best vs worst is daily vs cards: z=3.66, p=0.0002 — real, and it survives the multiple-comparison threshold.

r3-daily
12.5% · 51/408
r3-instant
12.2% · 73/598
r3-codex
11.9% · 49/412
r3-time
11.1% · 45/407
r3-dev
10.5% · 51/485
r3-doors
10.0% · 40/402
r3-honest
9.7% · 38/392
r3-boss
9.6% · 37/385
r3-combat
8.7% · 30/343
r3-cards
5.0% · 19/377
So: which R3 ads retained best?

daily and instant — and they are the only honest answer available. They are 1st and 2nd on the one powered metric (stayed ≥2 days: 12.5% / 12.2%), and top-4 on both D7 definitions, which no other ad manages. cards is the only clear loser (5.0%, less than half the leaders).

The caveat that matters more than the ranking: the gap between the best and worst ad's retention is worth ~7pp of a ≥2-day return, and nothing at all at day 7. Both R3 winners we carried forward (daily, instant) are confirmed as the right carry — but they were the right carry for CTR and completion reasons, and retention did not add information.

What it would actually cost to answer “which ad retains best”

Rather than keep staring at a column that cannot answer the question, here is the price of making it answerable. Using R3's own cost per player ($0.083) and the D7 gap we actually observed (2.0% vs 0.7%):

Compare just two ads on D7

n = 1,351 players per arm (80% power, α=0.05)
→ 2,702 players → $225

Affordable. But it answers exactly one question: are these two specific ads different?

Rank ten ads on D7

45 pairwise comparisons → α=0.05/45
n = 2,895 per arm × 10 ads
→ 28,950 players → $2407

6.9× the entire Round-3 budget, spent to rank creatives on a metric whose total spread is about one percentage point.

The decision this forces

Per-ad D7 is not "a number we don't have yet." It is a number this business model cannot afford to buy, and would not profit from owning: even a perfect ranking would let us pick between a 2% and a 0.7% D7 — on a base of ~4,000 players, that is the difference between ~48 and ~28 returning players per $320. The D7 problem is not on the ad account. It is in the game.

Part ② · the round that just finished

Round 4 — the 3×3 hook test

Three concepts that R3 proved (access / habit / trust) × three hooks each: the reigning R3 champion unchanged as a control, plus two fresh hooks. $350.01 split ~evenly ($38.34–$39.43 per ad), same audience, same landing, one variable: the hook.

Reddit Ads Manager, R4 ad groups
Reddit Ads · ad-group level, final billed. 599,463 impressions · 4,473 clicks · 0.746% CTR · $350.01.

The results, on the metric we pre-declared

The R3 report committed in advance to one win metric, precisely so we couldn't cherry-pick after the fact: cost per player who reaches the path-picker — CTR × click-quality, the metric that survives the “boss” clickbait trap. Sorted by it. Picker % is measured from path_picker_shown, instrumented since Jun 29, so it is comparable across both rounds.

AdImpr.CTRClicksPlayersReached pickerCompl.$/player$/picker ★
r4-honest-a r4-honest-a 66,471 0.936% 622 533 53.7%286/533 18.0% $0.072 $0.134
r4-daily-a r4-daily-a 63,147 0.820% 518 414 50.2%208/414 18.8% $0.094 $0.187
r4-instant-a r4-instant-a 58,677 0.825% 484 393 48.9%192/393 17.0% $0.100 $0.204
r4-instant-ctrl r4-instant-ctrlctrl 62,825 0.767% 482 368 42.9%158/368 17.9% $0.104 $0.243
r4-daily-b r4-daily-b 62,029 0.830% 515 348 36.5%127/348 14.1% $0.113 $0.310
r4-honest-ctrl r4-honest-ctrlctrl 73,048 0.661% 483 361 34.6%125/361 13.3% $0.107 $0.310
r4-honest-b r4-honest-b 70,480 0.675% 476 357 33.1%118/357 9.0% $0.110 $0.333
r4-instant-b r4-instant-b 69,199 0.652% 451 338 33.1%112/338 8.6% $0.117 $0.352
r4-daily-ctrl r4-daily-ctrlctrl 73,587 0.601% 442 323 29.1%94/323 7.7% $0.119 $0.408
Round total599,4630.746%4,4733,43541.3%1,420/3,435$0.102$0.246
The uncomfortable comparison: our new champion is worse than our old one

The pre-declared metric lets us price both rounds' champions on the same ruler. It does not flatter us:

instant r3-instant $0.088/picker → r4-instant-a $0.204/picker +132%
daily r3-daily $0.151/picker → r4-daily-a $0.187/picker +23%
honest r3-honest $0.195/picker → r4-honest-a $0.134/picker -31%

ROUND BEST  r3-instant $0.088 → r4-honest-a $0.134 +52%

Only honest genuinely improved (-31%). instant got twice as expensive per picker. And Round 4's best ad is +52% dearer than Round 3's best ad was. Part of that is real (R4 ate a learning phase R3's numbers never paid for; the R3 figures are reconstructed from that report's own per-ad cost-per-player) — but the bulk of it is simply that the auction repriced us 12% higher between the two rounds. This is the single strongest argument in this report against a Round 5: we spent $350.01 writing better hooks and still ended up buying pickers more expensively than the week before.

What the test was worth

The round paid the blended rate of $0.246 per picker and bought 1,420. But it found a rate: at r4-honest-a's $0.134, the same $350.01 would have bought 2,611 pickers — 84% more. That delta, repeated on every future dollar, is the entire return on this round.

The trap in this round's raw numbers: the learning phase

Before ranking anything, one finding that nearly broke this analysis. All nine ad groups launched on Jul 10 — straight into Reddit's delivery learning phase, where a fresh ad group is shown to a broad, unoptimised audience while the platform figures out who to target. The evidence is unmistakable, because we ran the same creative in both rounds:

r4-instant-ctrl · identical creative · picker rate by landing day
07-11: 26% · 07-12: 14% · 07-13: 49% · 07-14: 42% · 07-15: 65% · 07-16: 54%

That is not an ad getting better. That is Reddit learning who to show it to. The first three days of Round 4 bought materially worse traffic than the last four, for every ad:

AdPicker · learning (Jul 10–12)Picker · settled (Jul 13–17)Δ
r4-honest-a 50.0%106/212 56.1%180/321 +6.1pp
r4-daily-ctrl 29.3%41/140 29.0%53/183 -0.3pp
r4-instant-a 38.7%48/124 53.4%143/268 +14.6pp
r4-honest-b 28.0%42/150 36.7%76/207 +8.7pp
r4-daily-a 47.0%78/166 52.2%129/247 +5.2pp
r4-daily-b 24.8%34/137 44.1%93/211 +19.3pp
r4-instant-ctrl 18.8%19/101 52.1%139/267 +33.2pp
r4-instant-b 39.4%52/132 29.1%60/206 -10.3pp
r4-honest-ctrl 30.7%51/166 37.9%74/195 +7.2pp
Why this does not invalidate the round — but does change how we read it

The ranking is safe: all nine ads launched the same day and shared the same learning phase, so it hits every arm equally. The absolute rates are not: Round 4's blended picker rate is dragged down by its own first three days, which is a large part of why this round looks worse than R3 on every efficiency metric.

So every winner below is tested twice: once on all traffic, and once on settled traffic only (Jul 13–17). A hook only counts as a winner if it wins both. All three do — see the per-lane cards.

The obvious objection — and the data that settles it

Only the three controls had already been shown to these communities (for a full week, in R3). The six new hooks were unseen. So “beats its control” could easily mean nothing more than “it's a picture they haven't scrolled past yet” — and that bias runs one way, always favouring the challenger. If it were large, the pre-declared win rule would be structurally rigged toward replacing incumbents and most of this report would be worthless. So we measured it.

Measured: prior exposure is real, and it is far too small to matter

Telemetry can never answer this (we see clicks, never the impressions that were ignored). Reddit's Reach and Frequency can — because Reddit de-duplicates reach, so inclusion-exclusion on the account totals yields the cross-round overlap directly:

R3 reach 423,125 + R4 reach 382,533 + climax 16,883 = 822,541 de-duplicated union (account totals, no filter) = 719,614 OVERLAP = 102,927 people saw more than one campaign ⇒ only 26.9% of R4's audience had EVER seen an R3 ad. 73.1% were brand new. ⇒ per-ad FREQUENCY: R3 1.14–1.26 · R4 1.1–1.21 — almost nobody saw any ad twice.

Two facts kill the worry outright. Frequency ≈ 1.2 means there is barely any repetition to fatigue anyone with. And 73% of R4's audience had never seen us at all.

The bias, per lane, as an absolute ceiling

Sharpen it: how many of a control's R4 viewers had seen that exact creative before? Treating the two rounds' draws as independent from a pool of size P, the observed overlap implies P ≈ 1,572,564, and the per-creative overlap follows. The ceiling assumes every re-exposed viewer converts at 0% — nobody is that dead, so the true drag is smaller still.

LaneR3 ad reachControl reachSaw BOTH= % of control's audienceCeiling dragObserved dropExplains
instant 55,670 53,304 1,887 3.5% ≤2.2pp 9.3pp 23%
daily 52,756 65,025 2,181 3.4% ≤1.7pp 23.0pp 8%
honest 52,005 62,444 2,065 3.3% ≤1.4pp 3.6pp 38%

Only ~3.4% of a control's viewers had ever seen that creative. Prior exposure can account for at most 1.4–2.2pp of declines that run to 23pp — it explains 8–38% of the smallest drop and almost none of the largest. Conservative by construction: if Reddit re-targets the same high-propensity users, real overlap exceeds random for a given P, so solving this way under-estimates P and therefore over-states every ceiling above.

So: subtract the ceiling and re-score

New hookvs its control
settled traffic
− exposure ceiling
worst case
= adjusted liftvs its sibling (a vs b)
CLEAN — both unseen
Winner?
r4-instant-a +1.3ppp=0.764 −2.2pp -0.9pp +24.2ppp=1.3×10⁻7 Tie
r4-daily-a +23.3ppp=1.4×10⁻6 −1.7pp +21.5pp +8.2ppp=0.082 Winner
r4-honest-a +18.1ppp=6.5×10⁻5 −1.4pp +16.8pp +19.4ppp=1.4×10⁻5 Winner
Verdict: the winners stand

honest-a (+16.8pp adjusted) and daily-a (+21.5pp adjusted) both survive the worst-case correction with room to spare. instant-a (-0.9pp) remains what it was before the correction: a tie with its incumbent.

Recorded for honesty: before this data existed, we bracketed the bias by assuming 100% of each control's decline was prior exposure, and scored each hook against the champion's own R3 rate. That bound read instant -8.0pp · daily +0.3pp · honest +14.5pp — i.e. it would have thrown out daily-a as a tie and instant-a as a loss. The reach data refutes that bound. The controls did decline, but not because anyone had seen them.

And a bonus: a vs b is clean regardless

Both siblings are new, both launched Jul 10, neither ever seen. No exposure asymmetry, no incumbency, identical learning phase — this test needs no correction at all:

instant a 53.4% vs b 29.1% +24.2pp p=1.3×10⁻7 a wins
daily a 52.2% vs b 44.1% +8.2pp p=0.082 tie
honest a 56.1% vs b 36.7% +19.4pp p=1.4×10⁻5 a wins

The “-a” copy pattern beats the “-b” pattern in 2 of 3 independent lanes (daily's siblings tie on settled traffic). That replication is the round's most transferable lesson.

The finding that actually should stop the next round

The reach data was pulled to answer a narrow question, and answered a much bigger one on the way. If the controls' decline is not prior exposure and not the product and not the channel — what is it?

We are running out of Reddit, not running out of ideas

The overlap arithmetic implies an effective addressable pool of ~1,572,564 people. Across both rounds we have now reached 719,614 of them — 46% of the entire pool.

That reframes everything. R3 didn't just run first — it creamed off the most responsive 423,125 people in the pool. R4 then had to find ~279,606 new people, and those come from the less-responsive remainder. The same creative gets worse not because anyone is tired of it, but because the people left are harder to move. That fits every fact we have: frequency ≈ 1.2 (no repetition), 73% new faces, a step down at the round boundary rather than a slope, and a rising eCPM.

Why it does not bias the winners: pool depletion hits all nine R4 ads identically — they all draw from the same depleted remainder. The within-round ranking is fair. But it does put a hard ceiling on scaling: we are 46% through this audience after $700.01, and every further dollar buys from what is left. A Round 5 would not just re-buy a known lesson — it would buy it from the worst half of the pool.

The three winners

Ranked within each lane against its own control — the R3 champion, still running unchanged. The pre-declared rule: a new hook only wins if it beats its incumbent, not just its two siblings.

INSTANT — Zero-friction access

r4-instant-ctrl
r4-instant-ctrl The base to beat
“You are one tap from a run. A real roguelike deckbuilder. No download, no account, free.”
Discovery — the door pull
42.9%picker
$0.243/picker
0.767%CTR
r4-instant-a
r4-instant-a Tied with its control
“You are one tap from a run. No app store, no account.”
Immediacy — play right now
48.9%picker
$0.204/picker
0.825%CTR
vs control · all traffic: +5.9pp · z=1.64 · p=0.102
vs control · settled only: +1.3pp · p=0.764
r4-instant-b
r4-instant-b Loses to its control
“From a link to floor one in ten seconds. No install, free.”
Discovery — what is behind the door
33.1%picker
$0.352/picker
0.652%CTR
vs control · all traffic: -9.8pp · z=-2.68 · p=0.007
vs control · settled only: -22.9pp · p=5.5×10⁻7

DAILY — Shared daily ritual

r4-daily-ctrl
r4-daily-ctrl The base to beat
“Today’s door is the same for everyone. One daily run, one shared seed. Free, no download.”
Competition — shared ritual
29.1%picker
$0.408/picker
0.601%CTR
r4-daily-a
r4-daily-a Beats its control
“Everyone plays the same run today. How far do you get?”
Competition — the shared climb
50.2%picker
$0.187/picker
0.820%CTR
vs control · all traffic: +21.1pp · z=5.79 · p=7.1×10⁻9
vs control · settled only: +23.3pp · p=1.4×10⁻6
r4-daily-b
r4-daily-b Beats its control
“A new door opens every day. Come back for tomorrow’s run. Free.”
Community / ritual — a daily habit
36.5%picker
$0.310/picker
0.830%CTR
vs control · all traffic: +7.4pp · z=2.04 · p=0.042
vs control · settled only: +15.1pp · p=0.002

HONEST — Anti-F2P / trust

r4-honest-ctrl
r4-honest-ctrl The base to beat
“Build a deck. Open doors. Die. Go again. No energy timers, no gacha, no ads. Free in your browser.”
A real game
34.6%picker
$0.310/picker
0.661%CTR
r4-honest-a
r4-honest-a Beats its control
“No energy bars. No lives to refill. Play as much as you want, free.”
Unlimited play
53.7%picker
$0.134/picker
0.936%CTR
vs control · all traffic: +19.0pp · z=5.60 · p=2.1×10⁻8
vs control · settled only: +18.1pp · p=6.5×10⁻5
r4-honest-b
r4-honest-b Tied with its control
“No gacha. No ads. A real roguelike deckbuilder, free in your browser.”
Strategy — a real deckbuilder
33.1%picker
$0.333/picker
0.675%CTR
vs control · all traffic: -1.6pp · z=-0.45 · p=0.656
vs control · settled only: -1.2pp · p=0.798
★ Carry forward: r4-honest-a · r4-daily-a · r4-instant-a — but for three very different reasons

All three “-a” variants top their lanes, so these are the three to run — and after correcting for prior exposure (see the audience model), two of the three are proven better hooks. Read the reasons, because they decide how much to bet on each:

  • honest-a — a genuine win. Bet on this one. $0.134/picker, 53.7% picker rate, the round's highest CTR (0.936%). It is the only hook that beats its exposed control (+18.1pp), the champion at its freshest (+14.5pp, p=0.0001), and its unseen sibling (+19.4pp) — and the one lane that improved round-over-round ($0.195 → $0.134/picker, -31%).
  • daily-a — a genuine winner, and it rescues a collapsing lane. +23.3pp vs its control (p=1.4×10⁻6), still +21.5pp after subtracting the worst-case exposure drag. The old daily creative is now the worst ad we own (29.0%), so this lane needed rescuing — but daily-a earned the lane on its own numbers, not just by outliving its opponent.
  • instant-a — NOT proven. A tie, carried on judgement. +1.3pp vs its control (p=0.764), and -0.9pp once the exposure ceiling is subtracted — a dead heat either way. By the pre-declared rule the incumbent keeps the lane; we carry instant-a because it crushes its own sibling (+24.2pp, p=1.3×10⁻7) and the incumbent's CTR is sliding (-20%). This lane is the best candidate for a genuinely new idea — it is the one place R4 found nothing better than what we had.

Also worth banking: daily-b beat its control too (+15.1pp settled, p=0.002) and is a legitimate 4th creative if we widen the set. instant-b is the round's only proven loser — it loses to its own control by -22.9pp (p=5.5×10⁻7). Kill it.

What the winning hooks have in common

Three lanes, three independent tests, and the same shape of line won all three. The “-a” hooks all name what the player is freed to do, right now:

✅ honest-a “No energy bars. No lives to refill. Play as much as you want, free.”
✅ daily-a  “Everyone plays the same run today. How far do you get?
✅ instant-a “You are one tap from a run. No app store, no account.”

❌ honest-b  “No gacha. No ads.” — a list of what we aren't. Nothing to do.
❌ daily-b   “Come back for tomorrow's run.” — asks for commitment before the first run.
❌ instant-b “From a link to floor one in ten seconds.” — a spec, not an invitation.

The losers are not badly written; they are addressed to someone who already cares. honest-b and instant-b both take the winner's own barrier and state it as a feature rather than a permission, and both fall ~20pp. That is the most actionable creative lesson in the round — and, unlike the retention read, it is backed by three independent replications.

What we can — and cannot — statistically confirm

Same discipline as the R3 report: a metric is only usable if it has enough events to beat noise. This round is over-powered at the top of the funnel and under-powered at the bottom. Read each tier accordingly.

CTR  Confirmed

58,677–73,587 impressions per ad. Extremes are unambiguous.

honest-a 0.936% vs daily-ctrl 0.601% → a 60% relative gap on ~60k impressions each. Real.

Picker rate / $ per picker  Confirmed — the decision metric

94–286 picker-reaching players per ad — enough to separate tiers cleanly, and it holds on settled traffic alone.

honest-a 53.7% [49.4–57.9] vs daily-ctrl 29.1% [24.4–34.3] → non-overlapping.
6 new-hook-vs-control tests → Bonferroni α = 0.0083. Clearing it: honest-a (p=2.1×10⁻8), daily-a (p=7.1×10⁻9), instant-b as a LOSS (p=0.007).
Not clearing it: instant-a (p=0.102), honest-b (p=0.656), daily-b (p=0.042 on all traffic; 0.002 settled — borderline).

Completion rate  Directional

25–96 completers per ad. The top/bottom band separates; individual ranks do not.

The spread (7.7% → 18.8%) tracks picker rate almost perfectly, which is the tell: the ad is choosing who arrives, and the game converts them the same way regardless of which ad sent them.

Per-ad D1 / D7 for Round 4  Do not read

R4's D7 cannot mature until ~Jul 24 (last cohort Jul 17 + 7 days). Every R4 D7 cell currently reads 0% — that is “too early,” not “zero.” D1 lands at 5–18 returners per ad.

The R3 report predicted this exactly, in advance: “Do not read per-ad D1/D7. At ~9 ads × ~$36, each ad gets ~15–25 D1 returners and ~0 usable D7. That is noise; looking at it will only mislead.” Round 3's matured data (Part ①) now proves the prediction was right.

Our champions got worse — and it is not what it looks like

Round 4 re-ran the three R3 champions unchanged, which accidentally gave us a clean second measurement of the same creative one week later. All three got worse. Picker rates below use settled R4 traffic only, so the learning phase isn't doing the talking.

ChampionCTR · R3CTR · R4ΔPicker · R3Picker · R4 settledΔ
instant 0.957% 0.767% -20% p=0.0002 61.4% 52.1% -9.3pp p=0.010
daily 0.813% 0.601% -26% p=3.8×10⁻6 52.0% 29.0% -23.0pp p=2.0×10⁻7
honest 0.770% 0.661% -14% p=0.020 41.6% 37.9% -3.6pp p=0.398

First: they really were the same ads

Before claiming “the same ad got worse,” the claim has to survive the obvious objection — that R4 re-exported, re-cropped or tweaked the art. It doesn't need to survive it, because the files are byte-identical: same sha256 for the image R3 uploaded and the image R4 uploaded, and the headline copy is unchanged.

instant sha256 d2dc262f690d… IDENTICAL FILE
daily sha256 8b281d3d8c57… IDENTICAL FILE
honest sha256 b1d325bc8861… IDENTICAL FILE

But the shape says it is not simple wear-out

Fatigue predicts a slope — an ad grinding down as exposure accumulates. What we actually see is a step: flat inside R3, flat again at a lower level inside R4. And instant visibly climbs back once delivery settles, which is a cold-start signature, not a wear-out one. Violet = R3 · gold = R4 · red line = the campaign boundary · days with <15 landers dropped.

instant
03
04
05
06
07
08
09
11
12
13
14
15
16
daily
03
04
05
06
07
08
09
10
11
12
13
14
15
16
honest
03
04
05
06
07
08
09
10
11
12
13
14
15
16
R3 campaignR4 campaign (same creative)new campaign starts
Ruling out the two boring explanations

The channel did not collapse: blended picker rate across all paid traffic went up over the two rounds (35.5% on 07-03 → 45.1% on 07-16). The product did not regress: the drop tracks the individual creative, not the calendar — a build regression would have hit every ad on the same day, and none did.

Honest verdict: half of this is measured, half of it is a hypothesis
  • CTR decay — SUPPORTED. All three champions lost CTR on ~60k impressions each; two of three clear p<0.05 and daily clears p<1e-5. Same creative, same communities, fewer clicks. That is fatigue in the ordinary sense: same picture, same people, fewer clicks.
  • Picker-rate decay — REFUTED — it is POOL DEPLETION, not fatigue. The reach pull kills the repetition story: per-ad frequency is ~1.2, so almost nobody saw any ad twice, and only ~3.4% of a control viewer had seen that creative before — a ceiling of 1.4-2.2pp against drops of up to 23pp. The shape agreed all along: a STEP at the campaign boundary, not a slope. What actually changed is WHO is left. R3 creamed the most responsive 423k of a ~1.57M pool; R4 had to work the remainder, so the identical creative meets harder people.

Why this mattered: the two explanations have opposite symmetry, and therefore opposite consequences for every winner in this report. Repetition-fatigue would hit only the three controls (biasing every result); anything that hits all nine ads equally biases nothing.

RESOLVED — frequency + reach pulled from the Reddit Ads Manager 2026-07-17. Frequency ~1.2 and 73% brand-new audience ⇒ not repetition fatigue. Overlap arithmetic on de-duplicated reach ⇒ a ~1.57M pool, 46% already burned ⇒ pool depletion. The verdict is pool depletion — symmetric, so the winners stand — and a far more serious problem than fatigue, because you cannot rotate your way out of it.

Grading the Round-3 prediction

The R3 report published a falsifiable forecast specifically so it could be graded. Here is the scorecard, unedited. 3 hits, 3 misses — and the pattern in which ones is the most useful thing on this page.

MetricR3 actualR4 targetR4 actualNote
Blended CTR 0.775% 0.85% – 1.0% 0.746% Miss Fell slightly. No hook lifted blended CTR into the target band.
Cost per player $0.076 $0.060 – $0.070 $0.102 Miss Went UP 22%, not down. Media cost rose (eCPM $0.51 → $0.58).
Players for ~$320 4,206 4,600 – 5,300 3,435 Miss 18% FEWER players for the same money.
Best new hook vs incumbent ≥1 concept improves 2 of 3 lanes beat their control Hit honest-a and daily-a both beat their incumbent decisively; instant-a did not.
D7 ~1% (proxy) ~1% — unchanged 1.14% (R3 matured) Hit Called correctly. The 3× D1 lift did NOT carry to D7.
Per-ad D1/D7 significance none still none still none (χ²=5.4, p=0.80) Hit Certain in advance, confirmed.
The pattern: we predicted the game correctly and the market incorrectly

Every claim about our own funnel (D7 won't move, per-ad retention won't reach significance, at least one hook will beat its incumbent) was right. Every claim about what the ad auction would sell us (cheaper clicks, more players, higher CTR) was wrong — and wrong in the same direction, because eCPM rose 12% ($0.52 → $0.58) between rounds. Better creative bought better quality; the auction took the savings. Forecast our funnel; do not forecast the auction.

The bottom line, in the only unit that matters

Cost per click, cost per player, even cost per picker are all intermediate. The business question is what a dollar of Reddit buys in players who are still here next week. Round 3 is the only round old enough to answer, and now it has.

Round 3 · fully matured

$350.00 spent → 4,209 players landed  ($0.083 each) → 433 came back at all  ($0.81 each) → 48 came back on day 7  ($7.29 each)

Round 4 · projected

$350.01 spent → 3,435 players landed  ($0.102 each) → ~39 projected day-7 players → ~$8.93 each (projection at R3's matured rate — real read ~Jul 24)
$7.29 per day-7 player — and no hook moves it

Two rounds and $700.01 have now established the shape of this channel with real confidence: Reddit will sell us players at ~$0.07–0.09 all day long, and roughly one in ninety of them is still here a week later. Creative work moves the top of that funnel a lot — honest-a buys a picker for $0.134 where the round's worst ad pays $0.408, a 3.0× spread — and moves the bottom of it by nothing measurable.

The R3 report called this a round early: “Round 4 should be the last wide creative round. After it, the leverage moves off the ad account and onto the product.” The matured D7 now confirms it. We have a proven best hook per lane. There is no third creative round that changes the $7.29.

Next steps

One thing to stop, one to keep running cheaply, one date on the calendar, and one place to put the effort instead.

1
Now · this week

Consolidate onto the three winners and stop testing hooks

Kill the six losers, keep honest-a · daily-a · instant-a running at a maintenance budget as the standing acquisition set.

  • Kill immediately: instant-b (proven loser, -22.9pp vs its own control), honest-b, and all three controls — including daily-ctrl, which is now the worst ad in the set and measurably burnt out.
  • Keep on the bench: daily-b — it also beat its control and is the natural rotation partner when daily-a fatigues.
  • Do not run Round 5. The hook question is answered: name what the player is freed to do, right now. A third wide round would re-buy a lesson we already own.
2
✅ Done · 2026-07-17

Frequency + reach pulled — the exposure objection is closed

The question that gated everything: were the controls only losing because people had already seen them? No. Frequency ≈ 1.2, 73% of R4's audience was brand new, and the worst-case exposure drag is ≤2.2pp against drops of up to 23pp. honest-a and daily-a survive the correction; instant-a stays a tie.

  • It surfaced something worse: a ~1,572,564 pool, 46% of it already burned. That is the real ceiling — see Burning the pool.
  • Still fix the test design: never score a fresh challenger against an already-exposed incumbent. The bias was small this time only because frequency was low; at a higher budget or a narrower audience it would not be. Judge sibling-vs-sibling (a vs b) — it needs no correction at all.
3
Standing · budget for it

Treat creative rotation as a running cost

This round proved a champion decays measurably within ~2 weeks of exposure (CTR -20% / -26% / -14% across the three). A creative is not an asset we buy once.

  • Plan on refreshing the art/line of each lane every ~2 weeks, from the proven “-a” template, rather than re-testing the concept.
  • Watch CTR as the early-warning signal — it decayed on all three champions before picker rate did.
4
Fixed calendar checkpoint 2026-07-24

Read R4's matured D7 — and grade this report too

R4's last cohort landed Jul 17, so its D7 matures Jul 24. Re-run the readout then. The falsifiable prediction: R4's blended D7 will land in 0.86%–1.51%, statistically identical to R3's 1.14%, and no per-ad D7 will reach significance.

  • If that prediction holds, the case is closed: retention is not purchasable through this channel, at this budget, with any hook.
  • If R4's D7 comes in meaningfully above 1.51%, that is not the ads — that is the product work that shipped between the rounds, and it deserves a proper look.
5
Where the leverage actually is

Move the effort onto the game

Both rounds now agree on where the money leaks, and it is not the ad account.

  • The title screen still eats 59% of everyone we buy (2,015 of 3,435 R4 players never reached the path-picker). It is improving — blended picker rate rose from ~37% in R3 to ~46% on R4's settled days — which is the strongest evidence yet that product work moves this number when creative cannot.
  • D7 ≈ 1% is the ceiling on everything above it. Every acquisition dollar is multiplied by that number. Doubling D7 is worth more than halving CPC, and unlike CPC it is entirely within our control.
  • The honest framing for the next planning session: paid acquisition is now a measurement instrument, not a growth strategy. It reliably delivers ~4,000 strangers for ~$320 to test product changes against. That is genuinely valuable — it is just not a growth channel until D7 moves.

Method & provenance

Data. 958,653 telemetry events over 2026-06-18 → 2026-07-17, 9,432 distinct players, internal/dev devices excluded. Spend, impressions, clicks and CTR are typed from the Reddit Ads Manager (final billed) — spend never enters telemetry by design, so the cost math is done outside the pure layer.

Reproducibility. Generated by src/dashboard/campaignReadout.gen.test.ts (gated: AOD_CAMPAIGN=1), which recomputes per-creative cohorts and asserts they reproduce the dashboard's own acquisitionContents() percentages — so these counts cannot silently diverge from /dashboard.html. Every figure on this page is interpolated from campaign-stats.json; none is hand-typed.

A trap worth recording. npm run dashboard:fetch caps at the newest 500k rows, which today starts 2026-07-10 — Round 3 is entirely absent from a default export, and its stragglers get re-dated to the window edge, which would have produced a plausible-looking but fictional “R3 D7.” This analysis pulls the full 958,653-row history via the backfill sweep, and the generator now fails loudly if any R3 cohort lands on the export's first day.

Definitions. Dn = active on exactly day n after first seen, counting only devices old enough to know (firstDay + n ≤ last day of data) — the dashboard's own rule. D7+ = active on any day ≥ 7. Stayed = seen on ≥2 distinct days. Picker = fired path_picker_shown (instrumented since Jun 29; comparable across both rounds — unlike title_engaged, which only exists from Jul 9 and therefore cannot be compared between R3 and R4). Significance: two-proportion z-tests (pooled), Wilson score intervals, χ² homogeneity across ads; Bonferroni where families of comparisons are read together. iOS-web storage eviction undercounts returns, so all retention figures are floors.

Reddit Ads Manager · r4-refinement public/events.json · full backfill src/dashboard/retention.ts src/dashboard/adQuality.ts ← Round 3 report

← Back to the forum