Atlas of Doors · Field Log · Experiment 01

Take one thing off the first screen.
More players walk in.

A new player's first sight of Atlas of Doors is one screen: a giant New Run door, and beneath it a smaller gold seal — Your Atlas, the home of the meta-game. We ran a live three-way test on brand-new players: show that seal at full size, shrink it, or remove it entirely. Removing it won.

+4.4pp
Activation, removed vs. shown 32.0%36.4%
Read at ~400 sessions / arm

Same door. Same first tap. The only thing that changed between players was how present the Your Atlas seal was, sitting just below the button.

Home screen with the gold Your Atlas seal shown at full size below the New Run button
Control
The gold seal at full size — the original screen.
32.0%
Activation
Home screen with the Your Atlas seal shrunk and dimmed
De-weighted
Seal shrunk, cooled and pushed down.
34.3%
Activation
Home screen with the Your Atlas seal removed, leaving only the New Run door
Removed
No seal at all. New Run is the only door.
36.4%
Activation
Winner · +4.4pp

Straight from our live dashboard. Activation is the gate; the columns to its right show what each version does to the run after a player is in.

atlas-seal-firstjoin · KPI: Activation · collecting

Dashboard table: control 32.0% activation, de-weighted 34.3%, hidden 36.4%; hidden also leads Floor 2, Full Run, Depth and D1; lift +4.4pp, still collecting at the 400 sessions per arm gate

One honest caveat. Everything past Activation in that table — Floor 2, Full Run, Depth, next-day return — is shown for completeness, but only a hundred-odd players per version have reached those stages so far (next-day return is 4 players versus 7). That is far too small to be statistically meaningful, so we don't read it as a result. The one column we trust today is Activation, where the sample is large enough to matter. The deeper reads will come as the numbers grow.

Activation is the number that turns ad spend into players. Lifting it costs nothing extra — the budget is identical. It just makes every acquisition dollar reach further.

+13.8%
more activated players from the same spend (36.4 ÷ 32.0)
−12.1%
lower cost per activated player — a free efficiency gain, no extra budget
Illustration at an assumed $0.50 to bring one new player to the first screen. Swap in your real number — the percentages above don't move.
Per $1,000 of spendControl · shownRemoved · winner
New players landed2,0002,000
Activation rate32.0%36.4%
Activated players640728
Cost per activated player$1.56$1.37
Same $1,000, same clicks: +88 activated players, at $0.19 less apiece. The budget never moves — the rate does the work.

A brand-new player has zero Keys. The Atlas — where Keys are spent — has nothing for them to do yet. So on that first screen the seal wasn't depth; it was a second gold object competing for attention with the one action that matters: starting a run.

Two doors on the first screen is one door too many.

Take the competition away and more players take the single obvious door. The Atlas isn't gone. It opens the moment they finish their first descent, when they've earned Keys and it finally means something. See it early, earn your way into it later. It just doesn't belong in the doorway.

Shown for transparency, not as findings. The samples that reach these later stages are still far too small to call — they only become readable as more players go deep.

Reached Floor 2
54%61%
n ≈ 125–140 / arm
Finished the run
38%50%
n ≈ 125–140 / arm
Avg. depth
24doors
n ≈ 120–130 / arm
Back next day
3.1%4.9%
4 vs 7 players
Treat these numbers as unreliable

Only the activation read at the top had the volume to trust. Everything downstream — Floor 2, finishing the run, coming back the next day — is drawn from ~130 players per arm: far too few to decide anything. We show them for honesty, not as proof. And the real lesson of this log isn't “removing the seal won.” It's this: at our traffic, most experiments can never gather enough data to be reliable — and dressing up a thin, noisy number as a finding is how you fool yourself. When the data can't decide, decide with an informed gut, and make a bet.

No data beats bad data worn with false confidence. If the sample can't call it, call it with your gut.

Before running any A/B test, put it through these rules (now baked into our internal game-council). If a test fails them, don't run it — trust your gut and ship a bet instead: your best judgment, live, with a guardrail to watch.

  1. 1
    Run the power check first — the “can this even be measured?” math
    Work out up front how many players per arm you'd need to actually see the effect. If our funnel can't deliver that within a few weeks, the test is a mirage — you'll collect noise and read tea leaves. Know the answer before you start, not after.
  2. 2
    Don't test retention — it sits at the far end of the funnel
    Next-day return is ~4%: a rare event that needs thousands per arm — months of full traffic — to move a single point. Watch it as a health trend; never A/B it. Test something early and common instead.
  3. 3
    Test the highest-rate proxy — the earliest, most common step
    Statistical power comes from a high base rate. Read the change on Floor-2 reach (~60%) or the very first tap — not the low, downstream outcome. A common step needs a fraction of the sample a rare one does.
  4. 4
    Fewer arms — control plus one bold change
    Every extra arm splits the traffic thinner and demands a bigger sample, and the “winner” of four noisy arms is usually just the luckiest. Two arms halves the wait and the self-deception.
  5. 5
    Only swing big — a bold bet, not fine-tuning
    A ~+10pp swing on a common proxy is testable in days; a few-point tweak never is at our size. If the honest expected effect is small, don't test it — make it a bet: ship your best call and watch a guardrail. Big bats, or gut and go.

The through-line: low throughput means you cannot buy certainty with a test. Spend your rare tests only on big, early, two-armed, powered questions — and decide everything else on gut feeling, shipped as a bet.

How we read it

Every new device is split at random across the three versions — uniform assignment, no targeting on who's likely to spend or stay. The KPI is activation: the first meaningful action a new player takes. The figures here are the live read at roughly 400 sessions per arm; the test is still collecting, so treat the exact points as directional. Activation is the only stage with enough volume to read with confidence — the deeper stages don't have the sample yet.

The test at a glance

Experiment
atlas-seal-firstjoin
Audience
Brand-new players only
Arms
Shown · De-weighted · Removed
KPI
Activation (first real action)
Status
Collecting · gate 400 / arm
Result
Removed leads · +4.4pp activation
Atlas of Doors ·A roguelike of doors ·atlasofdoors.com ·Experiment Log · July 2026

← Back to the forum