UncategorizedAug 21, 202610 min read

Why Most App Store A/B Tests Never Give You an Answer

OA
OWA AI
Author
Why Most App Store A/B Tests Never Give You an Answer

Ninety days ago you allocated the traffic. You built the screenshots, you wrote the hypothesis, you waited.

Today the test expired. There are two conversion rates on the screen, a lift figure, and no label attached to either version. Apple has not told you which one won, and it is not going to.

Nothing went wrong during those ninety days. The test ran exactly as designed. It simply never gathered enough evidence to say anything, and whether it ever could was settled before the first impression landed.

This post is about that ending: what causes it, how to know in advance whether your test can finish, and what to do instead when it cannot. If you have not run one of these before, how to run an App Store A/B test covers the mechanic from the start. Everything below assumes you already tried and came away with nothing.

What Apple was waiting for

Apple marks a winner with a label, either "Performing Better" or "Performing Worse." It only attaches that label once it is at least 90% confident the difference between your versions is real.

Real, in this context, has a specific meaning. Suppose your original page sits at 5.0% and your treatment at 5.4%. Is the treatment better, or did a slightly more interested group of people happen to land on it?

With small numbers you genuinely cannot tell, and this is not a flaw in the tooling. Flip a fair coin ten times and you will sometimes get seven heads. That does not make the coin biased. Two groups of visitors differ from each other by luck in exactly the same way, all the time, for no reason at all.

Confidence is the measure of how much of your gap is signal and how much could still be luck. It starts low and climbs as more visitors flow through, because luck averages out at scale the way seven-in-ten heads disappears over ten thousand flips.

So the label is not Apple's opinion of your creative. It is Apple telling you the counter got far enough for the difference to be trustworthy. Below 90%, it says nothing, because it does not yet know.

Your test did not fail to produce a winner. It failed to produce enough data to identify one.

Blog image

Why some tests can never get there

Three things decide whether a test can reach 90%: how much traffic you give it, how many treatments you split that traffic across, and how big a difference you are trying to detect. All three are set before the test opens, and none of them can be changed once it is running.

The third one is the one nobody thinks about, and it dominates the other two.

Here is the shape of it. On a 5% baseline conversion rate, at 90% confidence, roughly the impressions each version needs:

- a 30% relative lift: around 3,000

- a 20% lift: around 6,400

- a 10% lift: around 24,600

- a 5% lift: around 96,200

(These come from standard sample-size maths rather than from Apple, so read them as the shape of the problem, not as Apple's own figures.)

Read the shape, because the shape is the lesson. Halving the effect you are chasing roughly quadruples the traffic you need. A 5% lift is not slightly harder to prove than a 10% lift. It is about four times harder.

Blog image

Now add the second fact that governs everything: a test has a hard ceiling of 90 days.

A test with one treatment against your original, chasing a 10% lift on a 5% baseline, needs roughly 49,000 impressions in total. An app with 10,000 store page impressions a month needs about five months to get there. Apple's ceiling is three. That test cannot finish. It will run its full ninety days, expire, and hand back two numbers and no label, and the team finds out a quarter later.

The same test at 30,000 impressions a month concludes in about six weeks. At 100,000, in about two.

And the allocation you chose multiplies all of it. The share you assign to the test is divided evenly across your treatments, so a third treatment does not come free. It takes traffic from the other two and pushes every one of them further from the line.

This is not a rare failure, and Apple knows it. Before you launch, App Store Connect will estimate how many impressions you would need to reach an outcome at 90% confidence, using your own daily impressions and downloads. That feature exists because not reaching an answer is an ordinary result.

Which makes the single most useful instruction in this post a boring one. Run that estimate before you build a single asset. If it runs past ninety days, your test is not slow. It is impossible, and you have just saved yourself a quarter.

What to do when your test won't conclude

Everything below follows from that arithmetic. Low traffic does not mean stop testing. It means test differently.

Chase bigger differences, not smaller ones. This is the counterintuitive core of the whole thing. If a 5% lift needs four times the traffic of a 10% lift, then subtle refinements are precisely what a smaller app can never prove. Stop testing a warmer shade of blue. Test a genuinely different concept: a different first screenshot, a different promise, a different icon direction. Big differences need far less traffic to confirm, and they are the only differences you have the traffic to read.

Two treatments beat three. A third one draws traffic away from the other two and pushes all of them further from the line. On limited traffic, one clean comparison finishes while three starved ones expire together.

Allocate more, not less. A cautious 20% allocation on a small app is how a test guarantees its own failure. If you are going to spend the quarter, give it the traffic to conclude.

Do not split your localisations. Running the same test across five markets divides your data five ways. Run your single largest market alone and let it actually reach a verdict.

Accept that some things are untestable at your size. Decide those with other evidence: what converts for larger apps in your category, qualitative feedback, first principles. That is more honest than pretending a test that never concluded settled the question.

And do not ship the variant because it was ahead. A treatment leading at 61% confidence is not a small win. It is not yet anything. Confidence stopped short of the line precisely because the gap could still be luck. Shipping it is guessing with extra steps, and now the guess is wearing a number.

Where a competitor's test becomes your evidence

Follow all that honestly and you arrive somewhere uncomfortable. If your own traffic cannot conclude a test, you have a real problem. You need evidence to make a good decision and you cannot manufacture it. No amount of care fixes a sample size that is not there.

But the constraint is yours, not the category's. A competitor with ten times your traffic can conclusively test things you never could. Their finished test is a settled result in your market that somebody else paid for, and unlike yours, their setup is not private.

You do not go looking for it. A competitor's test announces itself, because the moment they open one their store page changes, and that change is public. OWA watches the pages of the apps you track and marks the day a test appears, so the first you hear of it is that it started, not that it finished six weeks ago.

Blog image

Then you open it. What is visible from outside is their current page against each of their treatments, the share of traffic each one is carrying, the device size the test is running on, and the dates: the day it started, the day it ended, and the day the winning assets went live.

What is not visible is their confidence figure. That number lives inside their App Store Connect account and never leaves it. So the winner is not read from Apple's verdict. It is read by watching which set of assets they actually keep once the test is over, which is slower and, for your purposes, just as useful.

Now read it with everything you just learned, because the same arithmetic that governs your test governs theirs, and their setup is visible while yours is private.

How many treatments they ran tells you how much traffic they had to spare. You now know a third one does not come free. A competitor running three of them is telling you they have the volume to starve all three and still reach a verdict.

The split tells you their appetite. An even division is the ordinary case. A protective split, where the original keeps most of the traffic, means they were unwilling to risk their current page and still expected to conclude anyway, which is only possible at scale.

What changed between the versions tells you how big a difference they were chasing. A completely replaced screenshot set is a big swing, the kind that reaches significance quickly. A reordering of the same assets is a fine adjustment, and a fine adjustment only concludes for an app with traffic to burn. If you see a competitor testing something subtle, you are looking at a company that can afford to.

How long it ran is a rough read on their volume. A test that started and shipped a winner inside three weeks belongs to an app with a lot of traffic. One that ran the full ninety days and shipped nothing may have hit exactly the wall you just hit.

None of this is their result. It is their reasoning, made visible.

And that is what you take. You cannot adopt their winner, because their audience is not yours and a lift measured on their traffic says nothing certain about yours. But a shipped winner is a concept that already survived a real test, and the problem this section started with was that you could not afford to generate one. So you borrow the direction rather than the outcome, and you test it as the kind of big, clearly different change your own traffic can actually resolve.

FAQ

Why did my App Store A/B test not show a winner?

Apple only attaches a "Performing Better" or "Performing Worse" label once it is at least 90% confident the difference between your versions is real rather than chance. If your test ended without a label, it never reached that threshold. That is usually not a problem with the creative. It means the test did not gather enough impressions to separate a real difference from luck, which is decided by how much traffic you allocated, how many treatments you split it across, and how large a difference you were trying to detect.

Can you run an App Store A/B test with low traffic?

Yes, but not for small differences. The traffic needed to prove a difference rises sharply as the difference shrinks, so a low-traffic app that tests a subtle change will usually run out of days before it runs out of doubt. The workable approach is to test larger, clearly different concepts, keep the test to a single treatment against your original, and allocate generously. App Store Connect will estimate the impressions you need before you launch, which tells you whether the test is finishable at all.

How long does an App Store A/B test take to reach significance?

It depends on your traffic and on how big the difference is, not on a fixed number of days. A large true difference reaches significance in weeks. A small one may never reach it. The hard ceiling is 90 days, after which the test expires whether or not it concluded. Apple's guidance is to run to significance rather than to a date, so a flat first week proves nothing except that the counter has not arrived yet.

Is it safe to ship a variant that is ahead but below 90% confidence?

It is a risk decision rather than a rule, and it depends entirely on what the change costs to reverse. A screenshot reorder you can undo tomorrow is not the same bet as an icon baked into a release. What you should not do is treat the number as a result. Below the threshold, a lead can still be luck, which is exactly why Apple declines to call it.

The part that was decided before you started

A test that cannot conclude is not a slow test. It is a test that was never going to answer you, built by a team that did not know that yet.

The good news is that this is knowable in advance, in about five minutes, before a single asset exists. Run the estimate. Pick a difference big enough for your traffic to see. Keep it to one treatment if you cannot afford two.

And when the honest answer is that your own traffic cannot settle the question, the cheapest evidence available is a test somebody else has already finished. You can see which of your competitors have one running right now, free to start.