A/B Test Sample Size Before the Experiment Starts
Decide the Smallest Useful Lift
An A/B test can be mathematically clean and still waste time if it is underpowered. Sample size planning asks a practical question before launch: how many observations does each variant need before the test has a reasonable chance of detecting the effect we care about? Without that step, teams often run tests that cannot distinguish noise from a meaningful change. The result is a dashboard full of movement and no decision worth trusting.
The Normal Approximation Behind the Estimate
For a conversion-rate test, each visitor either converts or does not. The observed rate is an estimate of the true rate, and estimates bounce around because samples are finite. Smaller effects are harder to detect because the two rates sit closer together. Higher confidence asks for stronger evidence before calling a winner. Higher power asks for a better chance of noticing the effect if it is real. Those goals all push sample size upward.
The working equation is n per group ~= 2 * pbar*(1-pbar) * (z_alpha/2 + z_power)^2 / effect^2.
The normal approximation uses a baseline conversion rate, a target rate, confidence, power, and the absolute difference between rates. The minimum detectable effect should be entered in percentage points. A move from 5 percent to 6 percent is a 1 percentage point absolute lift, not a 1 percent relative lift. The formula estimates the number of users per variant. Doubling that gives total sample for a two-arm test. The result should be rounded up because partial users are not helpful.
Model limit: Uses a normal approximation for planning. Sequential testing, multiple comparisons, and low counts require more care.
A Five-to-Six Percent Experiment
Baseline conversion should come from recent, comparable traffic. If weekday and weekend behavior differ, use the segment that matches the experiment. Minimum detectable effect should be the smallest change worth acting on, not the smallest change someone hopes to see. Confidence is commonly 95 percent, while power is often 80 or 90 percent. Raising both makes the test more demanding. If traffic is limited, the honest answer may be that the test cannot detect small changes in a reasonable time.
Inputs That Belong in the Protocol
A product has a 5 percent baseline conversion rate and wants 80 percent power to detect an absolute increase of one percentage point at 95 percent confidence. The target rate is 6 percent. With the calculator's pooled normal approximation, the required sample is about 8,159 observations per variant, or 16,318 total. At 2,000 eligible visitors per day split evenly, enrollment alone takes a little over eight days. A smaller detectable lift would require much more data because sample size grows roughly with one over the squared effect.
The estimate assumes independent observations, a fixed test plan, and a binary outcome measured consistently. If users can appear multiple times, analyze at the user level or use a model that accounts for clustering. Do not stop the first evening a conventional p-value crosses 0.05; repeated peeking raises false-positive risk unless a sequential design was chosen in advance. Add allowance for bot filtering, missing outcomes, guardrail metrics, and the need to span weekday cycles before publishing an experiment end date.
Peeking and Attrition Distort the Plan
A frequent mistake is peeking at the test every hour and stopping when the graph looks exciting. That changes the statistical behavior unless the test is designed for sequential monitoring. Another mistake is testing too many metrics or variants without adjustment. Low baseline rates also make normal approximations less reliable. Sample ratio mismatch, bot traffic, repeated users, seasonality, and instrumentation changes can damage a test more than the sample size formula can repair.
Sample per variant is the planning number. Total sample helps estimate calendar time from traffic. Target conversion reminds everyone what effect the test is powered to detect. If the required sample is huge, the team has options: accept only larger detectable effects, run longer, target a higher-volume metric, reduce variance, use a different design, or skip the test and make a product judgment. Not every decision deserves an experiment, and not every experiment is affordable.
From Sample Count to Calendar Time
Use the calculator before writing the experiment ticket. Put the sample size, planned duration, primary metric, decision rule, and guardrail metrics in the test plan. During the test, monitor data quality and sample ratio, but avoid changing the decision rule midstream. After the test, report the observed effect with uncertainty, not just "winner" or "loser." The calculator helps set expectations so stakeholders are not surprised when a small effect needs a long run.
A good A/B test note records baseline rate, minimum detectable effect, confidence, power, sample per variant, planned duration, primary metric, and stopping rule. The calculator does not make experimentation automatic, but it prevents one of the most common failures: running a test that never had enough information to answer the question. The best time to learn that is before launch, when changing the design is cheap.