A/B test sample size is the amount of information I plan to collect before deciding whether a controlled comparison can answer my product question. I treat it as a planning decision, not a number I search for after the test has started. A sample-size estimate depends on the outcome, the baseline behavior, the change I need to detect, the uncertainty I will tolerate, and the traffic available to the experiment.
A calculator can help with the arithmetic, but it cannot decide whether the question is worth testing or whether the metric represents customer value. I use the calculation to make tradeoffs visible. If the required sample is larger than the audience or time window makes reasonable, I revisit the design instead of quietly lowering the standard after seeing early results.
Start with the decision and outcome
I write the decision before I discuss sample size. What will I do if variant B improves the primary outcome? What will I do if it is neutral? What would make me stop for safety? These answers determine what the experiment needs to observe.
I choose one primary outcome that is close enough to the product decision to be useful. A click may be easy to count but too far from the customer value I care about. A completed workflow, successful activation step, renewal behavior, or quality signal may be more relevant, although it can take longer to observe. I define the unit of analysis as well: a person, account, session, transaction, or another unit that matches how the treatment is assigned.
I also identify guardrail metrics. A change can improve the primary outcome while increasing errors, support burden, cancellations, or another important cost. Guardrails do not replace the primary outcome, but they keep an attractive local result from hiding harm elsewhere.
Define the detectable change
The minimum detectable effect is the smallest change I want the experiment to reliably distinguish from ordinary variation under the chosen design. I do not pick it because it produces a convenient sample. I ask whether the change would alter my decision, justify the rollout cost, or matter to customers.
For a conversion metric, I start with a defensible baseline from a comparable period and population. For a continuous measure, I consider its typical spread and whether the average is the right summary. For a ratio or account-level metric, I make sure the denominator and aggregation rules are stable. If the baseline is uncertain, I record a range or run a measurement period rather than presenting a guess as a fact.
A very small detectable change may require more observations and a longer run. That is not a flaw in the calculator. It is a signal that the question may be too sensitive for the available traffic, that the outcome may be noisy, or that I need a different product decision.
Make the statistical choices explicit
The inputs commonly called significance level and power describe how I want the test to behave. The significance level sets a tolerance for false positives under the model. Power describes the chance of detecting the chosen effect if that effect is really present under the assumptions. I document the choices before launch and avoid changing them because an early result is exciting or disappointing.
I also record the allocation between variants, the expected traffic, the observation window, and whether the metric is binary, continuous, count-based, or time-to-event. Unequal allocation changes the information available in each group. A metric with delayed outcomes may require more calendar time than the raw event count suggests.
I check whether the experiment has enough independent units. Repeated events from one person do not automatically equal independent people. In a business product, accounts, teams, or organizations may be the unit that receives the treatment. If members of the same account influence one another, I account for that dependence in the design and interpretation.
Plan for time and exposure
I estimate the sample required per variant, then translate it into a calendar window using eligible traffic rather than total site traffic. I account for ramp-up, exclusions, time to experience the feature, and the delay before the primary outcome occurs. I want the test to include the normal operating conditions that could affect the decision, such as weekdays, billing cycles, or common usage patterns when they are relevant.
I do not stop the test simply because the current estimate crosses a target. I predefine the stopping rule and use a method appropriate for any planned interim looks. Repeatedly checking and stopping when the result is convenient changes the false-positive behavior of the analysis. If product safety requires an early stop, I make that an explicit rule rather than pretending the interruption was unrelated to the statistics.
I revisit the plan when exposure is not what I expected. An underfilled experiment may not be able to answer the original question. An unexpectedly large audience can create operational or ethical concerns. I document the change and its effect on interpretation instead of treating the original plan as if it happened unchanged.
Treat the estimate as a planning tool
A sample-size calculation depends on assumptions. Baseline behavior can drift, users can be excluded, tracking can fail, and the effect I care about may differ by segment. I check instrumentation before launch and compare the assignment, eligibility, exposure, and outcome counts. If those checks fail, more observations will not repair the measurement problem.
I do not use a large sample to turn a tiny, irrelevant effect into a product victory. Statistical evidence answers a question about compatibility between data and a model; it does not decide whether the change is valuable, safe, or worth maintaining. I combine the result with effect size, uncertainty, guardrails, qualitative context, and implementation cost.
I connect this planning work with the broader A/B testing guide for product managers and experiment design for product managers. Before relying on a metric, I also review product analytics instrumentation for PMs so the outcome and exposure events mean what I think they mean.
My bottom line
I plan A/B test sample size from a decision, not from a desired headline. I define the outcome, choose a meaningful detectable change, document the statistical and operational assumptions, and protect the test from convenient stopping. The final number is useful when it helps me decide whether the experiment can provide credible evidence—and when it makes clear what the experiment still cannot tell me.
Next step
For a structured foundation in product experimentation, I use the Product Manager Certification, a commercial Product HQ resource. The Product HQ newsletter shares practical product lessons and frameworks.