A/B testing for product managers

Kevin Lee
By
Kevin Lee
Kevin Lee
Kevin Lee
Kevin is a Co-Founder of ProductHQ. He has worked as a VC at Pear Ventures where he invested in and partnered with…
More About Kevin →
×





A/B testing for product managers | Product HQ



A/B testing for product managers

A/B testing is a controlled comparison in which different groups receive different versions of a product experience and I measure a predefined outcome. Product managers use it to learn whether a change caused a meaningful difference, not simply whether a new screen received positive reactions. A test can inform a decision, but it cannot rescue a vague problem statement or unreliable instrumentation.

I treat experimentation as one part of product discovery and delivery. It works best when the change can be isolated, the population is large and stable enough for the decision, and the outcome can be measured without harming customers. When those conditions are absent, interviews, usability studies, staged rollouts, cohort analysis, or operational evidence may be more honest.

Start with a decision and hypothesis

Before asking for an experiment, I write the decision it will support. Will we ship a new onboarding step, keep the current flow, roll back a change, or investigate a segment? Then I state a falsifiable hypothesis: “For new workspace admins, reducing the number of setup choices will increase completion of the first useful action because it lowers decision burden.”

I name the expected direction, the population, the intervention, and the reason. If I cannot explain what evidence would change my mind, I am not ready to test. I also consider whether the effect could be too small to matter. A statistically detectable change that does not improve a customer or business outcome may not deserve the engineering and complexity cost.

Choose the experiment unit

I decide what gets randomized: a visitor, account, user, session, device, team, or another unit. The choice depends on how the product is used and where contamination could occur. If teammates share a workspace, randomizing individual users may allow the treatment to spill into the control experience. In that case, a team-level assignment may be more appropriate, even if it reduces the available sample.

I define eligibility, assignment, exposure, and analysis rules before launch. I check whether returning users can move between variants, whether internal users are excluded, and whether a user can be counted more than once. I document these choices so the result is reproducible rather than a story assembled after the numbers appear.

Select a primary metric and guardrails

I choose one primary outcome tied to the hypothesis. It might be completion of a meaningful task, activation, retention, revenue, or another defined behavior. I write the numerator, denominator, time window, segment, and data source in plain language. “Engagement increased” is not a metric definition.

I add guardrails for quality and risk. Depending on the product, these may include errors, support contacts, cancellations, latency, accessibility signals, complaint rates, or downstream retention. I also check an actionable north-star metric and a relevant counter metric when the experiment could improve a local number while harming the broader experience.

I avoid adding a long list of primary outcomes after seeing the data. Secondary metrics can help explain the result, but the interpretation should distinguish confirmatory evidence from exploration.

Validate the measurement plan

Before exposing customers, I test event names, properties, timestamps, assignment, eligibility, and the connection between exposure and outcome. I inspect a small internal or staged sample when appropriate. I ask analytics and engineering to check for missing events, duplicate events, delayed reporting, and variant leakage.

A clean-looking dashboard can still answer the wrong question. If the new flow changes the opportunity to trigger an event, the measured lift may reflect instrumentation rather than behavior. I compare pre-test patterns when possible and record any product, marketing, pricing, or seasonal change that could affect the result.

Run the test with discipline

I agree on the minimum duration, sample or information threshold, stopping rules, and decision criteria before launch. I do not stop as soon as a promising number appears or continue indefinitely because the result is inconvenient. The exact method depends on the organization’s statistics practice, but the product judgment is consistent: make the rules visible before the outcome is known.

I watch for novelty effects, seasonality, outages, unequal assignment, selective exposure, multiple simultaneous tests, and changes in traffic mix. A result can be real for one segment and absent for another. I inspect meaningful segments carefully rather than slicing the data until a favorable pattern appears.

I also protect customers. I avoid testing changes that create unacceptable safety, privacy, accessibility, or financial risk merely to obtain a cleaner comparison. A control group is not permission to withhold a necessary fix.

Interpret results for a decision

When the test ends, I ask four questions: Did the experiment run as designed? What happened to the primary metric? What happened to guardrails? Is the effect large and durable enough to matter? Statistical significance, where used, is evidence about uncertainty under a model; it is not a guarantee that the change is valuable or caused by the feature in every context.

Possible decisions include ship to everyone, ship to a segment, iterate and retest, keep the control, roll back, or gather qualitative evidence. I document the result, limitations, and next action. A neutral result can be valuable if it prevents a broad rollout or reveals that the assumed problem was not the right problem.

When an A/B test is not the answer

I do not force experimentation onto small populations, long sales cycles, low-frequency outcomes, safety-critical changes, or experiences where treatment contamination is unavoidable. For those cases, I combine customer research, prototype evaluation, phased release, cohort comparisons, and expert review. In B2B products, a customer conversation or implementation signal may be more actionable than a shallow split test.

Common mistakes

I watch for testing multiple changes at once without a reason, choosing vanity metrics, ignoring guardrails, peeking repeatedly, and claiming causality from a before-and-after comparison. I also avoid using an experiment to make a predetermined decision look objective. Good experimentation makes uncertainty visible; it does not replace responsibility.

Next step

Before I randomize a test, I step back through experiment design for product managers so the hypothesis, guardrails, and decision rule are clear enough that the A/B result can actually change the roadmap.

I use A/B test sample size guidance for PMs to plan a detectable change before I start a controlled comparison.

I use statistical significance guidance for PMs to interpret test evidence without confusing a threshold with product value.

Build practical skills in metrics, discovery, experimentation, and product judgment in the Product Manager Certification. Subscribe to the Product HQ newsletter for frameworks and templates that help me turn product evidence into better decisions.


Kevin Lee
Kevin Lee
Kevin is a Co-Founder of ProductHQ. He has worked as a VC at Pear Ventures where he invested in and partnered with early-stage founders on product & growth to help them build the foundations of category-defining companies. He has worked as a Product Manager at AltSchool (backed by Andreessen Horowitz, Founders Fund, First Round Capital, Mark Zuckerberg, John Doerr and other exceptional investors). Previously, he was a Senior Product Manager at Kabam (acquired by NetMarble and Fox for a combined $1bn+), where he worked on products through all lifecycles in San Francisco, Vancouver, and Beijing and helped grow one of the company’s products to become the third largest revenue generating product in the company portfolio. In a former life, he worked in Technology Investment Banking at Merrill Lynch. He is also the author / co-author on 10+ gaming patents.