Feature flags for product managers

Kevin Lee
By
Kevin Lee
Kevin Lee
Kevin Lee
Kevin is a Co-Founder of ProductHQ. He has worked as a VC at Pear Ventures where he invested in and partnered with…
More About Kevin →
×

Feature flags let a team separate deploying code from exposing a capability to users. For a product manager, that separation can make releases safer and learning faster. It can support a gradual rollout, an experiment, an internal preview, a customer migration, or a quick rollback when a change behaves badly.

A flag is not a launch strategy by itself. It introduces a control surface that needs an owner, a purpose, a targeting rule, a measurement plan, and a removal date. Without those decisions, flags accumulate as invisible product complexity and nobody knows which behavior customers are actually receiving.

What I use feature flags for

I use a release flag when the team wants to deploy behind a controlled switch while operational checks continue. I use a rollout flag when exposure should expand in stages, such as internal users, a design-partner group, a small percentage, and then the broader audience. I use an experiment flag when variants and decision rules have been defined before exposure begins.

Other flags can support a migration between systems or temporarily protect a risky dependency. These purposes are different. A migration flag may need to stay active until data reconciliation is complete. An experiment flag should have a defined analysis window. Calling every switch an “experiment” makes governance and measurement weaker.

I write the purpose in the flag description and link it to the product decision or release record. Someone opening the control months later should understand why it exists and what evidence will allow us to remove it.

Define the release contract first

Before the first customer sees a flagged capability, I agree with engineering, design, analytics, support, and go-to-market partners on the release contract. It states who is eligible, what the old and new experiences are, what success and guardrail signals matter, how support should identify the state, and who can pause or reverse exposure.

I also ask what happens when the flag service is unavailable or the configuration is missing. The safe default depends on the capability. For a nonessential enhancement, the old experience may be the safer fallback. For a critical workflow, the team needs a tested behavior rather than an assumption.

A flag should not bypass normal quality work. Automated tests, accessibility review, security review, data checks, documentation, and operational readiness still apply. The switch changes exposure, not the standard for a trustworthy product.

Target users intentionally

Targeting is useful only when the audience definition is clear. I prefer stable, understandable attributes such as account, role, plan, geography where appropriate, or an explicitly assigned cohort. I avoid rules that a customer-facing team cannot explain or that change unpredictably as data changes.

For an early preview, I create an allowlist with a named owner and a way to remove an account safely. For a percentage rollout, I want assignment to be sticky so users do not see the experience change on every session. I document exclusions, especially when a capability has contractual, regulatory, or support implications.

I check that analytics and support can identify the active variant. A customer who reports a problem should not have to prove which experience they saw before the team can investigate.

Measure before expanding

I define the decision rule before looking at results. The primary signal should represent the intended customer outcome, not just clicks on the new interface. I add guardrails for reliability, errors, support contacts, cancellation risk, or other harms that a narrow success metric could hide.

I inspect results by meaningful segment and by exposure time. A rollout can look healthy overall while harming a small but important customer group. I also distinguish a flag-controlled release from a properly designed experiment. If assignment, sample, timing, or analysis is not suitable for causal inference, I describe the result as an observation rather than claiming the flag proved impact.

The next action should be explicit: expand, hold, revise, roll back, or retire. A dashboard that never changes exposure is monitoring, not a decision system.

Plan rollback and communication

I treat rollback as a product and communication decision, not merely a technical button. The team should know which behavior returns, whether customer data remains compatible, how support will explain the change, and what evidence triggers a pause. We practice the path for high-risk changes instead of discovering it during an incident.

Gradual rollout reduces blast radius, but it does not eliminate risk. A defect can spread through a targeting rule, affect shared data, or create confusing differences between teammates. I want logs of flag changes, approvals for sensitive controls, and clear access boundaries.

When exposure changes materially, I align release notes, help content, sales guidance, and customer communication. Hidden variation is especially damaging when a customer is trying to follow documentation that describes a different experience.

Prevent flag debt

Every flag gets an owner, creation date, purpose, current state, review date, and removal condition. I include cleanup in the delivery work rather than creating a vague future task. After the new behavior is stable, I remove the old code path, tests that only protect the old path, targeting rules, dashboards, and documentation that no longer apply.

I review the inventory regularly. Stale flags create branching combinations that are hard to test and can make incident diagnosis slow. A flag that has become a permanent configuration may belong in a supported product setting with clearer ownership and lifecycle rules.

Common mistakes

I watch for flags used to hide unfinished work indefinitely, flags with overlapping targeting rules, and experiments that change midstream without recording what happened. I also watch for release toggles controlled by too many people, missing audit history, and analytics that cannot distinguish exposure from eligibility.

Another mistake is using a flag as a substitute for a reversible product decision. If the team does not know what evidence it needs or what it will do next, the switch only postpones the hard conversation.

A practical flag checklist

Before release, I can answer: What is this flag for? Who owns it? Who can see the capability? What is the safe default? How will exposure be measured? What are the guardrails? What pauses or rolls it back? How will support identify the state? When will the flag and old path be removed?

Those answers turn feature flags from mysterious switches into a disciplined release practice. Used well, they give product teams room to learn while protecting customer trust. Used carelessly, they create a second product surface nobody maintains.

Next step

I use a spike for PMs when a technical uncertainty could change the product decision or the safest rollout path.

A controlled internal rollout can support dogfooding for PMs without treating employee access as proof that a broader launch is ready.

Feature flags can support controlled exposure; experiment velocity for PMs helps me connect staged delivery with a clear learning decision.

Build stronger product delivery, experimentation, and launch judgment in the Product Manager Certification. Subscribe to the Product HQ newsletter for practical frameworks and career-ready product lessons.

Kevin Lee
Kevin Lee
Kevin is a Co-Founder of ProductHQ. He has worked as a VC at Pear Ventures where he invested in and partnered with early-stage founders on product & growth to help them build the foundations of category-defining companies. He has worked as a Product Manager at AltSchool (backed by Andreessen Horowitz, Founders Fund, First Round Capital, Mark Zuckerberg, John Doerr and other exceptional investors). Previously, he was a Senior Product Manager at Kabam (acquired by NetMarble and Fox for a combined $1bn+), where he worked on products through all lifecycles in San Francisco, Vancouver, and Beijing and helped grow one of the company’s products to become the third largest revenue generating product in the company portfolio. In a former life, he worked in Technology Investment Banking at Merrill Lynch. He is also the author / co-author on 10+ gaming patents.