Rendered at 16:23:28 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
liampulles 4 minutes ago [-]
Like many "cloud native ideas" there is a core 20% to the idea of feature flags that is useful to 80% of teams, and then a remaining 80% of periphery concerns that are relevant to 20% of teams (and I'm erring generously on those figures).
Specifically, the idea of putting toggles on risky functionality is a good idea that is useful broadly. Being able to dynamically toggle feature flags without a restart, or apply some features to some segment of users is unlikely to be useful and entails a lot of complexity and potential footguns.
My approach to feature flags is even simpler than what is suggested here: feature flags are just a section in your config. Done. Even if you are confident that you need more advanced capabilities, you should not "advance" to that level until you have demonstrable proof that your team can manage simple config values for it.
jameshart 4 hours ago [-]
In my experience these (configuration based feature flags and feature flag services) are actually two complimentary capabilities that solve two completely different problems but that happen to share the same name: feature flags.
The config approach is critical for using feature flags as a software development lifecycle tool. It is how you manage having a codebase which contains the unfinished code for new feature x, but can still be deployed and pass all tests without feature x being turned on.
In this model you need a mechanism which allows a developer who is working on feature x to enable it for local testing, and for your CI system to be able to interact with the flag system to test that the application works in both states - with x turned on and off.
This is ideal for trunk based development models; feature branches are an alternative approach that doesn’t really benefit from this (indeed it adds complexity to working in feature branches).
Meanwhile feature flag services are to solve the problem that different people using the same software need different features turned on. That can be as simple as internal testers or beta users, it can be holding features to roll out in fixed update windows per tenant, or it can be part of a risk management strategy where features are rolled out through progressive exposure.
It can also be tempting to mix up your feature flags system with an A/B testing system - you can use a feature flag service to expose a feature to a test cohort and measure performance changes.
There can be reasons for doing that but it’s really important not to tie all these things together: not every development lifecycle change is an ab test hypothesis. Not every ab test hypothesis is a development lifecycle change (often it is really about testing changes in data, and feature config is just one piece of data you might want to change). Similarly some other data changes than code changes need to roll out progressively to mitigate risk.
So all these things might be feature-flag shaped, but that doesn’t mean you can substitute different feature flag solutions in and solve the same problems.
This post is saying ‘it’s okay not to have an exposure control solution; you can have a config file’ - which is obviously true, if what you have is a config management problem, not an exposure control problem.
jdwyah 4 hours ago [-]
Strongly agree that Config & FeatureFlags are two different things. But as I've tried building tools in the space: man is it difficult to really nail down what the difference is.
For Quonfig I landed on:
- Data model wise they are identical. Flags and configs can both be targeted. They can both use segments. They can both do partial rollouts. They can both have the same range of values (bool/number/duration/json/json-w-schema/etc)
- Its fine to use Flags for your Experiments/AB Tests, but that's just the "allocation engine". The rest of experimentation is the exposure tracking and goal tracking. Those ought live in your product analytics stack, because they are really just events. But its helpful to have a single place (flags) for the experiment allocation because then you can re-use segments for things like hold out groups / and just general targeting.
- The only real difference is that Flags are intended to be ephemeral and configs are intended to be permanent. A good UI should show you how long a flag has been alive and help you clean it up (or convert it to a config) if it's been true for everyone for too long.
- The use cases are different enough that it's worth keeping them in two separate UI.
scott_w 2 hours ago [-]
One way I’ve thought of it is to ask what happens if the function is blocked, how is it enabled mechanically and why would that be changed? Once you answer these, you start to see the groupings more clearly, I’ve found.
scott_w 2 hours ago [-]
I 100% agree with everything you said and more. I spent a significant amount of time at my job arguing constantly with the entire engineering department that “Feature Flags” are not all the same thing and had to push back on efforts to mix them all up.
Examples:
Experiment flags for A/B testing
Permission checks at a user level
Product flags for feature gating by plan
Rollout flags to launch a new feature gradually
Configuration for an account/system/feature
The number of times I had to push against the “just put it behind hasFeatureFlag(user, flag)” is more than I can count at this point! I think it comes from a misunderstanding of the DRY principle, to be honest.
esafak 1 hours ago [-]
Why would you not use the same service to handle them? Less is more.
hasyimibhar 23 minutes ago [-]
Having worked at a large company that misuses feature flag service for everything, here are some examples why you shouldn’t:
- one team uses feature flag for product gating. Feature flag service goes down. Users temporarily got locked out of the features they paid for.
- one team uses feature flag for dynamic pricing by leveraging targeting rules (how hard is it to write a bunch of if else in code?). It’s evaluated against all users, even if they are not active (for analysis reasons). Feature flag service charges by MAU. We have millions of users. Our feature flag service bill is now 6 digits per year.
- one team uses feature flag as literal json store instead of a proper db (god knows why). Someone updated the value but the “schema” is wrong. Shit breaks.
jdwyah 4 hours ago [-]
I am on my second feature flag startup, but I also somewhat agree with this.
Every project should have flags, but many projects need just the basics and a service is overkill.
Rolling your own JSON still feels like something we ought avoid though. Yes to start out it’s 95% booleans. But then you want a rollout. And then you want some targeting rules. And then you want non booleans… maybe some json. Oo wouldn’t it be nice if the json could conform to a schema… and then eventually you are like damn I really want to change these without deploying. Or you want to read the same flag from multiple services.
I’ve tried to incorporate this lowest common denominator into https://quonfig.com Use it totally free & open source as SDK, and it’s just loading JSON that you can track in git. Agents love it, hot reloads, SDK in lots of languages. But vs rolling your own you’ve got a lot of headroom on the design. A bunch of targeting operators. Segments etc. And then if you do want to get a nice UI / delivery network for real time updates, then you can use the paid side of things.
Your second startup dedicated to feature flags? Aren't these just some booleans with optional support to toggle them at runtime? Maybe I've just not been in the right domain, but I would have thought that if someone's use of feature flags is so complicated that they need a whole company to support it then they've massively overengineered their code.
gejose 2 hours ago [-]
The author literally spells out why this isn't always the case:
> Yes to start out it’s 95% booleans. But then you want a rollout. And then you want some targeting rules. And then you want non booleans… maybe some json. Oo wouldn’t it be nice if the json could conform to a schema… and then eventually you are like damn I really want to change these without deploying. Or you want to read the same flag from multiple services.
ambicapter 3 hours ago [-]
Launchdarkly and Statsig are both well-established companies that basically do feature flags+add-ons.
aleksiy123 3 hours ago [-]
You can do much more with feature flags like ab testing/experiments, integrations with analytics.
There can be whole UIs and tooling and infrastructure to manage around them and that’s what the sass offer
horizonwingtech 3 hours ago [-]
[flagged]
jjice 1 hours ago [-]
The biggest problem I've found with feature flags of this nature (mostly for internal testing and the like) is that they can easily live too long and bloat code. This is also very dependent on what the feature flag on/off code paths do.
1. For reducing outliving, I always create a ticket to remove the flag at the same time it gets added in. This way there's documented work that will get scheduled. YMMV depending on how your team does planning.
2. For flag branch implementations, it's often fine to do an if/else, but I've seen this blow up into a real headache. When I can, I like having two implementations of an interface that get swapped between. Code calls the interface as normal the the underlying implementation is the same. Works for React components and larger implementations/changes that already have an interface. Don't want to force an interface where it feels wrong.
retired 1 hours ago [-]
It might depend on your programming language but what worked for my team was to have a file with all the feature-flags as enums/constants. That way you have a clear view of how many feature flags you currently have in your code and by CTRL+clicking you automatically see all the code associated with that flag.
maxidog 4 hours ago [-]
I’m currently reverse engineering a large enterprise app and the feature flag bloat is truly astounding. Imho excessive feature flag complexity is a symptom of management who are indecisive and mistrusted by the developers.
not_a_bot_4sho 2 hours ago [-]
> Imho excessive feature flag complexity is a symptom of management who are indecisive and mistrusted by the developers.
In my experience, engineers aren't using them to account for managerial dithering, they're doing it for safe deployments and experiments and rollouts and such. A product with millions of users can easily have a tens or even hundreds of active switches at any moment (I'm assuming a large engineering team behind said product), and that's not necessarily a bad thing.
However, as someone else noted here, you absolutely MUST delete and clean up your flags/gates/whatever when you've completed that effort. That part can be tricky because not everyone has the discipline to pay off tech debt.
Usually, a flag/gate should not live in code for more than a few months. If it does, it should have robust justification.
Gigachad 4 hours ago [-]
You need to be actively deleting them after the feature has gone live
chasd00 31 minutes ago [-]
i could see it getting complex with a lot of flags, inevitable you'll run into a situation where more than one flag is combined.
then, down the road, you remove flag4. Hopefully the above would result in a compile error but you may be in a language or situation where a missing flag4 is interpreted as boolean false. That could cause all kinds of havoc in logical combinations like that are scattered all over the codebase. Even worse would be REST APIs retrieving flag values because who knows what's happening in that service code? Plus, you'd never know there was an issue until users start reporting missing or extra features showing up unless you have e2e tests for every possible combination of feature flags...
forgotaccount3 4 hours ago [-]
Create story to add feature flag controlling access to the feature.
Immediately create story to remove said feature flag controlling access to the feature and review it during backlog refinements.
Lucasoato 4 hours ago [-]
This of course depends on the quantity and the depth of these feature flags. It’s ok to hard code few of them, maybe it’s not when it becomes a practice and you have tens of them. I think you would have a much easier life if there was a standardized, well documented, way to define them.
I just wanted to say I haven't been doing C# dev until as of late, and I forgot how flipping amazing the tooling is. One of the things the C# compiler (which is used for IDE tooling) can do out of the box is it can compile multiple versions of the file (syntax, and symbol resolution included), meaning you can still make sure the statically disabled codepaths are still valid, and all the symbols are consistently available in every feature flag combo.
When something does go wrong, the compiler degrades gracefully, and still can make sense of most of the code.
MS gets a lot of flak (sometimes deservedly so), but they can sometimes also show what a large organization of skilled individuals working under consistent direction is capable of.
alexpotato 2 hours ago [-]
Speaking as a SRE/DevOps guy with almost 20 years of experience:
There are no solutions, only tradeoffs.
I'll give some examples:
Firm 1 had a single server running tron [0] for ALL scheduling of processes.
Pros: very easy to see what should run when and made audits a breeze
Cons: the central server died and it was a giant outage response to make sure that when the server came up it didn't start killing processes that should be running
Firm 2 used a management gui to create custom cron entries on each machine
Pros: each node had a local copy of the schedule and could keep going even if the central system died
Cons: each node had a local copy of the schedule which could "drift" from other nodes, a node could be forgotten etc
So, more generally, I agree that it's good to label things as "ok" so that we don't get into flame wars etc. That being said, the more important point is to say that if you are going to pick a strategy, do the work to support, build tooling and plan for outages related to that strategy.
gps372 4 hours ago [-]
Hardcoded feature flags becomes a big issue if your release process is long and error prone. Original article also suggest to enhance incrementally. But you don't want the interval between increments to be too big.
aksappy 4 hours ago [-]
Unclear how the arguments against feature flags can be used to justify hardcoding them or using a JSON file to manage them. Feature flag decisions are at the end of the day, tradeoffs to solve a problem that may occur for another person at a different place/time. If you are a solo developer, hardcode, use bytes or mail your feature flags - it does not matter. If you are not, there is a lot to take care of beyond a single hardcoded json file.
gwbas1c 3 hours ago [-]
> Hardcoded feature flags ... start with a simple JSON file
Any configuration read out of a JSON file is not hardcoded. Hardcoded means you need to recompile to change it.
(And yes, hardcoded, as in flags set with #define or equivalent, are totally fine depending on what you're doing.)
odyssey7 3 hours ago [-]
Depends on how you deploy. If changing it in prod requires a PR merge and a release approval, I’d consider that pretty hardcoded.
kalcode 3 hours ago [-]
That's a bit too rigid of a definition. Just means it requires a code update to change it, that's hardcoded.
PaulHoule 3 hours ago [-]
Brother
The blogspam marketing behind them is so strong
Yep, right up there with Triplebyte back in the day and the drumbeat about "it's so hard to bill customers" and "JWT sucks"
cassianoleal 4 hours ago [-]
Why wouldn't it be? The post does a good job explaining why not only it's not a problem, but also likely desirable in many (I'd wager the vast majority) of circumstances.
Simple, effective, cheap, easy to understand and manage. Not dependent on an external service, not dependent on a third-party.
drdexebtjl 4 hours ago [-]
To do safe deployments, you must have a way to run versions X and X+1 side by side serving an arbitrary proportion of your users while you inspect metrics.
That covers a lot of uses of feature flags without the bloat.
Xirdus 4 hours ago [-]
Hardcoding feature flags prevents you from deploying the same build with different flag sets side by side and comparing results. Depending on the nature of the project, it might be irrelevant or it might be a deal breaker. Especially in cloud environments, it's usually the latter. And everything is in the cloud nowadays.
jasonjayr 4 hours ago [-]
Do all modern compilers do tree-shaking at this point? Unless you use compiler-based flags, If you hard code them into your code, the dead code paths may still be in the shipped binary. Depending on what you are doing, the security issues mentioned in the article are still present, or you may prematurely reveal an upcoming feature.
andy_ppp 4 hours ago [-]
I've never found feature flag systems complex and you can build one yourself in a few functions. The services I have used in the past were extremely simple to integrate so I don't understand the issue either way honestly.
floki165 3 hours ago [-]
[flagged]
michael_luog 5 hours ago [-]
[dead]
devecycle-marke 4 hours ago [-]
[flagged]
mhog_hn 5 hours ago [-]
note that this is a "pre-november 2025 opus 4.5" post folks
Specifically, the idea of putting toggles on risky functionality is a good idea that is useful broadly. Being able to dynamically toggle feature flags without a restart, or apply some features to some segment of users is unlikely to be useful and entails a lot of complexity and potential footguns.
My approach to feature flags is even simpler than what is suggested here: feature flags are just a section in your config. Done. Even if you are confident that you need more advanced capabilities, you should not "advance" to that level until you have demonstrable proof that your team can manage simple config values for it.
The config approach is critical for using feature flags as a software development lifecycle tool. It is how you manage having a codebase which contains the unfinished code for new feature x, but can still be deployed and pass all tests without feature x being turned on.
In this model you need a mechanism which allows a developer who is working on feature x to enable it for local testing, and for your CI system to be able to interact with the flag system to test that the application works in both states - with x turned on and off.
This is ideal for trunk based development models; feature branches are an alternative approach that doesn’t really benefit from this (indeed it adds complexity to working in feature branches).
Meanwhile feature flag services are to solve the problem that different people using the same software need different features turned on. That can be as simple as internal testers or beta users, it can be holding features to roll out in fixed update windows per tenant, or it can be part of a risk management strategy where features are rolled out through progressive exposure.
It can also be tempting to mix up your feature flags system with an A/B testing system - you can use a feature flag service to expose a feature to a test cohort and measure performance changes.
There can be reasons for doing that but it’s really important not to tie all these things together: not every development lifecycle change is an ab test hypothesis. Not every ab test hypothesis is a development lifecycle change (often it is really about testing changes in data, and feature config is just one piece of data you might want to change). Similarly some other data changes than code changes need to roll out progressively to mitigate risk.
So all these things might be feature-flag shaped, but that doesn’t mean you can substitute different feature flag solutions in and solve the same problems.
This post is saying ‘it’s okay not to have an exposure control solution; you can have a config file’ - which is obviously true, if what you have is a config management problem, not an exposure control problem.
For Quonfig I landed on:
- Data model wise they are identical. Flags and configs can both be targeted. They can both use segments. They can both do partial rollouts. They can both have the same range of values (bool/number/duration/json/json-w-schema/etc)
- Its fine to use Flags for your Experiments/AB Tests, but that's just the "allocation engine". The rest of experimentation is the exposure tracking and goal tracking. Those ought live in your product analytics stack, because they are really just events. But its helpful to have a single place (flags) for the experiment allocation because then you can re-use segments for things like hold out groups / and just general targeting.
- The only real difference is that Flags are intended to be ephemeral and configs are intended to be permanent. A good UI should show you how long a flag has been alive and help you clean it up (or convert it to a config) if it's been true for everyone for too long.
- The use cases are different enough that it's worth keeping them in two separate UI.
Examples:
Experiment flags for A/B testing
Permission checks at a user level
Product flags for feature gating by plan
Rollout flags to launch a new feature gradually
Configuration for an account/system/feature
The number of times I had to push against the “just put it behind hasFeatureFlag(user, flag)” is more than I can count at this point! I think it comes from a misunderstanding of the DRY principle, to be honest.
- one team uses feature flag for product gating. Feature flag service goes down. Users temporarily got locked out of the features they paid for.
- one team uses feature flag for dynamic pricing by leveraging targeting rules (how hard is it to write a bunch of if else in code?). It’s evaluated against all users, even if they are not active (for analysis reasons). Feature flag service charges by MAU. We have millions of users. Our feature flag service bill is now 6 digits per year.
- one team uses feature flag as literal json store instead of a proper db (god knows why). Someone updated the value but the “schema” is wrong. Shit breaks.
Every project should have flags, but many projects need just the basics and a service is overkill.
Rolling your own JSON still feels like something we ought avoid though. Yes to start out it’s 95% booleans. But then you want a rollout. And then you want some targeting rules. And then you want non booleans… maybe some json. Oo wouldn’t it be nice if the json could conform to a schema… and then eventually you are like damn I really want to change these without deploying. Or you want to read the same flag from multiple services.
I’ve tried to incorporate this lowest common denominator into https://quonfig.com Use it totally free & open source as SDK, and it’s just loading JSON that you can track in git. Agents love it, hot reloads, SDK in lots of languages. But vs rolling your own you’ve got a lot of headroom on the design. A bunch of targeting operators. Segments etc. And then if you do want to get a nice UI / delivery network for real time updates, then you can use the paid side of things.
Local use description: https://docs.quonfig.com/docs/how-tos/open-source-local
> Yes to start out it’s 95% booleans. But then you want a rollout. And then you want some targeting rules. And then you want non booleans… maybe some json. Oo wouldn’t it be nice if the json could conform to a schema… and then eventually you are like damn I really want to change these without deploying. Or you want to read the same flag from multiple services.
There can be whole UIs and tooling and infrastructure to manage around them and that’s what the sass offer
1. For reducing outliving, I always create a ticket to remove the flag at the same time it gets added in. This way there's documented work that will get scheduled. YMMV depending on how your team does planning.
2. For flag branch implementations, it's often fine to do an if/else, but I've seen this blow up into a real headache. When I can, I like having two implementations of an interface that get swapped between. Code calls the interface as normal the the underlying implementation is the same. Works for React components and larger implementations/changes that already have an interface. Don't want to force an interface where it feels wrong.
In my experience, engineers aren't using them to account for managerial dithering, they're doing it for safe deployments and experiments and rollouts and such. A product with millions of users can easily have a tens or even hundreds of active switches at any moment (I'm assuming a large engineering team behind said product), and that's not necessarily a bad thing.
However, as someone else noted here, you absolutely MUST delete and clean up your flags/gates/whatever when you've completed that effort. That part can be tricky because not everyone has the discipline to pay off tech debt.
Usually, a flag/gate should not live in code for more than a few months. If it does, it should have robust justification.
something like enabledFeature = (flag1 || flag2 || flag3) && flag4
then, down the road, you remove flag4. Hopefully the above would result in a compile error but you may be in a language or situation where a missing flag4 is interpreted as boolean false. That could cause all kinds of havoc in logical combinations like that are scattered all over the codebase. Even worse would be REST APIs retrieving flag values because who knows what's happening in that service code? Plus, you'd never know there was an issue until users start reporting missing or extra features showing up unless you have e2e tests for every possible combination of feature flags...
Immediately create story to remove said feature flag controlling access to the feature and review it during backlog refinements.
When something does go wrong, the compiler degrades gracefully, and still can make sense of most of the code.
MS gets a lot of flak (sometimes deservedly so), but they can sometimes also show what a large organization of skilled individuals working under consistent direction is capable of.
There are no solutions, only tradeoffs.
I'll give some examples:
Firm 1 had a single server running tron [0] for ALL scheduling of processes.
Pros: very easy to see what should run when and made audits a breeze
Cons: the central server died and it was a giant outage response to make sure that when the server came up it didn't start killing processes that should be running
Firm 2 used a management gui to create custom cron entries on each machine
Pros: each node had a local copy of the schedule and could keep going even if the central system died
Cons: each node had a local copy of the schedule which could "drift" from other nodes, a node could be forgotten etc
So, more generally, I agree that it's good to label things as "ok" so that we don't get into flame wars etc. That being said, the more important point is to say that if you are going to pick a strategy, do the work to support, build tooling and plan for outages related to that strategy.
Any configuration read out of a JSON file is not hardcoded. Hardcoded means you need to recompile to change it.
(And yes, hardcoded, as in flags set with #define or equivalent, are totally fine depending on what you're doing.)
Simple, effective, cheap, easy to understand and manage. Not dependent on an external service, not dependent on a third-party.
That covers a lot of uses of feature flags without the bloat.