Why Consolidating Feature Flags and Experiments Cuts Hidden Costs
Feature flags and experiments do two different jobs. A flag controls delivery. It decides whether a feature is on for a given user, and lets you switch it off fast when something breaks. An experiment is about measurem
Feature flags and experiments do two different jobs.
A flag controls delivery. It decides whether a feature is on for a given user, and lets you switch it off fast when something breaks.
An experiment is about measurement. Split users into groups, show them different variations, and find out which performs better on a metric that matters.
Take a new checkout button. The flag delivers it: on for these users, off for those. The experiment measures it: does the new button convert better than the old one? Both halves describe a single idea — try a different checkout button and see if it helps — but because the jobs are different, teams run them in different places. Delivery lives in a flagging tool or a config service. Measurement lives somewhere else: a dedicated A/B platform, or a home-grown mix of an SDK and an analytics pipeline.
Describe that one idea in two systems, and keeping them in agreement about which users are in which group becomes your job, not the tooling's. That's the seam — the join between delivery and measurement — and it carries a cost in money and engineering time that's easy to overlook. Now that AWS AppConfig runs A/B tests natively on the same feature flags you already deploy, the seam is worth a second look. This post is about the hidden tax of keeping delivery and measurement apart, and what you get back when they share one control plane.
The multi-service tax
You pay in four ways.
Integration. Two systems to wire into your application, each with its own setup, authentication, and configuration. Two things to provision and keep correct across every environment.
SDK sprawl. A flag-evaluation SDK here, an experiment or analytics SDK there, multiplied across every language and every service you run. More dependencies to ship, more versions to keep current, more surface area when something breaks.
Operational overhead. Two control planes. Two sets of dashboards. Two permission models to manage and audit. Two things every engineer has to learn before they can safely change what customers see. And the gap between "what's flagged" and "what's being measured" becomes its own category of confusion.
Assignment drift. This is the subtle one. When a flagging system and an experiment system bucket users independently, keeping a single user in a consistent variation across both gets fiddly. A small mistake here quietly pollutes your experiment data. It's the worst kind of bug, because nothing errors — you just end up deciding on bad numbers.
None of these is dramatic on its own. Together they're a standing tax on every change worth measuring.
What changed
In mid-2026, AWS announced the general availability of experimentation tools in AWS AppConfig. A/B and multivariate testing are now built into the feature flags you already deploy, with no separate experimentation infrastructure to build or run. You define variations, target an audience with a rule builder, set the traffic split, ramp exposure, and promote the winner — all anchored to the flag itself — through the console, CLI, API, or CDK. It's available in all commercial AWS Regions.
The timing matters for a second reason. Amazon CloudWatch Evidently, AWS's dedicated A/B testing service, reached end of support on October 17, 2025. So "where does A/B testing live now?" has been a live question for anyone who relied on Evidently, not a hypothetical one.
The practical effect is simple: the flag that delivers a feature is now the same flag that runs the experiment measuring it.
What you get back
Walk back through the same four costs.
One control plane. The flag is the experiment surface. An experiment selects an existing feature flag, and each treatment is a variation of that flag's configuration. The thing that controls the rollout is the thing that runs the test, and defining a variation and its traffic split is just part of the experiment definition. Here it is in the CDK, where each treatment carries the flag values it sets and the weight that splits traffic:
from aws_cdk import aws_appconfig as appconfig
AttrValue = appconfig.CfnExperimentDefinition.AttributeValueProperty
Treatment = appconfig.CfnExperimentDefinition.TreatmentProperty
# app, env, and profile are the L1 CfnApplication, CfnEnvironment, and
# CfnConfigurationProfile for your feature flags, defined elsewhere in the stack.
# The checkout_button flag must already define a "color" attribute. Treatments
# set the value of an attribute the flag defines; they don't create new ones.
appconfig.CfnExperimentDefinition(self, "CheckoutButtonExperiment",
name="checkout-button-color",
application_identifier=app.ref,
environment_identifier=env.ref,
configuration_profile_identifier=profile.ref,
flag_key="checkout_button",
# Editor-syntax expression. $country is a caller-supplied context attribute,
# not an AWS Region; here, everyone in the US is eligible.
audience_rule='(eq $country "US")',
# Baseline: today's blue button, half the traffic. AppConfig generates the
# treatment key; the Agent returns it as "_variant". The control comes back
# as "__control__" and treatments as "__t1__", "__t2__", and so on.
control=Treatment(
enabled=True, weight=50,
attribute_values={"color": AttrValue(string_value="blue")},
),
# Variation under test: green button, the other half.
treatments=[
Treatment(
enabled=True, weight=50,
attribute_values={"color": AttrValue(string_value="green")},
),
],
)
The CfnExperimentDefinition construct is only in recent versions of aws-cdk-lib (2.268.0 or later).
One agent. Experiments are delivered through the AWS AppConfig Agent, which runs on Amazon EC2, AWS Lambda, Amazon ECS, Amazon EKS, and on-premises servers. If you already use the agent to evaluate flags, that's one component per runtime to ship and keep current, not two. If you currently call the AppConfig API directly, budget for adding the agent: experimentation requires it.
One workflow. Define your variations, target the audience, set the traffic split, then start at 0% exposure and ramp up while watching your metrics. When a treatment wins, you promote it through a standard AppConfig safe rollout — the same gradual, monitored deployment you'd use for any other configuration change. One detail worth getting right: update and deploy the flag with the winning values while the experiment is still running, then stop it. In that order, users go straight from the split to the full rollout. Stop first, and the flag reverts to whatever is currently deployed (usually the default) until you redeploy.
Consistent assignment by construction. The system doing the flagging is the system doing the bucketing. There's no second service to disagree with, so "did both systems put this user in the same group?" stops being a question you can get wrong. It isn't zero work. You retrieve a treatment with the same agent call you already use for flags, plus two headers: an Entity-Id that identifies the user, and a Context that carries the attributes your audience rule evaluates. But given a stable entity ID, the agent returns the same treatment for that entity for the life of the run — consistency is the system's job, not yours. And because treatment allocations lock once a run starts, the split can't quietly shift underneath you partway through.
Where the savings stop
Consolidation covers delivery, assignment, and exposure control. Analysis is still yours. Configuring an analytics platform to capture experiment data is an explicit step in the workflow, and AppConfig is deliberately agnostic about which one: CloudWatch, Snowflake, Datadog, or your own warehouse. You're unifying where experiments run, not how you measure them. That's the right trade — your metric definitions are usually the last thing you'd want to migrate — but it does mean the analytics side stays with you.
It also isn't free. AppConfig experimentation is pay-as-you-go, billed per experiment-run hour from the moment a run starts until you end it, and standard AppConfig configuration-request charges still apply on top while the experiment runs. (See the pricing link below for current rates.) The argument isn't that consolidation costs nothing. It's that you stop paying twice for one idea.
When consolidating isn't the answer
This isn't a case for ripping out a system that's serving you well.
If you've already invested in a dedicated experimentation platform — one your analysts have built tooling around, with its own statistics engine, holdout groups, and a workflow people rely on — the math is different. AppConfig's AI-assisted experiment design validates your setup against experimentation best practices at design time, which helps you reach sufficient statistical power, but the analysis itself happens in your own tools, and very advanced programs may still want a platform that does both.
The same goes if your flags are evaluated in a browser or mobile client. The documented path runs through the agent on server-side compute, so client-side experiments mean your backend does the assignment and hands the result down. That's a design decision, not a blocker, but one to make deliberately.
The value of consolidation is highest when your second system is mostly overhead — when it exists only because flagging and experimentation historically couldn't share a home. The case is strongest for teams already running AppConfig feature flags, where a separate experiment system duplicates flags they already have.
A quick gut check
Four questions to size up your own tax:
- Do you deliver a feature in one system and measure it in another, describing the same change twice?
- How many SDKs or agents does a single feature touch as it evaluates end to end?
- When an experiment produces a winner, can you promote it through your existing deployment safety net, or is that a separate manual step?
- Is keeping a user in a consistent variation across systems something you actively babysit?
The more "yes" answers, the more there is to reclaim.
Why this matters
For teams building on AWS, the split between delivering a feature and measuring it used to be structural. AppConfig flagged, something else experimented, and stitching the two together was the cost of doing business.
It isn't structural anymore. It's a choice — and consolidating saves more than dollars: fewer moving parts, less drift, and one safe path from testing an idea to shipping the winner. It's the kind of cost that's easy to keep paying precisely because no invoice ever arrives for it, which is exactly why it's worth adding up on purpose.
Further reading
- AWS AppConfig experimentation (User Guide)
- AWS AppConfig launches managed experimentation tools for A/B testing
- Example experiment workflow
- Promoting a winning treatment
- AWS AppConfig pricing (AppConfig appears on the AWS Systems Manager pricing page; experimentation is billed per experiment-run hour)
- Experiment-run pricing model, explained (see the "Experiment run" definition)
- Support for Amazon CloudWatch Evidently ending soon
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.