Your Feature Flag Bill Is a Cache Key Problem

A client exceeded their feature flag vendor’s monthly request quota by 100%.

The two fixes everyone proposed were delete the dead flags and stop calling flags inside loops. Both are sensible. Both save exactly zero requests.

That quota isn’t a line item on one service. It’s a single account-wide meter, and the whole application drains it: four backend JVMs, every worker dyno behind them, the browser, plus a weekly cron and a build-time CLI call nobody remembered were there. No individual surface looks guilty, which is a large part of why it stayed unfixed.

What actually drives the meter isn’t in any of the code anyone was looking at. It’s the product of three numbers, and not one of them appears at a call site.

The vendor here is Flagsmith, and the specific mechanics are theirs, but the failure mode belongs to any metered API you’ve put a cache in front of.

Java SDK 7.4.3, JavaScript SDK 9.1.0, Spring Boot 3.5. Client details are anonymised throughout; the SDK behaviour is all from published sources.


1. Both obvious fixes save nothing

When a usage graph goes red, the instinct is to reduce the thing you can see. You can see flags in the console and flag checks in the code, so you reduce those. Neither one is what the vendor counts.

“Delete the dead flags.” Flagsmith meters HTTP requests: /flags, /identities, /traits, /environment-document. It does not meter how many flags exist.

The part that surprises people is that no flag name ever crosses the network. Here’s the entire call path:

Flags flags = flagsmithService.getIdentityFlags(identity, traits);
boolean value = flags.isFeatureEnabled(feature.getValue());
  • Line one is the HTTP call. It takes an identity and its traits, and nothing else.
  • Line two picks your flag out of the response, in memory.
  • So one request returns every flag for that identity, and 47 flags cost exactly what 20 flags cost.

The SDK’s cache key agrees: "identity" + identifier, with neither the flag name nor the traits in it. If your client genuinely called once per flag, that key would have to say so.

“Stop checking flags inside loops.” The Java SDK checks a local Caffeine cache before every HTTP call, in FlagsmithApiWrapper.identifyUserWithTraits. Five hundred iterations over the same user cost one request, not five hundred.

That instinct isn’t worthless, mind you. Those loops were genuinely wasteful, just not in the way anyone thought: the flag service logs at INFO on entry and again on every evaluation, so a 500-iteration loop emitted around a thousand log lines. Real money, wrong vendor.

2. What the vendor is actually counting

Two things don’t move the meter at all.

Doesn’t matterWhy
How many flags you’ve definedone request returns all of them
How many times you check a flagafter the first, it’s served from a local cache

Three things are the entire bill.

TermWhat it countsWhat ours was
identitiesone request per distinct identityone per active user
3600 / TTLrefreshes per identity, per hour360
processeseach JVM caches independentlyweb dynos plus workers

Multiply those three and you have the invoice:

requests/hour  =  identities  x  (3600 / ttl_seconds)  x  processes

Two of the three were set once during setup and never looked at again. What follows is the four things that were wrong, in the order this expression makes them matter.


3. Finding one: the cache key is the user’s email

The mechanism. Here’s the call every developer on the team writes:

if (featureFlagService.isFeatureEnabled(Feature.NEW_MATCHING_ENGINE)) {

That asks nothing about any user. It’s a question about the deployment. And here’s what it becomes, inside a wrapper nobody knows is there, because every one of those call sites injects the interface rather than the class:

return featureFlagService.isFeatureEnabled(
    feature, principal.getEmail(), tenantSlug);

The SDK’s cache key is "identity" + identifier, with traits deliberately excluded. So the identifier is the user’s email, and the cache holds one entry per active user.

Cache entries, one tenant with three users

TODAY                              KEYED BY TENANT

"identity" + amir@example.com      "identity" + acme
"identity" + sarah@example.com  →  (same entry)
"identity" + paul@example.com      (same entry)

3 entries, 3 requests              1 entry, 1 request

A tenant with 300 active users generates 300 requests where one would do.

Nothing about that is specific to feature flags. The key holds something the answer doesn’t depend on, so every distinct value of that something costs another trip to whatever sits behind the cache.

The fix is one line: stop passing the email, pass the tenant. Most calls that reached the service directly were already doing exactly that. Only the wrapper was putting the user back in.

What it buys: backend requests get divided by the average number of active users per tenant, and each flag check drops a database lookup on the way.

What to check first. This makes flag values tenant-wide, and none of it is visible in your codebase, because it lives in the Flagsmith console. Three things break if you have them:

  • a segment rule targeting the email trait,
  • an override set on one specific user,
  • a percentage rollout, which is the nasty one. The engine picks who’s in the bucket by hashing the identity, so “10% of users” quietly becomes “10% of tenants, all or nothing inside each”.

4. Finding two: the TTL is ten seconds

The mechanism. The SDK’s own default expiry is five minutes. FlagsmithCacheConfig.DEFAULT_EXPIRE_AFTER_WRITE is 5, TimeUnit.MINUTES. This codebase set ten seconds: thirty times more aggressive, for data that changes when a human clicks a toggle in a web UI.

The cache uses expireAfterWrite, so an entry dies a fixed time after it’s written no matter how often it’s read. Refresh cadence is 3600 / TTL per identity, per JVM, and it’s completely insensitive to load.

Refreshes per identity, per hour, per JVM
(assuming continuous activity across the full hour)

10s   ████████████████████████████████████  360
60s   ██████                                 60      -83%
5min  █                                      12

The fix is a default and two properties:

flagsmith.cache.ttl-millis=${FLAGSMITH_CACHE_TTL_MILLIS:60000}
flagsmith.cache.max-size=${FLAGSMITH_CACHE_MAX_SIZE:1000}

Neither key existed in the repository, and neither was set in the production environment. Those are two separate facts, and only the second one tells you what production actually runs. Once the knob exists, moving to five minutes later is config rather than a deploy.

What it buys: up to 6x fewer backend requests, with no change to what any flag means.

“Up to” is carrying weight in that sentence, so here’s where the 6x comes from and where it doesn’t.

The saving depends on how long someone stays active compared to the expiry:

  • Somebody clicking around for ten minutes needs 60 refreshes at a 10-second expiry, and 10 at a 60-second one. That’s the full 6x.
  • Somebody who loads one page and leaves costs one request either way. That’s nothing.
  • Everyone else lands in between, and you don’t get to choose which.

There’s also a detail in the cache that makes even that optimistic. It reads with a plain getIfPresent and nothing else, so nothing coordinates a miss: the instant an entry expires, every concurrent request for that identity sees an empty cache and every one of them fires its own call. Twenty threads hitting the same user at the wrong moment is twenty requests, not one.

What it costs is how fast a flag takes effect. Flip a switch in the console and it now takes up to a minute to reach every process, where before it took up to ten seconds.

That matters if any of your flags are emergency off-switches. This one had several. One puts a two-factor prompt back in front of users. Three stop events being published into a processing queue. One swaps a read path back to the older version. Every one of those gets flipped by a person watching a graph who wants something to stop happening, right now.

Sixty seconds is still the right call for all of them, since none is a security control where ten extra seconds changes the outcome. That’s a judgement to make against your own flags though, not a rule to copy.

5. Finding three: the browser refetches everything on every click

Check the dashboard split by SDK key before you believe any of the backend arithmetic. The browser uses a separate key. If it carries the volume, findings one, two and four do nothing and this one is the whole story.

The mechanism. In the vendor-facing portal, every client-side navigation cost one POST /identities/. Same user, same flags, same answer.

click "Invoices"   ->  1 request   <- full re-init, refetches every flag
click "Profile"    ->  1 request   <- full re-init
click "Invoices"   ->  1 request   <- full re-init
browser back       ->  1 request   <- full re-init

The waste isn’t how many flags a page reads. One request returns the whole set, so reading flags after initialisation is free. The waste is the client being rebuilt on every screen, and the cause is referential equality: the JavaScript version of comparing two Java objects with != instead of .equals().

The provider re-runs its init effect whenever its dependencies change, comparing them by reference. One dependency was an array built inline:

vendorTenantSlugs={context.vendorInfos.map((v) => v.tenantSlug)}

The chain from there:

  • .map() returns a new array every time it runs.
  • The provider above it calls usePathname(), so it re-renders on every navigation.
  • New array, identical contents, different reference, so the comparison says “changed”.
  • The whole client gets torn down and rebuilt with a fresh identity call.
  • Client-side flag caching defaults to off, so there’s no local fallback to soften it.

The fix is to stop passing an array:

vendorTenants={context.vendorInfos.map((v) => v.tenantSlug).join(",")}

Strings compare by value, so "acme,bob" equals "acme,bob" and the re-initialisation stops. No useMemo, which matters: a memo would work here and would also be exactly the kind of defensive memoisation that draws a review comment asking whether it’s needed. Passing a value type instead of a reference type is the smaller idea and the better one.

What it buys: one request removed per vendor-portal navigation.

The detail that made this satisfying: the SDK was already joining that array into a comma-separated string internally before sending it, so the fix just builds the same string one level earlier, where the comparison can see it. And the main tenant app, refactored at the same time, passes plain strings for its equivalent inputs and was never affected. Same refactor, two apps, one of them quietly paying per click.

6. Finding four: the number still scales with the business

Everything above lowers a number that keeps climbing. More tenants, more hires, more dynos, and it’s back. Exactly one change alters what the number depends on.

The mechanism. Local evaluation changes who does the work. Today Flagsmith evaluates your flags and sends back the answers for one identity at a time, which is why the bill tracks how many identities you have. Under local evaluation it sends the rules instead: every flag, every segment, every override, in one payload, refreshed on a timer. Your app then evaluates any identity it likes without asking anyone.

┌──────────────────────────────────────────────────────────────┐
│  TODAY - Remote evaluation                                   │
│                                                              │
│   ┌──────────┐      "who is user Y?"      ┌───────────┐      │
│   │ your app │ ─────────────────────────> │ Flagsmith │      │
│   │          │ <───────────────────────── │    API    │      │
│   └──────────┘   every flag, already      └───────────┘      │
│                  evaluated for user Y                        │
│                                                              │
│   > Flagsmith does the evaluating and sends the answers      │
│   > 1 call per identity, per cache expiry                    │
└──────────────────────────────────────────────────────────────┘

┌──────────────────────────────────────────────────────────────┐
│  WITH LOCAL EVALUATION                                       │
│                                                              │
│   ┌───────────┐   the rules (env document)  ┌──────────┐     │
│   │ Flagsmith │ ──────────────────────────> │ your app │     │
│   │    API    │    every 60s, per process   │          │     │
│   └───────────┘                             └────┬─────┘     │
│                                                  │           │
│                                                  v           │
│      your app evaluates any identity, in memory              │
│                                                              │
│   > your app does the evaluating                             │
│   > 1 call per process, per refresh interval                 │
└──────────────────────────────────────────────────────────────┘

The fix is four lines of builder:

FlagsmithClient.newBuilder()
    .setApiKey(secretKey)
    .withConfiguration(
        FlagsmithConfig.newBuilder()
            .withLocalEvaluation(true)
            .withEnvironmentRefreshIntervalSeconds(60)
            .build())
    .build();

What it buys: backend cost becomes processes x 1440 requests/day, flat, whatever the user count. The environment document is itself metered, so this makes the bill predictable rather than free, which is the property you actually wanted. It’s the only one of the four that stops the problem recurring, and it makes findings one and two irrelevant for billing.

Two traps, and both live in the SDK’s source rather than its documentation.

Trap A: it refuses to start without a server-side key

Flagsmith has two key types. The client-side one is public and can only ask questions. Only the server-side key, prefixed ser., may download the rulebook, and FlagsmithClient.Builder.build() throws rather than degrading when local evaluation is on without one.

  • In production: fine, as long as every service runs a ser. key, not just the main one.
  • In tests: fake keys like flagsmith or test are everywhere, and every Spring context that loads the client now dies on startup.
  • In local dev: an empty key is allowed today, logging “all features disabled” and booting anyway. That becomes a boot failure.

The fix: enable local evaluation only when the key actually starts with ser., otherwise build the remote client you have now. One condition, one place. Tests and local dev keep working, no properties files change.

Trap B: the first download blocks startup, and its failure is swallowed

This is the one worth the price of admission.

The polling manager fetches the rulebook in its constructor, not in a start() method, and build() constructs it. So creating the client bean performs a synchronous HTTP round trip to the vendor during Spring context refresh. Then updateEnvironment catches RuntimeException and logs it.

Put those together:

  • The vendor is briefly unreachable while your app boots.
  • The app starts successfully. Nothing alerts.
  • The environment stays null, lookups throw, the service layer catches and returns false.
  • Every gated feature reads false across every process until the next poll.

Deploy time is exactly when this bites, because that’s when every dyno cold-starts at once against the same endpoint. You’ve turned a soft dependency into a boot-time one and made it fail closed, silently.

The reassuring half: the exposure is first boot only. Once a rulebook has loaded, the SDK never discards it, and a later refresh that fails keeps the copy it already has.

The fix: hand the client a set of default flag values, so a missing rulebook falls back to known answers instead of false. That’s real work rather than a config line, because those defaults now have to be maintained against your flag enum. If you already keep a defaults file around for tests, check it before you trust it: ours covered about two thirds of the constants it was supposed to, and nothing had ever failed to tell us.

7. The part that isn’t about feature flags

Three mistakes here, and not one of them is a bug. Every line does exactly what it was written to do.

  • A wrapper adds the current user to a call that never asked about a user, because somebody wanted per-user targeting and this was the tidiest place to put it.
  • A ten-second expiry gets picked during setup, when correctness feels more pressing than cost, then never revisited because it never breaks.
  • A .map() in a prop is the most natural thing to write in React, and costs nothing until something downstream compares it by reference.

What they share is that each one is invisible from the place you’d go looking. The cache key isn’t at the call site. The array identity isn’t in the component that renders it. And the price of an expiry doesn’t appear anywhere near the line that sets it.

None of it would have mattered if someone had asked the first question:

requests/hour  =  identities  x  (3600 / ttl_seconds)  x  processes

Three numbers. Two of them were set once during setup by someone with no reason to think about a bill, and nobody had ever multiplied them together. It’s also why the fixes interact: local evaluation skips the cache entirely, so it quietly cancels two of the other three.

That’s not a Flagsmith problem. Any metered API with a cache in front of it has this same expression sitting behind it, and most teams can’t tell you a single one of the three numbers. If a monthly bill ever surprises you, look at the cache key first.

For the other half of the caching story, the JPA fetch types piece covers what happens when a cache key you can’t see multiplies queries instead of API calls. And for a different flavour of “the fix made it worse”, there’s the connection pool ceiling incident.

Related Posts

The Production Outage — Where the Fix Was Worse Than the Bug

One missing property line.

Two production outages.

The hot fix between them made the second one inevitable.

Read more

Threads, @Async, @Transactional, and Virtual Threads: What Actually Happens Inside a Spring Boot Backend

A webhook fires. One HTTP request comes in.

Ten seconds later, half the app is returning 503s.

The bug is not in the webhook.

Read more

What Jackson Doesn’t Deserialize for Free: The Empty Array That Took Down a Production Endpoint

A 503. Every 30 seconds, on the dot.

Sentry was empty. The application logs said the request was fine.

Then we looked at what Jackson was actually parsing.

Read more