System-design interview · Extended interviews
Design a feature-flag and configuration platform
Publish validated configuration and evaluate deterministic local rules, while making rollout freshness, offline fallback and rollback behavior explicit.
You will learn to
- Build a complete local flag evaluator before adding a distribution service.
- Explain stable cohorts, compatible snapshots and monotonic installation.
- Distinguish authentic configuration, current configuration and actual rollout exposure.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Caching: cache hits, misses, write policies and invalidation · Replication and durability · Real-time communication: polling, long polling, SSE, and WebSocket · Production readiness: SLI, SLO, observability, and recovery
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Separate publishing a decision from making it
A feature flag changes application behavior without deploying a new binary. The control plane edits, validates and publishes rules; the evaluation path uses those rules on live requests. Keeping evaluation local can make it fast and available during a control outage, but means configuration may be temporarily old.
We use recommendation flag F7. It assigns a stable ten-percent tenant cohort to a new experience. Snapshot C17 is one complete saved configuration, and tenant 54 is the targeting key for the example. The flow is publish validated snapshot → distribute it → atomically install it in the application → evaluate trusted context → execute the branch → record bounded exposure information.
Clarify what the flag controls. Here it selects recommendations, not authorization, spending permission or an irreversible schema migration. A disconnected application may use the previous configuration for a bounded grace period, then must return false. An instantaneous universal kill switch would need a different online authority contract.
Start with a file and a deterministic evaluator on one instance. The engineering problem becomes interesting when many instances update at different times or operators publish incompatible rules. A dashboard containing an enabled boolean does not establish safe fleet-wide behavior.
02Functional requirements
Agree on what the service must do before choosing its components.
- Author typed flags. Support boolean, numeric, string and structured flags, with separate development and production environments.
- Target an experience. Select explicit users or tenants, percentage cohorts and multiple variants from trusted context; return the chosen value, reason and generation.
- Publish and roll back. Validate and audit drafts/publications. A rollback publishes a new generation containing earlier intended values rather than moving the generation backward.
- Inspect and retire flags. Show fleet adoption and bounded exposure information; retire obsolete configuration and application branches. The worked flag selects recommendations, not authorization, spending permission or irreversible schema changes.
03Non-functional requirements
Use these hypothetical requirements for the worked interview. Confirm the assumptions with the interviewer; the numerical targets require testing and are not measured results. For latency, p95 and p99 mean that 95% and 99% of measured delays, respectively, are no greater than the reported value.
- Workload. Plan for 20,000 application instances and one billion peak evaluations/s. An SDK is the application library that loads and evaluates configuration; its actual runtime cost must be measured.
- Latency and propagation. Target local p99 below 50 microseconds for bounded rules, publication p95 below one second and propagation to healthy connected instances within ten seconds p99. Publication success alone does not prove fleet adoption.
- Deterministic, coherent results. The same snapshot and trusted context must produce the same result. Related flags in one request use one immutable snapshot; a delayed old download must not undo a newer generation.
- Outage behavior. For F7, use the last validated configuration during a temporary disconnection. After five minutes without successful synchronization, return typed false with reason
stale_config. This local timeout is not a strict five-minute global disable guarantee, and other flags may need different fallbacks. - Authenticity, access and privacy. Authorize environment writers and protect confidential rules/context. A checksum detects corruption; trusted transport or a signature under a trusted key establishes authenticity. Neither establishes that an authentic old snapshot is still current.
- Bounded resource use. Limit rule complexity, targeting data, compilation and telemetry work. Exposure reporting must not block application evaluation, and old snapshots remain usable by requests already holding them.
04Make stable allocation and whole-snapshot replacement work
Randomly choosing ten percent on each request makes the same tenant bounce between experiences. Instead define a stable hash over an unambiguous tuple of rollout seed and targeting key. A hash deterministically maps input data to a numeric value. The seed is a fixed per-flag rollout value; keeping it stable keeps the same targeting key in the same cohort. In the example, bucket = hash(tuple(seed, key)) mod 10,000; modulo takes the remainder, giving a bucket from 0 through 9,999. Include buckets below 1,000 for ten percent.
Tenant 54 maps to 731 and therefore receives F7 = true. Increasing the threshold to 2,000 preserves that assignment if the seed, targeting identity and hash specification remain unchanged. Changing the seed deliberately reshuffles the cohort. Multiple variants need non-overlapping bucket ranges and an explicit assignment contract.
One instance loads complete C17, validates types and dependencies, compiles rules and stores an immutable active pointer. Each request captures that pointer, evaluates ordered explicit rules and then the percentage fallback, and returns diagnostics alongside the selected value. Related flags use the same retained snapshot.
A new file is parsed and compiled separately, then replaces the pointer atomically after validation. A malformed update leaves the prior version active. Updating individual entries in a shared mutable map could expose a combination never validated together. This baseline proves deterministic decisions and consistent replacement before introducing fleet distribution.
Each request retains a whole validated configuration.
Read each connection in order
- syncCompile then atomic installValidated C17 → SDK snapshot pointer
- syncAcquire one snapshotTrusted tenant context → SDK snapshot pointer
- syncC17 plus contextSDK snapshot pointer → Local rule evaluator
- returnValue, reason and generationLocal rule evaluator → Trusted tenant context
05Compare local work with distribution traffic
At 20,000 instances and 50,000 evaluations/s each, a remote call for every evaluation would create one billion network calls/s. Even a 100-byte combined envelope implies 100 GB/s before transport overhead and puts network latency on each application path. Local evaluation concentrates network work on relatively rare configuration updates.
| Work | Calculation | Meaning |
|---|---|---|
| Full fleet update | 20K × 100 KB = 2 GB | Distribute cacheable immutable snapshots |
| Five-second polling | 20K/5 = 4K checks/s | Conditional requests still consume capacity |
| Ten updates/hour | 2 GB × 10 = 20 GB/hour | Regional caches reduce origin load |
| Evaluator CPU at 2 microseconds | 1B/s × 2 μs = 2,000 CPU-seconds/s | Local is not free at fleet scale |
| One-in-1,000 telemetry sampling | 1M events/s | Aggregate locally before sending |
Bound rule complexity and precompile expensive structures outside requests. Large targeting lists and unbounded expressions can dominate evaluation. Retaining old and new compiled snapshots during rollout costs memory, but permits in-flight requests to finish consistently. Measure publication bandwidth, compilation work and telemetry separately instead of assuming low lookup latency solves the whole service’s cost.
06Give drafts, generations and context distinct meanings
Edit a draft
PUT /prod/draft
| Request field | Purpose |
|---|---|
| Expected draft version | Detect an edit made since the operator last read the draft. |
| Proposed rules | The configuration to validate and save. |
Publish a validated draft
POST /prod/publish
| Request field | Purpose |
|---|---|
| Validated draft identity | Select the rules to publish. |
| Expected active generation | Detect a competing publication. |
Draft versions protect editors from lost updates; publication generations order committed releases. An optimistic version conflict asks the operator to review newer state rather than overwrite it.
Local evaluation call
getBoolean(F7, false, trustedContext)
| Returned field | Meaning |
|---|---|
| Typed value | The selected boolean value for F7. |
| Reason | Why the evaluator selected the value or a fallback. |
| Variant | The selected variant, where applicable. |
| Generation | The publication used for this decision. |
Missing flags, initialization failures and type mismatches follow documented fallback behavior. Define attribute types, merge precedence and missing-value handling so different SDK languages cannot interpret the same context differently.
The targeting key identifies the assigned user or tenant. Here all members of tenant 54 share one tenant-based assignment. The application derives the active tenant from authenticated membership, not an arbitrary request field. Country or plan can be additional typed rule inputs.
Control records
| Record | Information retained |
|---|---|
| Draft | Editable rules and draft version. |
| Immutable compiled-source snapshot | Schema/compiler version, canonical content hash and dependencies. |
| Environment generation pointer | The committed active publication for that environment. |
| Audit record | The recorded edit/publication history. |
A publication changes the active pointer only after its immutable bytes are durable. Notifications are recoverable announcements of that pointer, not the only record that a publication occurred.
07Identify the limits of the one-instance design
The first limit is operational: many editors and environments cannot safely overwrite files without version checks, audit and validation. The second is distribution: 20,000 instances can miss notifications, download at different speeds or disconnect. Publishing centrally does not prove a particular application is already using the new value.
Consider app8 starting a C17 download. C18 arrives and installs first; C17 then finishes. “Last download completed wins” moves the instance backward. A genuine rollback to C16’s content must be published as generation 18, so the same ordering rule works for ordinary updates and intentional reversals.
A partial-map race is different: a request reads F7 from the new configuration and dependency F8 from the old one. Locking each flag independently does not give the request one consistent configuration. The request must retain one whole snapshot across related evaluations.
Finally, app9 is partitioned while an operator disables F7. No push protocol can deliver information across a broken connection instantly. The product must choose how long old behavior may continue and what follows. These failures motivate versioned publication, recoverable distribution, atomic installation and freshness policy as separate mechanisms.
08Add a control plane and recoverable delivery
Publication sequence
- Authorize and validate. Authenticate the environment writer; check rule types, bounds, missing references, dependency cycles and complexity.
- Store the snapshot. Make the complete immutable bytes durable.
- Publish conditionally. Commit the active-generation pointer and audit/publication event only if the expected generation still matches.
A scheduled publication revalidates its assumptions at activation rather than blindly applying an obsolete draft.
A stream announces that a newer generation exists. Instances fetch immutable content through regional caches and periodically poll the authoritative active pointer to recover missed announcements. The notification prompts a check. Cached bytes provide the configuration, but only the publication service can confirm which generation is current. This keeps distribution efficient without relying on uninterrupted streaming.
Installation sequence
- Verify. Check environment, schema, checksum and trusted origin/signature.
- Compile. Prepare the fetched rules outside the request path.
- Compare and install. Atomically replace the current pointer only when the incoming generation is newer; ignore older/equal generations.
- Handle a race. If another installer changes the pointer, retry the comparison against that current value.
Application requests continue using their captured references until completion. Cleanup waits for those references before reclaiming old compiled structures. The cost is temporary duplicate memory and careful concurrency. Local evaluation remains independent of telemetry: aggregate and bound exposure reporting so a metrics outage does not become a recommendation outage.
The publisher stores a verified snapshot before changing the active generation. Notifications prompt downloads through regional caches; polling recovers missed notifications. Each application request evaluates one installed local snapshot, without a network call to the control plane.
Read each connection in order
- syncSubmit validated changeEnvironment writer → Control API and publisher
- syncStore complete snapshotControl API and publisher → Immutable snapshots
- syncCommit generation and eventControl API and publisher → Active generation and audit
- asyncAnnounce committed generationControl API and publisher → SDK fetch and install
- syncPoll current generationSDK fetch and install → Active generation and audit
- syncFetch named snapshotSDK fetch and install → Regional snapshot caches
- syncFill missing bytesRegional snapshot caches → Immutable snapshots
- syncEvaluate installed snapshotLocal application requests → SDK fetch and install
09Choose a practical stale-configuration policy
A signed C17 can remain authentic after C18 is published. A successful download proves that a configuration is available, not that the fleet has installed the latest publication. Keep those observations separate in the SDK and the operations dashboard.
For this recommendation experiment, the practical policy is last-known-good operation during a short outage. Periodic synchronization reads the current environment generation and fetches a replacement when needed. Track successful synchronization with monotonic elapsed time. After five minutes without it, return false and report stale configuration. Cached bytes after a restart do not establish a new successful synchronization.
This policy reduces disruption without claiming an instantaneous kill switch. A request already using C17 can finish with its captured decision. Different application instances may briefly disagree during propagation. A delayed reply may leave an application unaware of C18. The five-minute local timeout therefore does not prove that every instance disables F7 within five minutes of publication.
If a feature must stop sensitive actions immediately, check a current online authority at that action boundary and accept its latency and outage behavior. A strictly bounded global stale-use promise requires additional timing and renewal rules. Keep that stronger protocol in the advanced discussion; ordinary recommendation rollout does not need it.
10Prove rollback and request consistency with C17 and C18
An operator validates F7 against draft 16 and publishes C17/generation 17. app8 fetches, verifies and installs it. Request req61 derives tenant 54 from trusted context, retains C17, evaluates bucket 731 below 1,000 and uses the enabled recommendation branch. It records an exposure only if the decision actually influenced the experience.
Now generation 18 contains the intended older disabled values. If its download finishes before 17, installation leaves 18 active and ignores the later 17 result. If 17 finishes first, 18 replaces it afterward. Both orders converge to 18. An existing request retaining 17 may finish consistently under 17, while the next request selects 18; the documented stale-configuration policy still applies.
If C19 is malformed or uses an unsupported schema, validation fails before pointer replacement and 18 stays active. If two publishers race, the expected-active-generation check permits one transition and gives the other a conflict. Restoring an earlier boolean value is still a new publication, so its generation must increase.
Exposure data names flag, variant and generation actually used. Assignment, evaluation and exposure are different observations. A debug evaluation or hidden branch should not automatically count as a user seeing the experimental experience.
Publication generation orders installation independently from flag values.
Read each connection in order
- asyncC18 is availablePublisher → Fetch worker
- syncValidate and install 18Fetch worker → Active pointer
- syncDelayed C17 arrives; reject olderFetch worker → Active pointer
- syncPin generation 18Application request → Active pointer
- returnEvaluate one coherent snapshotActive pointer → Application request
11Measure fleet adoption and bound configuration risk
Observe publication-to-install lag, active-generation distribution, time since successful synchronization, typed fallback rate, compilation failures, evaluation latency and outcome metrics by cohort. A healthy fleet average can hide an enabled cohort with errors. A rollback is not complete merely because the control API returned committed; inspect adoption and recovery of the intended product metric.
Use the same fixed input/output test vectors in every supported SDK language. Test seed and targeting-key migrations intentionally because either can reshuffle users. Protect environment credentials and audit production publications. Browser clients must not receive confidential rules or other tenants’ attributes simply to evaluate a flag; server-side evaluation can keep those details inside a trusted boundary.
Drill concurrent editors, reversed download order, malformed configuration, disconnection, restart with old bytes and delayed synchronization responses. Roll compatible code/data changes separately; a flag cannot undo irreversible writes or make an old schema valid for new code.
Retire flags after rollout by removing obsolete application branches as well as configuration. Experiments additionally need stable assignment, genuine exposure and statistical analysis. A percentage slider supplies controlled assignment, not proof that a treatment improves outcomes. Revisit remote evaluation when current central authority or confidential rule access outweighs local latency and offline continuity.
12Check the design against its requirements
Before closing, check the final design against the agreed requirements. FR means functional requirement and NFR means non-functional requirement; the numbers refer to the lists above. These are proposed validation checks, not test results.
| Requirement | Mechanism in the final design | Validation and remaining limit |
|---|---|---|
| FR 1, 3; NFR 3, 5 | Versioned drafts and validated immutable publications protect edits and releases. | Race two publishers, publish a type error and roll back values under a new generation. Verify authorization and audit records. |
| FR 2; NFR 3 | Stable seed/key hashing and one captured snapshot make decisions repeatable. | Run common SDK test vectors; tenant 54 stays in bucket 731. Reverse C17/C18 downloads and evaluate related flags during installation. |
| NFR 1, 2, 6 | Local compiled evaluation and cached snapshot distribution separate request work from updates. | Benchmark p99 evaluation, publication and fleet propagation at the assumed workload; include compilation and telemetry pressure. Arithmetic alone does not establish the targets. |
| NFR 4, 5 | Authoritative synchronization tracks freshness; the local outage policy returns a typed fallback. | Disconnect an instance, restart with old cached bytes and delay a response. Confirm the five-minute local policy without claiming an instantaneous global switch. |
| FR 4; NFR 6 | Bounded exposure reporting and adoption metrics show actual use and remaining old generations. | Lose telemetry, hide an evaluated branch and retire F7. Evaluation must continue, and a hidden branch must not be counted as exposure. |
13Rapid revision
Rehearse the numbered functional requirements and non-functional targets first. Use this table to recall the mechanisms, then close with the requirements check above.
Remember: Newer generation wins, not the last download.
| Prompt | Recall |
|---|---|
| Why local evaluation? | Avoid a network call per decision; state how long older configuration may be used |
| Why a stable hash? | The same seed, key and hash rules give the same user assignment |
| Why immutable snapshots? | A request evaluates related flags from one complete, validated configuration |
| Why increasing generations? | An old download cannot replace a newer release; rollback gets a new generation too |
| What proves authenticity? | Trusted transport or a verified signature; a checksum alone is insufficient |
| What confirms the current version? | Check with the environment’s publication service; cached bytes cannot answer that question |
| What happens offline? | Use the last configuration for the allowed interval, then return the declared typed fallback |
| What counts as exposure? | The variant was actually used to determine the user experience |
Close with: “I keep publication audited and evaluation local. Stable cohorts make rollouts repeatable, whole snapshots prevent mixed rules, and monotonic installation handles download reordering. A documented last-known-good policy and local outage timeout govern stale use. I accept bounded offline fallback and measure actual fleet adoption rather than calling a central write an instantaneous global switch.”
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
Why avoid a remote call for each flag check?
Reveal a model answer
At the assumed billion evaluations/s, network calls become enormous work and add a service dependency to application latency. Local compiled snapshots make lookups cheap and let requests continue briefly during control outages. That benefit requires an explicit stale-use and fallback policy rather than claiming instant global updates.
Interviewer follow-up
When would you choose remote evaluation?
Reveal the follow-up answer
When current authority or confidential rule access justifies the network and availability cost.
What the answer must demonstrate: At the assumed billion evaluations/s, network calls become enormous work and add a service dependency to application latency.
Why does expanding ten percent to twenty preserve tenant 54?
Reveal a model answer
Its bucket 731 remains below both thresholds 1,000 and 2,000 when seed, targeting key, encoding and hash algorithm stay unchanged. A new random draw each request would not preserve assignment. Tenant-based targeting gives all authorized users of that tenant the same cohort.
Interviewer follow-up
What does a seed change do?
Reveal the follow-up answer
It deliberately changes hash inputs and can reshuffle the cohort; treat it as a migration.
What the answer must demonstrate: Its bucket 731 remains below both thresholds 1,000 and 2,000 when seed, targeting key, encoding and hash algorithm stay unchanged.
Why is locking each flag separately insufficient?
Reveal a model answer
A request can still read F7 after its update and F8 before its dependent update, creating a configuration never validated together. Compile a complete immutable snapshot and retain one reference throughout related evaluations. Atomic pointer replacement changes later requests while old references safely finish.
Interviewer follow-up
When may old compiled data be freed?
Reveal the follow-up answer
Only after no active request or retained policy reference still needs it.
What the answer must demonstrate: A request can still read F7 after its update and F8 before its dependent update, creating a configuration never validated together.
C18 installs before a slow C17 download. What happens?
Reveal a model answer
The installer validates bytes then compares generations under an atomic update. Since 17 is older than 18, it is ignored. A real rollback also uses a higher generation containing the desired older values. Thus content reversal never requires reversing publication order.
Interviewer follow-up
What if two installers read the same old pointer?
Reveal the follow-up answer
Compare-and-swap lets one succeed; the other rechecks against the changed pointer.
What the answer must demonstrate: The installer validates bytes then compares generations under an atomic update.
Why cannot a successful C17 cache fetch renew freshness?
Reveal a model answer
A cache can return authentic C17 after C18 was published. Periodic synchronization must check the environment publication authority, not merely download old bytes again. The main design tracks synchronization and falls back after a long outage; it does not claim a strict publication-to-disable deadline.
Interviewer follow-up
Does successful synchronization prove every instance installed the current generation?
Reveal the follow-up answer
No. It describes that synchronization, not the entire fleet. Observe active-generation distribution and actual rollout adoption across instances.
What the answer must demonstrate: A cache can return authentic C17 after C18 was published.
What happens when an instance cannot hear the off switch?
Reveal a model answer
The application uses its last validated snapshot during a short outage, then returns the documented false fallback after five minutes without successful synchronization. That is an outage policy, not an instantaneous command or a proved global five-minute disable bound. Sensitive actions need their own current permission check.
Interviewer follow-up
What happens after restart with saved C17?
Reveal the follow-up answer
Require a new authority confirmation; do not reset old bytes to a fresh age.
What the answer must demonstrate: The application uses its last validated snapshot during a short outage, then returns the documented false fallback after five minutes without successful synchronization.
Is every evaluation an experiment exposure?
Reveal a model answer
No. The decision must influence the experience under the agreed exposure definition. Debug calls, hidden branches and requests that never render can evaluate without exposing a variant. Record actual flag, generation and variant used, and analyze outcome guardrails in addition to adoption.
Interviewer follow-up
Does a percentage allocation prove causal improvement?
Reveal the follow-up answer
No. It is one mechanism within an experiment requiring valid assignment, observation and analysis.
What the answer must demonstrate: No. The decision must influence the experience under the agreed exposure definition.
How do you validate a multi-language SDK rollout?
Reveal a model answer
Run identical fixed context/seed inputs against expected bucket and rule outputs in every language. Test type and missing-attribute behavior, snapshot installation order, offline expiry and restart. Observe generation spread and fallback rates during a limited rollout. Rule schema/compiler compatibility must be checked before activation.
Interviewer follow-up
Can a flag undo an incompatible database migration?
Reveal the follow-up answer
No. Data and code compatibility need their own staged migration and recovery plan.
What the answer must demonstrate: Run identical fixed context/seed inputs against expected bucket and rule outputs in every language.
Blank-page exercise · 45 minutes
Build the answer yourself
Roll F7 to ten percent of tenants, then roll it back while downloads reorder and one instance is disconnected.
- Agree the numbered functional requirements and non-functional targets, including flag purpose, evaluation load, propagation and the offline fallback.
- Calculate fleet checks and update traffic.
- Explain deterministic bucket assignment.
- Design draft/publication/context APIs.
- Prove whole-snapshot and generation races.
- Check typed behavior, deterministic snapshots, latency, propagation, privacy and outage fallback against the numbered requirements; distinguish authenticity, freshness and exposure.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a feature-flag and configuration platformapp8 installs C18, then its slow C17 download finishes. Which version stays active?Recall first, then reveal
C18 stays active because installation accepts only a newer generation. A rollback also gets a new generation, even when it restores older values. Publication, installation and user exposure are separate events.
Newer generation wins, not the last download.
Return to lessonDesign a feature-flag and configuration platformHow can rollback restore older values without accepting an old download?Recall first, then reveal
Design a feature-flag and configuration platformDoes downloading a cached snapshot prove it is current?Recall first, then reveal
No. Check the current generation with the publication service; cached bytes alone cannot show whether a newer version exists.
Ask which version is current; cached bytes cannot answer.
Return to lessonFinal revision
Summary and interview notes
Applications evaluate complete local configurations quickly, but may briefly disagree during updates. Stable user assignments, increasing generations and an offline fallback keep behavior predictable. Publishing a change does not mean every instance has installed it.
Remember these points
- Agree the numbered functional requirements and non-functional targets before designing components; validate the final design against them.
- Keep control and evaluation paths separate.
- Use stable targeting and documented hash semantics.
- Install complete newer snapshots atomically.
- Renew freshness only from current authority.
- Measure real cohort exposure and retire obsolete branches.
Interview tips
- Use the C18-before-C17 race to explain generations.
- Name the offline fallback before promising a kill switch.
Important qualifications
- Ordinary flags do not replace authorization.
- Experiment quality and schema compatibility need separate checks.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Atomic freshness renewal implementation
Expand locking and timing details when strict bounded stale use is challenged.
- SDK context merge and hook semantics
Cross-language integrations need a precise provider contract.
- Large snapshots and validated deltas
Useful only when full-distribution cost justifies more complex reconstruction.
- Multivariate experiment analysis
Allocation alone does not establish unbiased outcome measurement.
Technical references
- OpenFeature provider specificationStandard interfaces, typed defaults, resolution details, and provider lifecycle; it does not promise one hosting/freshness model.
- OpenFeature evaluation contextDefines targeting identity and contextual evaluation inputs.
- LaunchDarkly percentage rolloutsProvider-specific rollout allocation documentation; compare rather than assume its algorithm is universal.
Practice marks stay in this browser.