Authentication verifies an identity. Authorization decides whether a principal may perform an action on a resource. Isolation enforces the boundaries between users, tenants or workloads. Authentication alone does not authorize access to every object a service can reach.
A principal is an acting identity, such as a person, service or delegated worker. A tenant is a customer or organizational boundary. An access-control list (ACL) associates a resource with permissions for principals or groups. One tenant can contain many users with different ACLs.
Interview scope: private team learning material
A learning platform stores shared study guides, private interview recordings, mentor feedback and organization-level reports. A learner can read their own recording; a mentor needs an active sharing relationship; an organization administrator cannot automatically read every private recording.
Functional requirements
- Sign in people and authenticate service-to-service calls.
- Grant and revoke resource-specific access for learners, mentors and teams.
- Carry trusted identity through retrieval, generation, exports and delayed jobs.
- Recheck authority before sensitive reads and actions.
- Explain allow/deny decisions through protected audit records.
Non-functional requirements
- Deny access when required identity or policy information cannot be established.
- Define a measurable revocation bound for each path, including cached results.
- Keep policy checks within a stated latency budget and monitor their availability.
- Prevent a tenant parameter, guessed ID or model output from granting authority.
- Rotate credentials and deploy policy changes without silently expanding access.
Start with one enforcement path and a small permission matrix. Add role or relationship complexity when the product requires it; do not start with an unexplained universal administrator role.
Choose the authorization model
| Model | Definition | Learning-platform example | Tradeoff |
|---|---|---|---|
| RBAC | Permissions are assigned to roles; principals receive roles | Curriculum editor can publish guides | Simple capabilities; many object-specific roles become difficult |
| ABAC | A policy evaluates subject, action, resource and environmental attributes | Mentor may view an active assignment in an allowed region | Expressive; attributes need authoritative update paths |
| ReBAC | Access follows defined relationships among principals and resources | Recording owner shares it with a particular mentor | Natural sharing model; relationship changes and traversal need care |
These models can be combined. A role may grant the ability to review, while an assignment relationship limits which recording may be reviewed. Establish default-deny behavior and a consistent enforcement point. OWASP authorization guidance.
Identity protocols are not interchangeable
| Mechanism | Purpose | Common mistake |
|---|---|---|
| OAuth 2.0 | Delegated authorization framework | Treating an arbitrary access token as proof of an end-user identity |
| OpenID Connect | Authentication layer built on OAuth 2.0 | Accepting an ID token as an access token for any API |
| JWT | Claims representation that can be signed and/or encrypted under its profile | Decoding a token and trusting its fields without verification |
| Server session | Server-managed login state referenced by a client credential | Assuming logout clears every copied token or worker credential |
| API key | A credential identifying/authorizing a caller under the service contract | Giving all keys the same broad permission |
Use an established identity implementation. For an interactive authorization-code flow, apply PKCE and the protocol's transaction-binding/CSRF protections, exact redirect handling and token validation. Do not build a new password or token protocol for the interview application. Sources: OpenID Connect Core, OAuth security BCP.
For a JWT accepted by this API, enforce its intended profile: trusted issuer and keys, configured algorithms, audience, required time claims and token type. Bind issuer and subject to an application account; apply current session/revocation policy. Use configured key discovery, not an arbitrary key URL supplied inside an untrusted token. A valid signature does not make a token intended for another service acceptable. JWT BCP.
Separate policy decision from enforcement
The policy decision point (PDP) evaluates an access request. The policy enforcement point (PEP) prevents the operation unless the decision permits it. They may be library functions in one service or separate components; the vocabulary does not require a remote microservice.
Read diagram source
flowchart TD
C[Caller credential] --> I[Validate and derive trusted principal]
I --> E[API enforcement point]
E --> P[Policy decision using current roles and relationships]
P -->|Deny or unavailable| D[No protected operation]
P -->|Allow scoped read| R[Authorized source and index query]
R --> M[Model receives permitted evidence]
M --> T[Proposed action without self-assigned authority]
T --> X[Executor checks current policy and business state]
X --> S[Scoped operation and protected receipt]
E --> A[Audit decision metadata]
X --> A
The trusted principal travels outside generated text. A model may select a candidate recording ID, but the executor verifies the caller may access that recording. A service identity with broad database access must enforce the delegated user's narrower authority where it acts on that user's behalf.
Policy combination must be explicit
A policy engine cannot return on its first allow if its declared combination rule gives later denies precedence. That bug turns rule ordering into privilege escalation. Different engines have different combination rules; deny-overrides is a policy choice, not the definition of ABAC. AWS IAM provides a concrete system where explicit denies override applicable allows, subject to its documented policy combinations. IAM evaluation rules.
This self-contained exercise combines already evaluated decisions for a deliberately small deny-overrides policy. Real engines must also evaluate scope, conditions, errors and obligations correctly.
def combine(decisions):
known = {"allow", "deny", "not_applicable"}
if any(d not in known for d in decisions):
return False # Includes unavailable/indeterminate results.
return "deny" not in decisions and "allow" in decisions
assert combine(["allow", "not_applicable"])
assert not combine(["allow", "deny"])
assert not combine(["deny", "allow"])
assert not combine([])
assert not combine(["allow", "error"])
Protect all four paths
- Retrieval: derive tenant and permission constraints server-side. Enforce them before content is exposed to a model or user. Trusted retrieval code can revalidate candidate IDs against the authoritative ACL; do not send forbidden text to a reranker or trace before that check.
- Actions: authorize the resource and operation, then enforce current business prerequisites. Reading a recording does not authorize publishing it.
- Caches: use the required identity/permission scope and content/policy versions. Revalidate or invalidate on access changes; a long TTL is not a revocation strategy.
- Artifacts and telemetry: protect stored answers, exports, memory, recordings, traces and download URLs. A private page with a public attachment is still an exposure.
A shared vector index can be appropriate when mandatory filtering and revalidation meet the requirements. Separate collections or databases reduce some shared risks but add routing, migration and operational cost. Physical separation does not itself prove that the application routes each user correctly.
In PostgreSQL, superusers and BYPASSRLS roles bypass row security; table owners normally do too unless forced to obey it. Test with the actual application role and correctly scoped connection state. Pool reuse must not carry a previous request's tenant context. PostgreSQL row security.
Two kinds of API key
| Key | What the application needs | Storage approach |
|---|---|---|
| Client key your service issues | Verify a presented high-entropy secret | Keep a verifier, key ID, owner, scopes, expiry and status; return the secret once |
| Upstream provider key | Recover a secret to authenticate to the provider | Restrict recoverable secret storage and inject it into the authorized executor |
A one-way hash cannot be forwarded as the original provider credential. Random API tokens and human passwords also need different storage considerations: passwords require a suitable slow password-hashing scheme, whereas a cryptographic verifier can protect a sufficiently random token.
The lifecycle is issue → validate → rotate → revoke → audit. Use secure randomness, safe comparison, expiry/revocation checks and scoped access. Rotation may use a bounded overlap while clients migrate; a suspected compromise may require immediate revocation. Invalidate validation caches and inspect usage by key ID. Never log raw keys or let a caller choose another tenant's secret reference.
Revocation and the stale-approval trap
Suppose a mentor loses access while an export job is queued. Store a trusted principal and proposal reference, then check current access at execution. A queue message is not permanent authority. Standing delegation can authorize routine work, but only within its scope, validity and resource conditions.
An approval, when required, is tied to a particular operation and version. If the target, amount or recipient changes, the old approval cannot authorize the changed operation. Rechecking permission does not prevent a duplicate effect; use a separate idempotency and durable-execution design.
For a hypothetical 60-second revocation target, a five-minute positive-decision cache is insufficient on its own. Use prompt invalidation/version checks or authoritative revalidation on protected paths, and measure propagation including offline workers. Existing downloaded data cannot be recalled merely by changing an ACL. Short-lived signed URLs bound future access only according to their expiry and validation contract.
Test the matrix and close the design
| Request | Expected result | What it proves |
|---|---|---|
| Learner reads their own recording | Allow | A legitimate path remains usable |
| Learner guesses another recording ID | Deny | Object-level checks work |
| Mentor reads a currently shared recording | Allow | Relationship policy is applied |
| Mentor repeats the request after revocation | Deny within the stated bound | Caches and workers honor changes |
| Tenant admin reads another tenant's trace | Deny | Local administration is not global authority |
| Worker changes the approved export target | Deny | Delegation binds the intended operation |
| Required policy service fails | Deny protected operation; report unavailable | Failure does not silently grant access |
Audit principal, tenant, resource, operation, policy version, decision, time and operation ID. Restrict audit access and retention; hashes of predictable payloads are not automatic privacy. Compare latency, availability, revocation speed and operational burden when deciding between local evaluation and a remote policy service.
Interview tip: Follow one permission change through the index, cache, saved answer, download URL and background worker. This reveals more than naming RBAC and drawing a login box.
Interview questions with developed answers
Q1: How do you implement multi-tenant isolation in a RAG system?
Sample answer: I derive tenant and user identity from authenticated server state and enforce the access policy at the data boundary before content reaches the model. I preserve document-level permissions, scope caches and conversation history, and protect saved artifacts and asynchronous jobs. Database or index isolation can use separate resources or a carefully enforced shared design, depending on scale and requirements. I test cross-tenant queries, role changes, revoked access, and privileged service roles. An output filter is useful defense in depth, but it is too late to be the primary control for forbidden content already supplied to the model.
Follow-up: How do you test background workers? Execute jobs after revocation and with tampered identity references, and verify that access is denied.
Q2: How do you manage API keys for an LLM service?
Sample answer: I distinguish client keys we issue from upstream provider secrets we must use. Client keys can have stored verifiers, scopes, expiry, and revocation checks. Provider credentials need protected recoverable storage because the service sends them to the provider. Both need least-privilege access, rotation, auditing, and redaction from logs. Credentials should not be embedded in prompts or client code. I would also separate environments and service identities so a development or tenant-level credential cannot become a production administrator credential.
Follow-up: What changes during a compromise? Revoke or rotate promptly, inspect usage, and address the path that exposed the key.
Q3: When would you choose RBAC versus ABAC?
Sample answer: RBAC works well for stable groups of capabilities, such as billing administrators and support readers. ABAC expresses conditions involving resource, user, and request attributes, such as region or refund amount. I choose a policy that is understandable, testable, and consistently enforced, often combining roles with resource attributes. More expressive rules can become harder to reason about, so policy tests and clear ownership matter. The model may propose an action, but it must not decide its own authorization by interpreting a natural-language role description.
Follow-up: Who maintains the attributes? They need authoritative sources and update paths, or the policy can use stale facts.
Q4: Why must approval be checked again at execution time?
Sample answer: Approval records a decision about a particular proposal, while authorization and business state may change before execution. The user could lose their role, the order could already be refunded, or the proposal amount could change. I bind approval to the exact proposal and version, set expiry where appropriate, and revalidate current permissions and prerequisites before acting. This avoids treating a historical approval as unlimited future authority. Duplicate approval events also need safe handling so one decision cannot trigger repeated independent actions.
Follow-up: Does rechecking permission replace idempotency? No; permission and duplicate-effect prevention solve different problems.
Q5: What should an authorization audit record contain?
Sample answer: It should identify the principal, tenant, requested action, target resource, policy version or decision context, decision, timestamp, and related task or approval. It should explain a denial or allow decision sufficiently for investigation without copying unnecessary sensitive content or credentials. I control access and retention for the audit store and test that critical paths actually emit events. The record supports accountability, but it does not prove that the policy itself was correct; policy review and negative tests remain necessary.
Follow-up: What if a privileged database role bypasses row policies? Restrict that role and test the effective permissions used by the application.
60-second interview answer
I separate authentication—who the caller is—from authorization—what they may do. The application derives tenant and user identity from trusted credentials, then enforces access at retrieval, tool execution, caching, and storage. A model cannot grant itself a role or choose an arbitrary tenant ID. Permissions need to propagate through background jobs and long-running agents, with revalidation before sensitive actions. I distinguish service API keys we verify from upstream credentials we must retrieve securely, and I test negative cases such as revoked access and cross-tenant cache hits.