2026-10-01
Identity and Access · Part 3 of 4
Building It: An Order of Work for Identity and Access
The implementation guide: which boundary to build first, how to keep the identity provider replaceable, how to model policy as data, how to check the door's documents on entry, and what the lifecycle of a grant actually requires.
Platform & Infrastructure · Identity and Access, Part 3 of 4 · ~13 min read · prev: How a Session Crosses Many Applications · next: Operating and Integrating It · assumes Parts 1–2: the four words, the five layers, the four flows
Part 2 described the machinery. This post is the order to build it in, and the reason for that order: each step makes the next one reviewable. Build the door before the policy and you will wire permission into the application. Build the application before the session and you will invent a second login.
The order below is the one we would follow again — with three honest exceptions, at the end.
Boundaries Before Features
What this shows: the build order — each step makes the next one reviewable.
If you have one week, and a provider that already signs people in, build steps 4 and 5 for your highest-risk application and stop: a check that verifies sessions and refuses anonymous readers closes the exposure we see most often — the page nobody meant to publish. If you have one month, add steps 1–3 and 6, so a person signs in once and reaches several applications. If you have a quarter, finish, and spend the last part of it on step 7 — because a door you cannot diagnose will consume the time you thought you saved.
Notice what is not in the list: a login form. The form is the smallest part of the work — and the part most teams start with.
The same order works as an audit of a layer you already have. Walk it once and note the first step your system cannot answer — that is where the next piece of work is, whether you built the system or inherited it.
The Identity-Provider Port
The first decision is not which provider, it is what shape the application sees. Put a port between your code and the provider: a narrow contract with three operations — start a sign-in with the provider, verify a session against the provider's published keys, administer identities. The vendor lives behind an adapter that implements the port, and nowhere else.
What this shows: the port is the contract; the adapter is the only place a vendor appears — which is what makes the provider a decision rather than a commitment.
This is the abstraction principle from Why Software Design Principles Are What Make AI-Assisted Work Reliable applied where it pays best. The cost is one interface and one adapter. The return is that replacing the provider becomes a bounded change instead of a rewrite — and that the rest of your code, and every test of it, never mentions the vendor at all.
The same principle applies inside the provider choice itself: prefer the parts of a provider that are shaped by standards — a standard sign-in, standard tokens, standard claims — over the parts that are shaped by its SDK. Standards are the only interface a future provider is guaranteed to speak.
One detail belongs in the port rather than in the adapter: the set of issuers your boundaries accept is configuration, not code. A trust decision is then a deployment variable, so a second identity source can be trusted — or a compromised one removed — without rebuilding anything.
The Authorization Model
Authorization answers "may this principal do this to that?", and there are three well-worn ways to answer it.
RBAC — role-based access control — grants permissions by role: an operator, a reviewer, an administrator. It is simple to reason about, simple to audit, and the right place to start. Its failure mode is proliferation: when roles multiply to express exceptions, nobody can say what a role means any more.
ABAC — attribute-based access control — decides from attributes: department, region, clearance, time of day. It expresses rules roles cannot, and its failure mode is the opposite one: policy becomes a programming language nobody can read, and an access review becomes an exercise in interpretation.
ReBAC — relationship-based access control — decides from relationships: this person is an editor of this document, and the document belongs to that project. It is the right model when access is genuinely relational, and it is the most expensive to operate.
Start with RBAC. Move to attributes or relationships only when you can name the question that roles cannot answer — because each move buys expressiveness with auditability.
Whatever the model, two decisions about placement keep it operable: where a request is refused, and where the decision is made.
What this shows: enforcement at the boundary, decision from data — the separation that lets you answer "why was this allowed?" months later.
The enforcement point (often abbreviated PEP) sits in front of the resource and refuses without knowing the policy; the decision point (PDP) reads policy and returns a decision without touching the request. Keep them separate and an access question becomes a query: which grant allowed this, granted by whom, valid until when. Merge them and the answer is a code reading.
Policy as data is the second of those two decisions. A grant should be a record with, at minimum: who it is for, what role, on which resource, in which tenant, when it starts, when it ends, whether it is revoked, who granted it, and when they did. Nine fields that turn "can we remove this person's access?" from a deployment into a row.
And one rule that saves an entire class of incident: authority does not belong in the token. A token states identity and expiry; embedding permissions in it creates a second copy of your policy that cannot be revoked before it expires. The token proves who; the store decides what.
The Boundary
The boundary is the component that verifies every request — the gateway for a page, each application's own service for its data. Its rules are few and non-negotiable.
Verify on every request. A previous page load proves nothing about this one, and a boundary that trusts a browser's claim of having been checked is not a boundary.
Deny by default. An unknown surface, a missing credential, an API path nobody has bound, an application nobody has declared — the safe answer is always "no". Fail closed means that when the boundary cannot reach a decision — the policy store is unavailable, the verification cannot complete — it refuses rather than admits. An outage becomes an access problem instead of a breach.
Keep the credential off the static origin. The bundle that serves files must never receive the session cookie, and must never be able to set one. A file server that holds a credential is a credential store with no access control.
Answer specifically. Every refusal should say which question failed and what would fix it. That message costs a few words at each refusal point, and it repays them the first time someone else is on call.
The Entry Check in Front of Applications
A browser application is a bundle of files, so the entry check belongs in front of it — as a layer that decides before the file is served. Three details separate an entry check from a decoration.
Check documents on entry, serve assets. The page needs a session; the stylesheet does not. Refusing assets breaks the very sign-in page you redirected the reader to.
Carry the return address. The redirect to sign-in must remember the page the reader asked for, and the sign-in must send them back to that page once it has checked that the return address is one it is allowed to use. Anything else strands people on a default screen and gets reported as "the login is broken".
Refuse before the page is sent. A check that serves the page and then hides it in the browser has already served it: anyone with the address can read the file without the script that hides it. The check is only worth having if the file never leaves your server.
Multi-Tenancy and Machine Identity
Two extensions arrive as soon as the platform is used by more than one customer, or by anything that is not a person.
Tenants are isolated slices of one shared system. The design requirement is not "filter by tenant", it is that a decision for one tenant is never made with another tenant's data — which means isolation enforced at the store (row ownership and keys), not at the caller's discretion.
Machine identity is the account an application uses to act on its own behalf. Treat it as a person with a name, a role and an expiry: a service account with its own least privilege, its own grants in the same policy store, and its own audit trail. "The application is trusted" is not an access model — it is an unbounded grant with no name.
The Lifecycle Nobody Plans For
Access is not granted once; it moves. The states below are the whole operational surface of an identity layer, and each one needs an operation rather than a procedure in someone's memory.
What this shows: a grant's life — and that its endings, expiry and revocation, are where access systems are actually judged.
Three distinctions matter more than the diagram. Disabling a login is not ending a session: turning off a person's ability to sign in leaves every session they already hold alive until it expires. If the requirement is "access ends now", the design needs a way to end live sessions, not just a way to block the next sign-in. Revocation is state, not deletion: keeping the revoked grant (with its timestamp and its actor) is what makes the audit trail answerable, and deleting it is what makes the same question unanswerable next quarter. Expiry is a feature: a grant with an end date limits the blast radius of the one nobody remembers to review.
Around those sit the operational habits an auditor will ask about: onboarding as a single reviewed grant rather than an accumulation; access reviews that read the store rather than a spreadsheet; and key rotation for the signing material your boundaries verify against. Then there is the break-glass path — a documented, loud, time-boxed way in when the normal path is broken — because the alternative is an undocumented one.
Vendor Independence in Practice
The exit drill is one question: "You must replace the identity provider next quarter. What changes?"
If the answer is the adapter, the credential exchange and the claims mapping, your coupling is right. If the answer includes the policy store or the boundary's contract, the coupling is wrong — and it is worth fixing before you need it, because the moment you need it is the moment you have least room to move.
What is worth owning is the part that is genuinely yours: the policy model, the door, the audit trail. What is worth buying is everything commodity: password resets, breach detection, the mechanics of a corporate directory. And what is worth refusing is an abstraction that exists only to look architectural. Wrapping a standard token verifier in three layers "for portability" does not make it portable; it makes it three layers deeper. The question to ask here is the one from Part 1: where does an abstraction earn its keep?
If You Need One Sign-In for Applications You Don't Build
Everything so far assumes you control both ends: the applications you build accept the session, and you decide what it may do. The moment a requirement arrives for one sign-in across software you do not build — a partner's portal, a supplier's tool, a customer's own system — the layer changes role. You stop being only the party that accepts identity proofs and become the party that vouches for them — in the standards' word, the issuer: applications you do not control trust an assertion your layer signs.
That is federation, and it is additive rather than a different design. It adds:
- An issuance surface. An OIDC or SAML endpoint other applications redirect to, plus the metadata they need in order to trust it.
- A client registry. Every application allowed to ask for an assertion becomes a record with an identifier, an allowed return address and a set of claims — the same records-not-code approach as the policy store, one row per application.
- Consent and release. Which attributes leave the estate, to whom, and on what basis. Inside your own applications the question does not arise; outside, it is the whole question.
- An ending that travels. Signing out of your estate should also end the other application's session, where the other side supports it — that has to be arranged across the boundary, not done locally.
- Provisioning. Where the other side keeps its own accounts, identity has to be pushed there and removed there: the grant lifecycle described above, performed in someone else's store.
What this shows: federation is the same layer in a second role — the identity it already proves is asserted to applications you do not control.
None of that is needed to run your own estate, and building it speculatively is how a layer doubles in size before it has a second consumer. It is also where the support surface grows: an issuer is depended on by people who cannot see your logs, so availability, key rotation and change notification become promises rather than internal matters.
What makes it a phase rather than a rewrite is that each part rests on something already built. The session is a standards-shaped token, so a stranger can verify it. The issuer set is configuration, so a second identity source needs no rebuild. Keys are the estate's own, with a rotation policy and a published key set. Applications are records, so a client registry is a second table rather than a second model. Assurance is a claim, so "how strongly was this person proved?" can be asserted rather than assumed. And tenancy is explicit in policy, which is the precondition for serving more than one organisation from one issuer.
What we deliberately did not build is the other half of the same decision: no issuance endpoint, no client registry, no consent screen, no provisioning agent, no separate identity store per tenant (one issuer serves every tenant, with isolation enforced in policy and data instead), and no published service commitment. Each is real work with an operational tail, and each waits for a requirement that names it.
What We Would Sequence Differently
Three decisions cost us more than they should have.
We treated the identity provider as architecture rather than as a dependency. It was not a vendor SDK in the domain — but it was close enough that replacing it would have touched more than the adapter. The fix was the port described above, and the lesson is that "we chose a provider" and "our code depends on a provider" are different statements, only one of which is reversible.
We built the refusals before we built their diagnosis. Every check worked, and none of them said which layer had spoken, so the first real incident was investigated by elimination rather than by reading a response. Instrumentation is the last item in the build order above and the one we would move earlier: a refusal that cannot be traced is a defect in the design, not an operational inconvenience.
A third, dated 2026-10-02: we shipped the door and its first operator application as one bundle, so the sign-in and an application shared a single origin. That was the wrong order: an application belongs behind the door, never beside it, and unpicking the two afterwards cost a week.
Part 4 is the operational half of that lesson: how to trace a refusal through the layers that can produce one, and what a new application must do to join the estate.
The Orchestrator's Takeaway
- Boundaries before features. Provider port, policy model, policy store, boundary, entry check, session, instrumentation — each step makes the next reviewable, and the login form is the smallest part of the work.
- The provider is a dependency, not the architecture. One port, one adapter, standards-shaped underneath, and the exit drill answered in a sentence.
- Policy is data with an end date. Grants as records with validity and revocation; authority out of the token; enforcement at the boundary and the decision from the store.
- The endings are most of the work. Disabling a login, ending a session, revoking a grant and expiring one are four different operations, and an access system is judged on how well it performs them.
- Federation is the next phase, not a fork. Issuing assertions to applications you do not control stays additive when the session is standards-shaped, the issuer set is configuration, and applications are records.