Authentication & authorisation
Who are you, and what are you allowed to do? Two questions, not one.
The shape of the decision
Read this diagram as text
- Identity provider
- Authenticates and issues tokens. May be yours or a federated third party.
- Client
- Holds a short-lived token and refreshes it, never the password.
- Gateway
- Verifies the token signature and rejects anything unsigned or expired.
- Service
- Checks authorisation per action, against the actual record. This is the check that matters.
- Data
- Reached only after both questions have been answered.
- Client → Identity provider (1 · authenticate)
- Identity provider → Client (2 · token)
- Client → Gateway (3 · request + token)
- Gateway → Service (4 · verified identity)
- Service → Data (5 · authorised access)
- Identity has to be established once and then trusted across every surface (web, mobile, third-party integrations) without each of them reimplementing it, and without a password ever travelling further than it must.
- An identity provider authenticates and issues a short-lived token. Services verify the token rather than calling back for every request. Authorisation is separate and evaluated per action against roles or policies, at the point where the resource is actually touched.
- Stateless tokens are fast precisely because nobody checks in, which means revoking one before it expires needs a mechanism you have to build. Short expiry plus refresh is the usual answer, and it adds a rotation path that has to work when the network does not.
- Do not put authorisation only in the gateway or only in the UI. The gateway knows who is asking, but not whether this particular record belongs to them. Hiding a button is not a permission check.
Sessions or tokens
I treat the choice between a session and a token as a question about revocation before it is a question about anything else, because almost everything else the two are compared on, throughput, statelessness, how well either fits behind a load balancer, can be approximated by the other with enough engineering effort, and revocation timing cannot. A session is a reference: the browser holds an opaque identifier that means nothing on its own, and the server holds the actual state that identifier points to, whether the user is signed in, what they are permitted, when they last acted. Revoking a session means deleting that one record, complete the instant the delete commits, because nothing else in the system holds a copy of the state that could disagree with it. A token carries the state itself, signed so it cannot be altered without detection but not secret, so any service holding the right key verifies it without asking anywhere else whether it is still good. That is what makes a token fast to check, and it is exactly what makes a token impossible to revoke on its own terms: once issued, it stays valid until it expires, and checking its signature can never tell a verifier that the issuer changed its mind a moment ago.
Sessions win wherever the infrastructure to check them cheaply already exists and the number of places that need to check it stays small. A single application, or a handful of services that already share one database or one fast cache, pays almost nothing extra to look up a session on every request, and in exchange it gets revocation that works the way a user expects: sign out on this device, force a sign out everywhere, invalidate on a password change, all of them a single write away from taking effect. The lookup that tokens exist to avoid is not expensive here, because the store being checked is already close to everything doing the checking.
Tokens earn their place once the number of independent verifiers grows past what a shared session store can serve cheaply: several services running in different regions, a public API called by clients outside the organisation, a mobile app that needs to keep working through a patchy connection without a round trip to a central store on every request. Here a signature check performed locally, with no network call to anywhere else, is the entire reason to choose a token over a session, and giving that up by adding a central check on every verification defeats the choice at the point it was made. The revocation cost tokens carry is worth paying specifically because the alternative, every verifier calling home before trusting anything, reproduces the session model's dependency on one shared store while losing the session model's clean revocation in exchange for nothing.
Revocation is the hard part
Short lifetimes paired with a refresh step are the default answer because they turn an unrevokable token into one with a bounded exposure instead of an unlimited one. The access token stays short-lived and fully stateless, verified locally on every request the way tokens are meant to be. The refresh token is longer-lived and is checked against server state, but only at the much lower frequency of minting a new access token rather than on every request, so the system pays a server round trip a small fraction as often as it would if every request needed one. Revoking access, in this design, means revoking the refresh token, and the access token already issued from it simply runs out on its own within whatever window its short lifetime allows.
A bounded window is not the same as immediate, and some situations, a token known to have been stolen, an account an administrator needs to lock out this moment, need the shorter answer anyway. A denylist supplies it: a store of identifiers for tokens that have been revoked but have not yet reached their own expiry, checked against every incoming token before its signature is trusted. That check is the exact server round trip stateless verification was designed to avoid, now paid on every request instead of only the revoked ones, which is why I keep a denylist small and short-lived on purpose, holding only tokens that have not yet naturally expired, rather than letting it grow into a permanent record of every revocation the system has ever issued.
Signing a user out everywhere is a different problem from revoking one token, because there is no single token that represents a user's access. Whichever device or integration authenticated first holds its own token, independently valid until its own expiry, and a request to end all of them at once has to reach every one without a list of exactly which tokens currently exist. The workable answer scopes revocation to the user instead of to any individual token: a counter or a timestamp stored on the user's own record, checked on every verification against whatever value was current when the token was issued. A token issued before that marker moved fails verification no matter which device it came from, which invalidates every session for that user in one write instead of one token at a time. It is a coarser tool than revoking a single session, since it cannot leave one device signed in while ending the rest, but that coarseness is what makes sign out everywhere achievable without keeping a permanent, ever-growing table of every token the system has ever handed out.
Role models that survive growth
Role-based access starts as the right model for almost everything, because it reduces a large set of individual permission questions to a much smaller set of named bundles, and auditing who can do something becomes a question about who holds a role, not a search through individual grants scattered across the system. The failure mode shows up as roles multiplying to capture every exception, a one-off grant for a single customer, a temporary elevation for one person covering someone else's work, until the count of roles approaches the count of people they were meant to simplify. That growth is a signal worth reading correctly: a specific permission has outgrown what naming a role can express, which argues for a different mechanism on that one permission, not for abandoning roles across a system where most permissions are still simple enough for them.
Attribute-based rules become necessary once a permission genuinely depends on a relationship between the caller and the specific resource that no fixed role can name in advance: allowed to edit this document if you are its owner, or you share a team with its owner, and it has not yet been published. Modelling that in roles alone means a role for every combination the relationship could produce, which is exactly the multiplication the previous paragraph describes. An attribute check evaluates the relevant relationship at the moment of the request instead, against whatever the caller and the resource currently are, rather than against a role assigned in advance that could not have anticipated every combination. I reach for an attribute check on the specific permission that genuinely needs one instead of replacing the role model everywhere for consistency, because most permissions in an application are not relational at all and gain nothing from a heavier mechanism applied uniformly.
Where the check happens matters as much as which model produces the answer. A single point invoked at the entry to every handler that touches a protected resource, a guard, a middleware, a policy evaluation called the same way everywhere, means a permission rule is written once and applied consistently regardless of which endpoint is being called. Leaving each handler to decide independently whether its caller is allowed means every new endpoint has to remember to add that check on its own, in whatever form the person writing it thought to use, and a missing one stays invisible until someone finds the specific route that never got it.
This is where the caution against stopping at the gateway earns its detail. A gateway can confirm who is making a request; it was never positioned to know whether the specific record being requested belongs to that particular caller, because that fact lives in the resource itself, not in the identity token the gateway checked. The boundary that decides has to sit at the point the record is loaded, where the caller and the row being touched are both available to compare, and a check placed anywhere else, in the gateway alone, in a UI that simply hides a button from someone lacking a role, is a convenience for a legitimate user rather than a barrier to anyone willing to call the endpoint directly.
Authorisation in a multi-tenant system
Everything the previous section says about checking a specific record still applies once a system serves more than one tenant, except that ownership now has an extra dimension sitting underneath role and relationship: which tenant the record belongs to at all. A user can hold a legitimate role, be the genuine owner of records within their own tenant, and still be issuing a request for a record that happens to belong to a different tenant entirely, distinguishable from their own only by an identifier the user never sees. Checking role and relationship without first confirming tenant answers a question that was never being asked.
The discipline that closes this is scoping every query that touches tenant-owned data by the tenant identifier at the query itself, as an explicit filter joined into the lookup, rather than as a check applied to whatever row the lookup happens to return. Loading a record first and rejecting it afterwards once its tenant fails to match still means the record was fetched: pulled out of the database, held in memory, potentially logged or captured in a stack trace, before the rejection was ever decided. A query scoped by tenant from the start never returns the row at all, which closes the leak at the point that matters, the moment before the data leaves the database, not the moment after it has already arrived in the application and is merely being discarded.
This fails quietly wherever a record's own identifier is unique across the whole system instead of scoped to a tenant, because a lookup by that identifier alone still runs and still returns a row for most inputs anyone tries while testing, since the identifiers a developer reaches for during testing usually belong to their own tenant in the first place. The bug is invisible until someone tests with two tenants deliberately, which is rarely the first thing anyone thinks to try. I default to composing the key used for lookups from the tenant identifier and the record's own identifier together, rather than trusting every call site to remember an additional filter, because a lookup that cannot be constructed without a tenant fails immediately in every environment, a mistake caught the first time it is written instead of a discipline that only holds for as long as everyone remembers to apply it.
A control enforced at the database itself, restricting which rows a connection can see to those matching the tenant it has been told about, gives a second layer beneath the application code that a forgotten filter cannot bypass, because the restriction applies before the query's own logic runs instead of depending on the query having remembered to include it. It sits alongside deliberate scoping at the application layer, not in place of it, catching the one query that skipped that discipline despite everyone's best effort. A system serving more than one tenant is worth building with both layers together rather than treating either alone as sufficient.
Let’s talk
Have a product, platform or delivery challenge? Let’s talk about turning it into a structured, scalable solution.
Open to technical leadership, product delivery and senior engineering roles, and available for architecture consulting, technical reviews and mentorship. Engagements run as project-based work, contracts, consulting, freelance engagements, remote collaboration and long-term partnerships.
Based in Cairo, Egypt, working remotely with clients across the MENA region and internationally.