Skip to main content
Ahmed Salama

Architecture Lab

Queues & async work

Take the slow work out of the request the user is waiting for.

The shape of the decision

System diagram: Queues & async work: Take the slow work out of the request the user is waiting for.Accepts the request, enqueues the work, responds immediately.SERVICEAPIDurable buffer. Absorbs spikes that would otherwise become timeouts.QUEUEQueueConsume and retry independently, scaled to the backlog rather than to traffic.SERVICEWorkersWhere a job goes after it has failed enough times to need a human.QUEUEDead letterThe slow or unreliable dependency the user should never wait on.EXTERNALThird partyenqueueconsumecallafter N retries
Queues & async work: Take the slow work out of the request the user is waiting for.
Read this diagram as text
APIService
Accepts the request, enqueues the work, responds immediately.
QueueQueue
Durable buffer. Absorbs spikes that would otherwise become timeouts.
WorkersService
Consume and retry independently, scaled to the backlog rather than to traffic.
Dead letterQueue
Where a job goes after it has failed enough times to need a human.
Third partyExternal
The slow or unreliable dependency the user should never wait on.

Connections

  • API → Queue (enqueue)
  • Queue → Workers (consume)
  • Workers → Third party (call)
  • Workers → Dead letter (after N retries)
The problem it solves
A request that sends three emails, generates a PDF and calls a payment provider is as slow as the sum of those, and it fails if any one of them does.
How it works
The request writes a job to a queue and returns. Workers consume jobs independently, retry on failure with backoff, and scale separately from web traffic. A spike becomes a longer queue rather than a timeout.
What it costs
The system becomes eventually consistent and considerably harder to reason about. You now need idempotent jobs, a retry policy, a dead-letter queue, and a way to tell a user that something they triggered has not finished yet.
When not to reach for it
The user genuinely needs the result before the response; you cannot queue an authorisation check. And a queue relocates a slow operation rather than fixing it. If the work is slow because it is badly written, it will be just as slow in a worker.

The contract a queue actually gives you

The contract a queue gives you by default is at-least-once delivery: a message is delivered one or more times, never zero, provided the queue itself stays up. That 'at least' is doing real work in the phrase. A consumer can crash after processing a message but before acknowledging it, the queue then assumes the message was never handled and redelivers it, and the same message can be processed twice by design rather than by accident.

Ordering is a separate promise, and most queues that offer it only offer it within a narrow scope: messages from a single producer, or within a single partition, or on a single queue with one consumer. The moment work fans out across multiple consumers, or across partitions chosen by a key, global ordering is gone, and a system that assumes messages arrive in the sequence they were sent will occasionally see some other sequence once consumers process at different speeds. Anything that depends on strict ordering needs a mechanism that survives that, a sequence number checked at the consumer, rather than a hope that the queue keeps things in line.

Exactly-once delivery is the claim I treat with the most suspicion, because delivering a message exactly once across a network that can drop, duplicate and reorder packets is a guarantee no queue can make entirely on its own. What some systems offer is closer to exactly-once processing, achieved by pairing at-least-once delivery with an idempotent consumer or a transactional write that only commits once. The guarantee lives partly in the queue and partly in how the consumer is written, not in the queue alone, and a team that reads 'exactly-once' on a product page and skips the idempotency work on the consumer side has bought half a guarantee while assuming the whole one.

It is worth separating the delivery guarantee from the durability guarantee. Durability answers whether a message survives a crash of the queue itself; delivery semantics answer how many times a surviving message reaches a consumer, and a queue can be extremely reliable about the first while still being only at-least-once about the second. Treating 'durable' and 'exactly-once' as the same claim is a common shortcut, and the gap between them is exactly where duplicate-processing bugs live.

Idempotency is the consumer's job

If at-least-once is the honest default, the obligation it creates falls entirely on the consumer: a handler has to produce the same outcome whether a message arrives once or five times, because the queue has already said it might be the latter. Writing a handler as though each message were unique and unrepeated is building on a promise the queue never made.

An idempotency key is the direct mechanism: a unique identifier carried on the message, checked against a record of keys already processed before any side effect runs, so a redelivered message is recognised and skipped rather than reprocessed. The key has to be chosen by whoever creates the message, from something genuinely stable across redelivery, an order id or a request id generated once at the edge, rather than anything computed fresh inside the handler, or every redelivery mints a new key and the check never catches a duplicate.

Some writes are naturally idempotent without any key at all, and recognising them matters because they need none of this machinery. Setting a record's status to 'shipped' produces the same end state whether it runs once or three times; the second and third runs simply confirm what the first one already did. A write that instead increments a counter, or appends a row, has no such property, and redelivering it silently double-counts or double-inserts unless something external stops it.

The practical habit is to prefer the state-transition shape wherever the domain allows it, an absolute assignment rather than a relative adjustment, because it survives redelivery without any extra mechanism. Where the operation genuinely has to be additive, a payment captured, an item appended to a list, that is exactly where an explicit idempotency key earns its cost, because nothing about the operation's own shape will protect it.

Dead letters and poison messages

A retry budget is a decision made in advance about how many attempts a failing message deserves before the answer changes from try again to stop and tell someone. Without that decision made explicitly, a message that will never succeed, a poison message whose payload the handler cannot parse or whose referenced record no longer exists, retries endlessly, consuming a worker slot on every attempt and pushing every message behind it further back in the queue.

Backoff between attempts matters as much as the count. Retrying immediately, at the same interval every time, is the shape most likely to hit a dependency exactly while it is still recovering from whatever caused the first failure, turning a downstream outage into a self-inflicted retry storm layered on top of it. Increasing the interval between attempts, and capping how far it grows, gives a struggling dependency room to recover instead of a queue hammering it at full speed throughout.

The dead-letter queue is where a message goes once its retry budget is spent, and the mistake I see made most often is treating it as storage rather than as a signal. A message sitting in a DLQ represents a specific failure that retrying did not fix, which means retrying alone was never going to fix it either, and it needs a person to look at why, not a scheduled job that quietly requeues the whole DLQ back onto the main queue overnight.

The DLQ is worth alerting on directly, at a low threshold, because a handful of messages landing there is diagnostic and cheap to investigate, while a large batch landing there overnight because of one bad deploy is the same failure discovered a day late. I would rather treat any message reaching the dead-letter queue as a question worth answering immediately, what changed, what do this batch of messages have in common, than as an appendix to be cleared out on a schedule.

The test for staying synchronous

The test is simple to state and easy to skip under pressure: does the caller need the result of this work to construct its response? If an authorisation check has to pass before an order is confirmed, queuing that check does not remove the caller's dependency on its answer. It just makes the caller wait on a queue and a worker instead of waiting on the check directly, which is strictly more moving parts for the same wait.

A queue does not remove the coupling between producer and consumer. It changes its shape: the API and the worker still depend on the same message format, the same failure semantics, the same eventual outcome. What changes is that the dependency is now mediated by infrastructure that can itself be slow, unavailable, or backed up, adding its own latency and its own failure modes on top of whatever the underlying work already had. Coupling that used to show up as a function call now shows up as a queue depth graph instead.

Async genuinely pays once the caller's own success stops depending on the deferred work finishing on the same timeline: a confirmation email, a thumbnail, a sync to a reporting system. The request can complete before any of that runs, and if the deferred step is slow or briefly unavailable, nobody waiting on the original response notices. The saving is the absence of that wait, not any change in how fast the deferred work itself runs.

Moving slow work into a queue without first asking why it is slow only postpones the question. A worker calling the same unindexed query or the same chatty third-party API a request handler would have called takes just as long to do it, and that cost is now hidden inside a queue depth metric instead of a request latency graph, where it would have been noticed sooner. The queue has not fixed the underlying problem. It has relocated the wait to a place fewer people are watching, and a backlog grows quietly until someone finally asks why yesterday's emails still have not gone out.

Let’s talk

Have a product, platform or delivery challenge? Let’s talk about turning it into a structured, scalable solution.

Open to technical leadership, product delivery and senior engineering roles, and available for architecture consulting, technical reviews and mentorship. Engagements run as project-based work, contracts, consulting, freelance engagements, remote collaboration and long-term partnerships.

Based in Cairo, Egypt, working remotely with clients across the MENA region and internationally.

Also on LinkedIn (opens in a new tab)