The Queue Transport Decision: Redis Streams vs. SQS for a Noisy Multi-Tenant Pipeline

Cover Image for The Queue Transport Decision: Redis Streams vs. SQS for a Noisy Multi-Tenant Pipeline
DevOps5 min read

"Just swap the queue" is the kind of fix that sounds cheap and turns out to be either exactly right or completely beside the point, depending on whether you've actually diagnosed what's failing. A multi-tenant processing pipeline had been running on a Redis Stream as its job queue for about a year, and when reliability problems kept showing up, the instinctive first move was to blame the transport. The useful second move — the one that actually mattered — was refusing to trust that instinct until every recent incident had been classified by root cause, and then refusing to trust the first classification either, once it turned out to be partly wrong.

The first analysis was wrong, and that's the interesting part

The initial write-up blamed three things: a metric the autoscaler was reading incorrectly, a stream-trimming setting that was dropping undelivered messages, and the general feeling that a purpose-built managed queue would just be more robust. Two of those three had, by the time anyone looked closely a second time, already been fixed directly on Redis — the autoscaler was corrected to read the actual pending-entries count instead of the wrong metric, and the trim limit was raised by two orders of magnitude. If those two patches were the whole story, the correct conclusion at that point would have been "stay on Redis, the problems are fixed." The real argument for migrating had to be something that patching Redis couldn't touch — and finding it required going back through the incident log and classifying every incident into exactly one of three buckets, not just the "transport's fault" bucket by default:

transport-caused   the queue's own semantics produced this; a different
                    transport with different semantics removes the failure class
job-caused          the job or consumer code is doing something no queue
                    can be blamed for — unbounded retries, no idempotency,
                    an unhealthy dependency
config-caused       correct transport, correct job design, wrong settings

What the reclassification actually found

Seven incidents, reclassified:

#IncidentBucket
1Stream trim dropped undelivered messages — a document got stuck "in queue" forevertransport-caused, structural
2A stalled consumer's message got reclaimed by a second worker — same document processed, and billed, twicejob-caused (see below)
3Autoscaler misread the wrong metric, undercounted real backlogconfig-caused (already fixed)
4Hundreds of "ghost" consumers accumulated over months — a replaced pod leaves a registry entry nobody cleans uptransport-caused
5One tenant's burst of uploads pushed every other tenant's jobs behind it in a single linear stream — adding workers didn't helptransport-caused, the biggest one
6The job queue's stream lived on the same cache node as the application's general-purpose session/response cache — ingest load degraded unrelated app cachingtopology-caused, fixed by moving it off that node
7No native dead-letter mechanism — a hand-rolled "move to a second stream" script had to be built and maintainedtransport-caused

Incident #2 deserves a closer look, because it's the one that looks transport-caused but isn't. A stalled consumer's claim on a message got reclaimed by a second worker because the event loop stalled long enough to stop the consumer's liveness signal. That's not specific to Redis Streams — a managed queue's visibility-timeout-plus-extension mechanism is the exact same shape of problem; migrating transport while keeping a 30-minute unit of work reproduces this bug completely unchanged on the new system. The only real fix is shrinking the unit of work below any plausible lease duration, which is a job-design fix, not a transport fix — the kind of distinction that's easy to blur if you're already looking for reasons to justify a migration you've half-decided on.

Net: four of the seven were genuinely structural to how a single Redis Stream shares consumption across consumers — not bugs to be coded around, but properties of the transport itself. That 4-of-7, not "a managed service feels more robust," is the actual argument for a migration.

The deciding mechanism: fairness across tenants, without building it yourself

The sharpest of the structural incidents was the noisy-neighbor problem: every tenant's jobs lived on one physical stream, and one tenant's traffic spike delayed processing for every other tenant sharing that consumer group. A consumer-group primitive gives you shared consumption, but fairness across logical tenants on top of a single physical stream isn't something the primitive expresses on its own — building it means sharding streams per tenant and hand-rolling a weighted scheduler and starvation protection, then maintaining all of that indefinitely.

A newer SQS feature — Fair Queues, available on standard (not FIFO) queues — solves exactly this, natively: tag every message with a MessageGroupId at send time, and SQS tracks each group's share of in-flight messages, biasing delivery toward groups that aren't already saturating their share, with zero consumer-side code changes and a dedicated CloudWatch metric for watching whether quiet tenants are being starved.

# Producer: tag by tenant, on a STANDARD queue.
sqs.send_message(
    QueueUrl=queue_url,
    MessageBody=json.dumps(job_payload),
    MessageGroupId=f"tenant-{tenant_id}",
)
# Consumer: unchanged. Fairness across MessageGroupId values
# is handled by SQS, not by application code.
messages = sqs.receive_message(QueueUrl=queue_url, MaxNumberOfMessages=10)

The queue type matters more than it looks here, and it's a trap specifically because the field name is identical on both. FIFO queues use the same MessageGroupId field, but its semantic is to strictly serialize everything within a group — enabling FIFO out of habit, because "ordered" sounds safer, would recreate the exact head-of-line blocking this migration exists to remove, just scoped down to one tenant instead of the whole stream. Standard queues with Fair Queues enabled are the ones that actually match the requirement: deliver fairly across groups, never serialize within one. Known limits worth knowing going in: fairness is based on in-flight count, so it's a coarse signal with very few concurrent workers; and it does not rate-limit a noisy tenant when nobody else is waiting — that's a job for admission-control quotas, deliberately kept separate.

This is the bar a transport migration should have to clear: not "the other option is generally considered more robust," but "here is a specific, recurring failure, and here is the exact mechanism the new transport already provides that removes it, as opposed to something application code would have to build and maintain forever."

The bug caught before shipping: two clocks that must agree

Ownership of a unit of work lives on two independent clocks — a lease expiry recorded in the application's own database, and the queue's own visibility timeout — and they have to agree or you get a subtle self-inflicted failure. An early version of the migration computed the queue's visibility timeout as roughly double each stage's own processing timeout. That's wrong, and the way it's wrong is worth internalizing: if the queue's visibility window is shorter than the database lease, the queue redelivers a message that the database still considers owned. The second delivery's claim correctly fails — no double-processing, no corruption — but the redelivery still counts against the queue's delivery-count limit. The practical effect: a stage that is simply running slowly, working exactly as intended, ends up moving itself to the dead-letter queue purely because it took a little longer than expected, with no actual failure anywhere in the chain. The fix was to derive the visibility timeout programmatically from the longest lease among the stages sharing that queue plus a fixed safety margin, and assert that relationship in a test against the actual infrastructure configuration — never hand-maintained as two numbers in two different files that can quietly drift apart.

What the delivery-count limit actually means

A detail that's commonly misunderstood badly enough to cause wasted tuning effort: the retry budget for a unit of work does not live in the queue at all. A worker deletes its message on every outcome, including "this needs to be retried" — a retry travels onward as a brand-new message carrying its own computed backoff, and the attempt counter that actually matters lives in the application's own database, where it survives redelivery regardless of which physical message carries the work next. So a delivery-count limit of five means "five different workers died while holding this exact message" — a crash loop — not "this unit of work was retried five times." Turning that limit up when documents start failing changes nothing about actual retry behavior, because the real retry budget was never there to begin with. It's a documented anti-pattern specifically because it's such an intuitive thing to reach for when retries aren't behaving as expected.

Queue boundaries have to match pool boundaries, not pipeline stages

One queue exists per worker pool, not per pipeline stage, because an autoscaler that scales a deployment from a single queue's depth needs that queue to represent exactly the unit it's scaling — mixing two unrelated stages onto one queue breaks that signal. Every queue gets its own dedicated dead-letter queue rather than sharing one across pools; a shared dead-letter queue would mix a permanently malformed job from one pool with a misconfigured job from a completely different pool, turning "which pool is stuck in a crash loop" from something you can see at a glance into something you have to go query for.

What didn't move, and what the actual cost looked like

The job-design and configuration incidents — the unbounded retry loop, the oversized-payload-in-the-queue pattern, the misconfigured backlog alerting — got fixed in place, before any transport migration work started. Migrating the transport first and fixing job design second would have made it impossible to tell which change actually improved reliability, because better numbers after the fact could have come from either one. What stayed on the in-memory store entirely, correctly scoped down rather than removed: short-TTL permission caches, a response-level semantic cache, idempotency/dedup markers, and distributed rate limits — none of which are job queues, and none of which benefit from a managed queue's delivery guarantees. Moving a handful of job queues off a shared cache node was, in practice, the cheapest way to give memory and CPU headroom back to that cache — cheaper than standing up a second instance just to isolate it. And the migration's actual infrastructure cost, at realistic message volume, landed in the single-digit dollars per month — explicitly not a factor in the decision even at two orders of magnitude more load, which is worth noting precisely because cost is so often assumed to be the deciding factor in a transport choice when it's usually nowhere close to the real constraint.

The reusable part isn't the answer, it's the method

Moving off a Redis Stream was the right call here, for a multi-tenant fairness problem that's structural to how a single stream shares consumption across consumers — and it wouldn't have been discoverable at all without refusing to trust the first, intuitive diagnosis, reclassifying every incident by actual root cause, and specifically separating "this is inherent to the transport" from "this would reproduce unchanged on any transport because the real problem is unit-of-work size." The expensive mistake to avoid isn't picking the wrong queue. It's migrating transports to fix problems that were never about the transport, and only finding that out after the migration is already done.