SQS Dead-Letter Queues: Redrive Policies, maxReceiveCount, and Safe Replay
Every SQS operator eventually meets the dead-letter queue: the safety net where messages go when normal processing keeps failing. It is also one of the most commonly misconfigured parts of an SQS setup — a maxReceiveCount set too low sends slow-but-healthy messages to the graveyard, while one set too high lets poison messages burn consumer hours in a retry loop. This post covers how DLQs and redrive policies actually work, how to choose a maxReceiveCount that matches your failure profile, and how to replay messages safely once the underlying bug is fixed.
What a Dead-Letter Queue Actually Is
A dead-letter queue is just another SQS queue. It has no special type — what makes it a DLQ is a redrive policy attached to a source queue that names it. When a message in the source queue exceeds the allowed receive attempts, SQS moves it to the DLQ instead of delivering it again.
import { SQSClient, SetQueueAttributesCommand } from "@aws-sdk/client-sqs";
const sqs = new SQSClient({ region: "us-east-1" });
// Attach a redrive policy to the source queue
await sqs.send(new SetQueueAttributesCommand({
QueueUrl: process.env.ORDERS_QUEUE_URL,
Attributes: {
RedrivePolicy: JSON.stringify({
deadLetterTargetArn: process.env.ORDERS_DLQ_ARN,
maxReceiveCount: "5",
}),
},
}));
Two things happen on move:
- The message body is preserved byte-for-byte, along with its message attributes — but the receipt handle changes, and the receive count keeps its history.
- No notification is sent. SQS does not tell anyone the message moved. If nobody is watching the DLQ, the failure is silent. A DLQ without monitoring is not a safety net; it is a black hole.
A standard queue can use a standard DLQ; a FIFO queue must use a FIFO DLQ. Mixing them fails at policy-attach time, not quietly later.
How maxReceiveCount Interacts with the Visibility Timeout
The counter that drives redrive is the ApproximateReceiveCount — incremented every time a message is received, including redeliveries after a visibility timeout expiry. This is where most misconfiguration lives, because the two mechanisms compound:
- A message that keeps failing consumes one receive per attempt and moves to the DLQ after
maxReceiveCountattempts — the intended behavior. - A message that is merely slow also consumes receives whenever its visibility timeout expires mid-processing. With a 30-second timeout, 40-second handlers, and
maxReceiveCount: 3, a perfectly healthy message can rack up three receives before it finishes — and land in the DLQ having never failed once.
The rule of thumb: before tuning maxReceiveCount, make sure the visibility timeout covers your p99 processing time (we cover that side of the equation in SQS visibility timeout tuning). Redrive policy tuning is meaningless while slow messages are being miscounted as failures.
A reasonable starting point is maxReceiveCount: 5 for handlers with transient upstream dependencies, lower (2–3) for handlers whose failures are almost always deterministic — a malformed payload will fail forever, and extra attempts just delay the quarantine.
Designing the DLQ, Not Just Attaching It
Three practices separate a useful DLQ from a decorative one:
1. Separate DLQ per source queue. Sharing one DLQ across five source queues means every replay requires separating mixed message types by hand. Per-queue DLQs keep replay mechanical.
2. Preserve failure context. The message body alone rarely explains why it failed. Enrich at the point of failure — wrap the payload with the last error before it moves:
import { SQSClient, SendMessageCommand } from "@aws-sdk/client-sqs";
const sqs = new SQSClient({ region: "us-east-1" });
// Dead-letter enrichment: forward with failure context instead of
// relying on the raw redrive, when you control the failure path
export async function deadLetter(
queueUrl: string,
message: { MessageId: string; Body: string },
error: unknown,
) {
await sqs.send(new SendMessageCommand({
QueueUrl: queueUrl,
MessageBody: JSON.stringify({
originalBody: message.Body,
failedAt: new Date().toISOString(),
errorMessage: error instanceof Error ? error.message : String(error),
stack: error instanceof Error ? error.stack : undefined,
}),
}));
}
3. Alert on DLQ depth, not on individual messages. A CloudWatch alarm on ApproximateNumberOfMessagesVisible for the DLQ (greater than zero, for a few minutes) is the minimum viable monitoring. The count of messages waiting is the signal; the messages themselves come next.
Replaying Dead-Lettered Messages
When the bug is fixed, replay means moving messages from the DLQ back to the source queue. The mechanics are simple — receive from the DLQ, send to the source, delete from the DLQ — but three rules make the difference between a replay and an incident:
- Replay in batches with a pause. Thousands of dead-lettered messages arriving at once is a traffic spike your consumers may not be sized for. Rate-limit the replay loop.
- Re-process idempotently. Some of those messages may have succeeded on a late attempt after the timeout expired (the at-least-once reality again). Handlers must tolerate the duplicate — same as they must for any redelivery.
- Sample before you replay everything. If the failure was a schema change, the first 10 messages tell you. Blind-replaying a DLQ full of now-permanently-invalid payloads just burns through maxReceiveCount a second time.
import {
SQSClient, ReceiveMessageCommand, SendMessageCommand, DeleteMessageCommand,
} from "@aws-sdk/client-sqs";
const sqs = new SQSClient({ region: "us-east-1" });
// Replay with bounded pacing
export async function replayBatch(dlqUrl: string, sourceUrl: string, limit = 25) {
const { Messages } = await sqs.send(new ReceiveMessageCommand({
QueueUrl: dlqUrl,
MaxNumberOfMessages: Math.min(limit, 10),
WaitTimeSeconds: 5,
}));
for (const m of Messages ?? []) {
if (!m.Body || !m.ReceiptHandle) continue;
await sqs.send(new SendMessageCommand({ QueueUrl: sourceUrl, MessageBody: m.Body }));
await sqs.send(new DeleteMessageCommand({
QueueUrl: dlqUrl,
ReceiptHandle: m.ReceiptHandle,
}));
}
return Messages?.length ?? 0;
}
If your failure enrichment wrapped bodies as shown above, strip the envelope back to originalBody during replay so the source queue sees the original payload.
Watching Redrive Health in a Dashboard
The two numbers that describe DLQ health on any queue are depth (messages parked, awaiting human attention) and age of the oldest message (how long the team has been ignoring them). A queue dashboard that surfaces both per queue — plus the source queue's redrive policy (target and maxReceiveCount) next to its live counts — turns "did anything dead-letter today?" into a glance instead of a script. QueueHub shows exactly this panel for SQS queues with a redrive policy attached, including the inflight and receive-count context that explains why messages moved.
Summary
A DLQ is where failed messages go to wait for a human. Set maxReceiveCount against your failure profile after fixing visibility-timeout miscounts, give every source queue its own DLQ, enrich messages with failure context, and alarm on DLQ depth. When it is time to replay, do it in paced batches with idempotent handlers. Treat the DLQ as a work queue for your team, not a landfill — every message in it is a real user-visible event that never completed.
Related Articles
BullMQ Worker Concurrency: How to Choose the Right Value
BullMQ worker concurrency decides how many jobs one worker runs at once — and the right value depends entirely on whether your jobs are I/O-bound or CPU-bound. Learn how the semaphore works, how to size it, and how to change it at runtime.
SQS Visibility Timeout: How Message Redelivery Works and How to Tune It
The SQS visibility timeout decides whether a slow consumer means the message waits or gets processed twice. Learn how the lease works under the hood, how to tune it against your real processing time, and how to spot redelivery failure modes before they become incidents.
BullMQ Repeatable Jobs: Scheduling with Every and Cron
BullMQ repeatable jobs put cron scheduling inside the queue itself. Learn how the repeat option works under the hood, when to prefer it over a separate cron service, and the timezone and removal pitfalls that trip up most teams.