SQS Visibility Timeout: How Message Redelivery Works and How to Tune It
The SQS visibility timeout decides what happens when a consumer runs slow, crashes mid-message, or silently stops responding — whether the message waits patiently or gets handed to another consumer and processed twice. Most teams set it once at queue creation and never revisit it, until duplicate charges or double-sent webhooks force a painful reintroduction to how redelivery works. This post explains what the SQS visibility timeout does, how to tune it to your real processing time, and how to spot the failure modes early.
What the SQS Visibility Timeout Does
When a consumer receives a message, SQS hides it from every other consumer for the duration of the visibility timeout. The intent is simple: give the receiving consumer temporary, exclusive ownership so it can process the message and call DeleteMessage when it is done.
- Default visibility timeout: 30 seconds
- Standard queues: 0 seconds to 12 hours (43,200 seconds)
- FIFO queues: 1 second to 12 hours
If the consumer deletes the message in time, the lease ends cleanly and the message never reappears. If it does not — the work was slow, the process crashed, the network died — the visibility timeout expires and the message becomes visible again, eligible to be received by any consumer, possibly a different one than the first attempt:
import { SQSClient, ReceiveMessageCommand, DeleteMessageCommand } from "@aws-sdk/client-sqs";
const sqs = new SQSClient({ region: "us-east-1" });
const { Messages } = await sqs.send(new ReceiveMessageCommand({
QueueUrl: process.env.QUEUE_URL,
MaxNumberOfMessages: 10,
VisibilityTimeout: 60, // overrides the queue default for these messages
WaitTimeSeconds: 20, // long polling: wait up to 20s instead of returning empty
}));
for (const message of Messages ?? []) {
await process(message.Body);
// The receipt handle is your proof of ownership — delete with it
await sqs.send(new DeleteMessageCommand({
QueueUrl: process.env.QUEUE_URL,
ReceiptHandle: message.ReceiptHandle,
}));
}
The Mental Model: A Lease, Not a Lock
The most common misconception is that the visibility timeout works like a mutex — that a message is "locked" by the first consumer and everyone else queues up politely behind it. It is not a lock. It is a lease with an expiry, and nothing stops that lease from expiring while the consumer is still working. Two consequences follow:
- Your work must be idempotent. SQS delivers at least once, and an expired visibility timeout is the most common way a message ends up delivered more than once. If processing a message twice produces two charges or two emails, no timeout value will save you — design the handler to tolerate duplicates.
- Two consumers can process the same message at once. Consumer A's lease expires mid-processing and Consumer B receives the message while A is still running. BullMQ teams know this exact failure mode as a stalled job, where a worker's processing lock expires and another worker picks the job up — see our guide to BullMQ stalled jobs for the Redis-side mechanics.
What Happens When a Message Is Redelivered
Every time a message becomes visible again after an expiry, SQS increments its receive count. Two details matter once redelivery starts:
- The receipt handle changes on every receive. A handle from an earlier receive is worthless after redelivery; calling
DeleteMessagewith it errors and the message stays in the queue. If your error handling treats that as "fine, it is deleted," you now have a duplicate in flight. - Repeated redeliveries eventually hit the dead-letter queue. With a redrive policy and
maxReceiveCountset, a message whose receive count exceeds the threshold moves to the DLQ instead of looping forever. That is your safety net — but it only catches messages that keep failing. A message that is redelivered and succeeds anyway (because it was already processed the first time) never reaches it.
How to Choose the Right SQS Visibility Timeout
There is no universal correct value, but there is a correct process: measure what your handler actually needs, end to end, and set the timeout comfortably above the tail of that distribution.
- Too short is the expensive failure. Work that is merely slow gets redelivered to a second consumer, so "slow" turns into "duplicated." If your p99 processing time is 40 seconds, the 30-second default guarantees a steady stream of double-processing.
- Too long is the quiet failure. A poison message — one the handler can never succeed on — occupies its full visibility window on every attempt, retrying slowly, polluting your inflight count, and delaying the DLQ redrive that would finally quarantine it.
The robust pattern for long-running or variable work is a heartbeat: keep the initial timeout short and extend it with ChangeMessageVisibility as long as the task is making progress:
import { SQSClient, ChangeMessageVisibilityCommand } from "@aws-sdk/client-sqs";
// Extend the lease for another 60 seconds while the task is still alive
await sqs.send(new ChangeMessageVisibilityCommand({
QueueUrl: process.env.QUEUE_URL,
ReceiptHandle: handle,
VisibilityTimeout: 60,
}));
Tuning rules that hold in practice:
- Base the timeout on p99, not p50 — the average hides the stragglers that actually trigger redelivery.
- Add headroom for downstream retries inside the handler. If the handler calls an API with three retries, the timeout must cover the worst-case chain, not the happy path.
- Remember the 12-hour ceiling. A message can be hidden for at most 12 hours, so a genuinely multi-hour task needs a heartbeat loop, not a giant static timeout.
- Use the per-receive override (
VisibilityTimeoutonReceiveMessage) when different message types in one queue need different budgets.
Signs Your SQS Visibility Timeout Is Wrong
Your queues will tell you when the timeout is misconfigured — you just have to watch the right signals:
- Receive counts climbing on healthy queues — messages are being redelivered before succeeding; your timeout is below real processing time.
- Duplicate side effects in consumers — same story, seen from the handler's side. Add idempotency keys and re-check the timeout.
- High inflight counts with low throughput — messages sitting inside visibility windows that never resolve, usually a crashed consumer whose lease has not expired, or a poison message with an over-generous timeout.
- Poison messages taking forever to reach the DLQ — each failed attempt waits out the full timeout; consider a shorter timeout on known-fatal error paths so the redrive happens sooner.
A redelivery storm also hammers whatever downstream system the handler calls. On BullMQ, the queue-level throttle in our rate limiting guide caps that burst; with SQS, pace it inside the consumer instead.
Summary
The SQS visibility timeout is a lease, not a lock: it hides a received message from other consumers for a bounded window, and when that window expires the message is redelivered — possibly to a different consumer, with a fresh receipt handle and an incremented receive count. Set the timeout against your p99 processing time rather than the default, use ChangeMessageVisibility heartbeats for long tasks, and write every handler as if it will run twice — with SQS's at-least-once delivery, it sometimes will. Monitor receive counts and inflight age per queue so redelivery loops surface before they become duplicate side effects.
Related Articles
BullMQ Worker Concurrency: How to Choose the Right Value
BullMQ worker concurrency decides how many jobs one worker runs at once — and the right value depends entirely on whether your jobs are I/O-bound or CPU-bound. Learn how the semaphore works, how to size it, and how to change it at runtime.
SQS Dead-Letter Queues: Redrive Policies, maxReceiveCount, and Safe Replay
SQS dead-letter queues catch messages that keep failing — but a misconfigured maxReceiveCount buries healthy ones. Learn how redrive policies work, how to design a DLQ worth monitoring, and how to replay messages without causing a second incident.
BullMQ Flows: Parent-Child Jobs and When to Use Them
BullMQ flows model fan-out/fan-in batches as parent-child job trees: one parent job that only completes when its children finish, with results collected automatically. Learn how flow trees work under the hood, when they beat manual job chaining, and the stalled-parent pitfalls to avoid.