BullMQ Job Retention: removeOnComplete, removeOnFail, and Cleaning Up Redis
Every job added to a queue has a lifecycle: it waits, runs, and eventually ends up completed, failed, or delayed for another attempt. What happens after that ending is a decision most teams never make consciously — and Redis makes it for them, badly. Without explicit retention settings, BullMQ keeps completed and failed jobs in Redis forever, and the dataset grows until used_memory forces the conversation. This post covers how BullMQ stores job history, which retention options control it, how to choose values, and why cleanup belongs in the queue's configuration rather than a weekend cron script.
How BullMQ Stores Job History
For every job, BullMQ maintains a hash in Redis with its data, options, and results, plus membership in sorted sets per state — wait, active, completed, failed, delayed. The hash keys look like bull:{queueName}:{jobId}. Nothing removes these automatically in the base configuration. A queue doing 100,000 jobs a day with no retention policy is adding ~100,000 hashes and sorted-set entries per day, indefinitely.
The consequences arrive in two forms:
- Memory pressure. Redis
used_memoryclimbs until eviction policies or OOM behavior intervene — none of which interact gracefully with a queue's data structures. - Degraded operations.
getJobs()scans grow slower as state sorted sets get longer, and dashboards listing recent jobs pay the cost first. If your queue UI has become sluggish over months, unbounded job history is the usual suspect.
Retention settings also interact with features that read job history later: a flow parent that needs its children's return values can be broken by aggressive removeOnComplete on the child queue, a pitfall we detail in BullMQ flows.
The Retention Options: removeOnComplete and removeOnFail
BullMQ lets you set retention per job at add() time, or per queue as defaults:
import { Queue } from "bullmq";
const queue = new Queue("emails", {
connection: { host: "127.0.0.1", port: 6379 },
defaultJobOptions: {
removeOnComplete: { age: 3600, count: 5000 },
removeOnFail: { age: 86400 * 7 },
},
});
await queue.add("send-email", { userId: 42 });
The two shapes mean different things:
removeOnComplete: true— the job is removed immediately on completion. Zero history.removeOnComplete: { age: 3600, count: 5000 }— remove when older than 3600 seconds and keep at most 5000 completed jobs, whichever prunes first. Using both together prevents both "history too old to matter" and "sudden burst fills memory between cleanup passes."
removeOnFail deserves more caution than removeOnComplete. Failed jobs are diagnostic evidence: the payload that produced the error, the stack trace, the attempt count. A team that auto-deletes failures after an hour loses the ability to answer "what exactly broke at 2am?" Set failed retention generously (weeks), and rely on alerting to know about failures fast — see our guide on BullMQ stalled jobs for detecting the related case of jobs that never report either way.
Retention by Job Type, Not Just by Queue
The most effective retention setups treat job classes differently inside one queue:
// Critical jobs: keep failure evidence for a long time
await queue.add("charge-card", { orderId: 99 }, {
removeOnComplete: { age: 86400 * 3 },
removeOnFail: false, // keep until explicitly cleaned after postmortem
});
// High-volume, low-value jobs: minimal footprint
await queue.add("track-analytics-event", { event: "page_view" }, {
removeOnComplete: true,
removeOnFail: { age: 3600 },
});
The charge-card example uses removeOnFail: false — failed payment jobs stay until a human decides, because the cost of keeping one is trivial next to the cost of losing the audit trail. The analytics event is the opposite profile: millions per day, individually worthless, and there is no investigation that needs yesterday's page views.
Cleaning Up Existing History: queue.clean()
Retention options only affect jobs added after them. Existing accumulated history needs an explicit cleanup:
// Remove completed jobs older than 24h, in batches of 1000
await queue.clean(86400, 1000, "completed");
// Remove failed jobs older than 30 days
await queue.clean(86400 * 30, 1000, "failed");
Two operational notes:
- Run it in batches, not once with a huge count.
clean()is a single Redis call; a multi-million-job clean in one shot blocks the event loop and stalls every queue sharing the instance. - Prefer off-peak windows for the first big cleanup of an accumulated backlog. On shared Redis instances, the clean pass is felt by everything else on the box.
A queue dashboard with per-queue retention visibility makes the before/after obvious — QueueHub's queue detail view shows state counts alongside the configured retention policy, so you can see a completed count that never shrinks and know exactly which knob is missing.
What Not to Do
- Don't set
removeOnComplete: trueon queues you debug. Immediate removal deletes the result payload too. On queues where "what did this job return?" is a reasonable question, keep an{ age, count }window. - Don't rely on Redis
maxmemoryeviction as your retention policy. Allkeys-LRU will happily evict the sorted-set entries tracking active queue state, not just old job hashes. Retention belongs to the queue layer, in config, where its semantics are explicit. - Don't clean during incidents. A mass cleanup destroys exactly the evidence you are about to need. Clean after the postmortem, not during the fire.
Summary
Unbounded job history is the most common way a healthy BullMQ deployment slowly becomes an unhealthy Redis instance. Set removeOnComplete with both age and count on every queue, keep failed jobs long enough to investigate, treat job classes differently where their value differs, and do the one-time backlog cleanup in paced batches. Retention is a queue configuration concern, not an emergency measure — and a dashboard that shows state counts next to retention policy makes drift visible before it becomes memory pressure.
Related Articles
BullMQ Worker Concurrency: How to Choose the Right Value
BullMQ worker concurrency decides how many jobs one worker runs at once — and the right value depends entirely on whether your jobs are I/O-bound or CPU-bound. Learn how the semaphore works, how to size it, and how to change it at runtime.
BullMQ Flows: Parent-Child Jobs and When to Use Them
BullMQ flows model fan-out/fan-in batches as parent-child job trees: one parent job that only completes when its children finish, with results collected automatically. Learn how flow trees work under the hood, when they beat manual job chaining, and the stalled-parent pitfalls to avoid.
BullMQ Repeatable Jobs: Scheduling with Every and Cron
BullMQ repeatable jobs put cron scheduling inside the queue itself. Learn how the repeat option works under the hood, when to prefer it over a separate cron service, and the timezone and removal pitfalls that trip up most teams.