·QueueHub Team·5 min read

BullMQ Job Retention: removeOnComplete, removeOnFail, and Cleaning Up Redis

BullMQRedisjob retentionremoveOnCompletecleanupmemory management

Every job added to a queue has a lifecycle: it waits, runs, and eventually ends up completed, failed, or delayed for another attempt. What happens after that ending is a decision most teams never make consciously — and Redis makes it for them, badly. Without explicit retention settings, BullMQ keeps completed and failed jobs in Redis forever, and the dataset grows until used_memory forces the conversation. This post covers how BullMQ stores job history, which retention options control it, how to choose values, and why cleanup belongs in the queue's configuration rather than a weekend cron script.

How BullMQ Stores Job History

For every job, BullMQ maintains a hash in Redis with its data, options, and results, plus membership in sorted sets per state — wait, active, completed, failed, delayed. The hash keys look like bull:{queueName}:{jobId}. Nothing removes these automatically in the base configuration. A queue doing 100,000 jobs a day with no retention policy is adding ~100,000 hashes and sorted-set entries per day, indefinitely.

The consequences arrive in two forms:

  • Memory pressure. Redis used_memory climbs until eviction policies or OOM behavior intervene — none of which interact gracefully with a queue's data structures.
  • Degraded operations. getJobs() scans grow slower as state sorted sets get longer, and dashboards listing recent jobs pay the cost first. If your queue UI has become sluggish over months, unbounded job history is the usual suspect.

Retention settings also interact with features that read job history later: a flow parent that needs its children's return values can be broken by aggressive removeOnComplete on the child queue, a pitfall we detail in BullMQ flows.

The Retention Options: removeOnComplete and removeOnFail

BullMQ lets you set retention per job at add() time, or per queue as defaults:

import { Queue } from "bullmq";

const queue = new Queue("emails", {
  connection: { host: "127.0.0.1", port: 6379 },
  defaultJobOptions: {
    removeOnComplete: { age: 3600, count: 5000 },
    removeOnFail: { age: 86400 * 7 },
  },
});

await queue.add("send-email", { userId: 42 });

The two shapes mean different things:

  • removeOnComplete: true — the job is removed immediately on completion. Zero history.
  • removeOnComplete: { age: 3600, count: 5000 } — remove when older than 3600 seconds and keep at most 5000 completed jobs, whichever prunes first. Using both together prevents both "history too old to matter" and "sudden burst fills memory between cleanup passes."

removeOnFail deserves more caution than removeOnComplete. Failed jobs are diagnostic evidence: the payload that produced the error, the stack trace, the attempt count. A team that auto-deletes failures after an hour loses the ability to answer "what exactly broke at 2am?" Set failed retention generously (weeks), and rely on alerting to know about failures fast — see our guide on BullMQ stalled jobs for detecting the related case of jobs that never report either way.

Retention by Job Type, Not Just by Queue

The most effective retention setups treat job classes differently inside one queue:

// Critical jobs: keep failure evidence for a long time
await queue.add("charge-card", { orderId: 99 }, {
  removeOnComplete: { age: 86400 * 3 },
  removeOnFail: false, // keep until explicitly cleaned after postmortem
});

// High-volume, low-value jobs: minimal footprint
await queue.add("track-analytics-event", { event: "page_view" }, {
  removeOnComplete: true,
  removeOnFail: { age: 3600 },
});

The charge-card example uses removeOnFail: false — failed payment jobs stay until a human decides, because the cost of keeping one is trivial next to the cost of losing the audit trail. The analytics event is the opposite profile: millions per day, individually worthless, and there is no investigation that needs yesterday's page views.

Cleaning Up Existing History: queue.clean()

Retention options only affect jobs added after them. Existing accumulated history needs an explicit cleanup:

// Remove completed jobs older than 24h, in batches of 1000
await queue.clean(86400, 1000, "completed");
// Remove failed jobs older than 30 days
await queue.clean(86400 * 30, 1000, "failed");

Two operational notes:

  • Run it in batches, not once with a huge count. clean() is a single Redis call; a multi-million-job clean in one shot blocks the event loop and stalls every queue sharing the instance.
  • Prefer off-peak windows for the first big cleanup of an accumulated backlog. On shared Redis instances, the clean pass is felt by everything else on the box.

A queue dashboard with per-queue retention visibility makes the before/after obvious — QueueHub's queue detail view shows state counts alongside the configured retention policy, so you can see a completed count that never shrinks and know exactly which knob is missing.

What Not to Do

  • Don't set removeOnComplete: true on queues you debug. Immediate removal deletes the result payload too. On queues where "what did this job return?" is a reasonable question, keep an { age, count } window.
  • Don't rely on Redis maxmemory eviction as your retention policy. Allkeys-LRU will happily evict the sorted-set entries tracking active queue state, not just old job hashes. Retention belongs to the queue layer, in config, where its semantics are explicit.
  • Don't clean during incidents. A mass cleanup destroys exactly the evidence you are about to need. Clean after the postmortem, not during the fire.

Summary

Unbounded job history is the most common way a healthy BullMQ deployment slowly becomes an unhealthy Redis instance. Set removeOnComplete with both age and count on every queue, keep failed jobs long enough to investigate, treat job classes differently where their value differs, and do the one-time backlog cleanup in paced batches. Retention is a queue configuration concern, not an emergency measure — and a dashboard that shows state counts next to retention policy makes drift visible before it becomes memory pressure.

Related Articles