BullMQ Worker Concurrency: How to Choose the Right Value
BullMQ workers are the processes that pull jobs off Redis and execute them, and the concurrency setting decides how many jobs each worker runs at once. It is a single number, and it is routinely set by vibes: copied from a tutorial, left at the default, or cranked up when a backlog forms. The right value depends on what the processor function actually does — a job that is 95% waiting on an HTTP call and a job that is 95% CPU work want completely different concurrency, and confusing the two produces either a starved queue or a melted event loop. This post covers how concurrency works mechanically, how to reason about the number for I/O-bound and CPU-bound jobs, and how to change it safely on a running system.
What Worker Concurrency Actually Does
A BullMQ worker maintains a semaphore of in-flight jobs. With concurrency: 10, the worker keeps up to 10 processor invocations running simultaneously — pulling a new job from the queue each time one finishes. The default is 1: one job at a time, in order.
import { Worker } from "bullmq";
const worker = new Worker(
"emails",
async (job) => {
await sendEmail(job.data.userId);
},
{
connection: { host: "127.0.0.1", port: 6379 },
concurrency: 20,
},
);
The key mental model: concurrency is the number of simultaneous async invocations, not the number of threads. Node.js still runs your JavaScript on one event loop. Concurrency above 1 only helps when the processor spends its time awaiting something external — a database, an API, the filesystem. While one job awaits, the event loop is free to progress others.
I/O-Bound vs CPU-Bound: The Decision That Matters
I/O-bound processors (API calls, DB queries, sending email) scale well with concurrency, because each in-flight job is mostly parked waiting on a socket. A worker doing nothing but await fetch(...) against a fast endpoint can reasonably run 50–100 concurrent jobs on one process. The ceilings are elsewhere: connection pool sizes, rate limits at the downstream service, and your worker's own memory. If your jobs call an API that tolerates 10 requests per second, worker concurrency above what saturates that budget just moves the waiting from the queue into your workers — while tripping the very 429s the queue was supposed to absorb. Pacing at the queue level with the rate limiter is the cleaner tool for that problem.
CPU-bound processors (image resizing, compression, parsing huge payloads) do not benefit from concurrency at all within a single Node.js process — the event loop can only execute one of them at a time, and setting concurrency: 8 just means 8 jobs are pulled into the worker and executed interleaved, each finishing later than it would have alone. Worse, a long CPU task blocks the loop so the worker misses its lock renewals, and the queue's stalled-job detection kicks in — the failure mode we explain in BullMQ stalled jobs. For CPU work, the answer is worker processes (or worker threads with proper offloading), not higher per-process concurrency.
The practical audit: open your processor function and ask what fraction of its runtime is await on something external. That fraction is your headroom for concurrency. 90% awaiting → try 20–50 and watch the queue. 10% awaiting → more concurrency buys you nothing; add processes instead.
Dynamic Concurrency on a Running Worker
One of BullMQ's quieter features is that concurrency is mutable at runtime — no restart required:
// React to backpressure: drain the backlog faster
worker.concurrency = 50;
// Shed load when a downstream dependency degrades
worker.concurrency = 5;
This turns concurrency into a live operational knob. A common pattern is pairing it with downstream health: if the API your jobs call starts returning 5xx or latencies climb, drop concurrency before the queue's retries compound the pressure — then raise it back when the dependency recovers. Doing this by hand requires watching the right signals (waiting count, throughput, job duration), which is exactly what a queue dashboard is for; QueueHub's live worker view shows per-worker concurrency and in-flight counts so the adjustment is grounded in numbers rather than instinct.
Concurrency and Ordered Processing
Concurrency trades ordering for throughput. With concurrency > 1, jobs from the same logical stream can finish out of order — job 102's email lands before job 101's. For most workloads that is fine. When it is not, the standard techniques are:
- Per-entity queues — put all jobs for one user/account on a queue with
concurrency: 1, and shard entities across queues. Order within an entity, parallelism across them. - Flows with sequential children — chain dependent steps so each starts only when the previous completes; see BullMQ flows.
- Priority instead of ordering — if what you actually want is "urgent jobs jump the line" rather than strict sequence, job priorities express that directly without giving up throughput.
Sizing in Practice
A process that has worked for sizing concurrency from first principles:
- Measure baseline job duration — p50 and p99 — at concurrency 1.
- Double concurrency and watch three numbers: job duration (if it inflates, you are contending somewhere), throughput (if it does not rise, you hit a ceiling), and downstream error rates.
- Stop when throughput plateaus — that plateau is the real limit (event loop, connection pool, or downstream capacity), and pushing past it only adds latency and retry pressure.
- Scale out, not just up — concurrency scales one process's awaiting capacity; more worker processes also add Redis polling parallelism and crash isolation.
Summary
Worker concurrency is one number, but choosing it is a classification exercise: I/O-bound jobs scale with it until a downstream budget says stop, CPU-bound jobs do not scale with it at all and need processes instead. Measure job durations, raise concurrency in steps, and watch throughput and error rates to find your actual ceiling. Use the runtime concurrency knob to react to dependency health, keep strict ordering confined to queues or flows that genuinely need it, and ground every adjustment in the queue metrics a dashboard can see.
Related Articles
BullMQ Job Retention: removeOnComplete, removeOnFail, and Cleaning Up Redis
Without retention settings BullMQ keeps completed and failed jobs in Redis forever. Learn how job history is stored, how removeOnComplete and removeOnFail work, choosing retention per job class, and cleaning accumulated backlogs safely.
SQS Dead-Letter Queues: Redrive Policies, maxReceiveCount, and Safe Replay
SQS dead-letter queues catch messages that keep failing — but a misconfigured maxReceiveCount buries healthy ones. Learn how redrive policies work, how to design a DLQ worth monitoring, and how to replay messages without causing a second incident.
SQS Visibility Timeout: How Message Redelivery Works and How to Tune It
The SQS visibility timeout decides whether a slow consumer means the message waits or gets processed twice. Learn how the lease works under the hood, how to tune it against your real processing time, and how to spot redelivery failure modes before they become incidents.