Benchmarks
Throughput, latency and concurrency scaling for sepp against beanstalkd, Faktory, NATS JetStream and BullMQ.
sepp is benchmarked against beanstalkd, Faktory, NATS JetStream and BullMQ with an identical workload. Brokers run one at a time in Docker under identical resource caps and each is driven by its native client over its native protocol.
Methodology
Three phases run in order against each broker:
- Enqueue. N producer connections enqueue 256 byte jobs one RPC at a time until 200,000 are stored. Reported as jobs/s.
- Drain. N worker connections reserve and acknowledge one job at a time until the queue is empty, so a drained job costs two RPCs. The timer stops at the last ack; idle reserve polls after the queue empties don't inflate the number.
- Latency. One connection runs sequential enqueue, reserve, ack round trips. Reported as p50/p99/p999 of the full round trip.
Each connection keeps exactly one job in flight. These are closed-loop numbers: they measure sustainable throughput, not behavior under overload. The headline tables use 50 producers and 50 workers.
Durability modes
A fair comparison requires equal durability guarantees, so the suite runs twice:
| broker | durable | buffered |
|---|---|---|
| sepp | persist_mode = "sync_data" | persist_mode = "buffer" |
| beanstalkd | binlog, -f 0 (fsync every write) | binlog, -f 86400000 |
| Faktory | none available | defaults (Redis RDB snapshots) |
| NATS JetStream | file storage, sync_interval: always | file storage, default flushing |
| BullMQ (Redis) | AOF, appendfsync always | AOF, appendfsync no |
In durable mode an operation is fsync-ed before the broker acknowledges it. Faktory persists only via periodic RDB snapshots (save 120 1, save 30 5, hardcoded in its storage layer), so it appears only in the buffered table. BullMQ is a client library over Redis, so the broker under test is Redis itself, driven by the official BullMQ Rust client with AOF as the only persistence.
Durable results
| broker | enqueue jobs/s | drain jobs/s | p50 ms | p99 ms | p999 ms |
|---|---|---|---|---|---|
| sepp | 15,325 | 6,980 | 4.820 | 9.784 | 12.755 |
| beanstalkd | 1,494 | 1,431 | 1.421 | 6.261 | 7.057 |
| Faktory | - | - | - | - | - |
| NATS JetStream | 716 | 724 | 3.202 | 23.750 | 34.420 |
| BullMQ | 10,939 | 5,566 | 5.113 | 10.381 | 13.352 |
Buffered results
Not the recommended use case
sepp is designed for durability. Although it can run in buffered mode, some architectural design decisions made to make durable mode fast (like a single committer thread) make buffered mode slower than most of the other brokers. It is still plenty fast enough for most applications.
| broker | enqueue jobs/s | drain jobs/s | p50 ms | p99 ms | p999 ms |
|---|---|---|---|---|---|
| sepp | 52,325 | 25,124 | 1.077 | 1.305 | 1.728 |
| beanstalkd | 44,324 | 30,530 | 0.275 | 0.462 | 0.545 |
| Faktory | 77,643 | 23,315 | 0.855 | 1.052 | 1.178 |
| NATS JetStream | 97,857 | 29,205 | 0.664 | 0.822 | 1.132 |
| BullMQ | 28,808 | 12,963 | 0.670 | 0.917 | 1.032 |
With durability off the field compresses to within roughly 3.5x and the ordering follows protocol overhead. The gap between the two tables is what durability costs: BullMQ keeps 38% of its buffered enqueue rate, sepp 29%, beanstalkd 3% and NATS under 1%.
Throughput vs concurrency
Durable mode with equal producer and worker counts.
Chart data, including drain
Cells are enqueue / drain jobs/s.
| connections | sepp | beanstalkd | NATS JetStream | BullMQ |
|---|---|---|---|---|
| 1 | 708 / 543 | 785 / 1,166 | 551 / 576 | 524 / 308 |
| 8 | 3,127 / 2,405 | 1,516 / 1,459 | 708 / 704 | 2,329 / 1,165 |
| 64 | 19,011 / 8,289 | 1,498 / 1,446 | 720 / 718 | 14,522 / 7,214 |
| 256 | 37,184 / 18,596 | 1,460 / 1,445 | 724 / 721 | 20,453 / 10,132 |
At one connection every broker pays a full fsync per job and all four land within 1.5x of each other. sepp's committer batches every operation waiting in line into a single fsync (group commit), so throughput grows with the number of concurrent connections. Redis batches AOF fsyncs across clients the same way, so BullMQ scales too, though it flattens past 64 connections at about half of sepp's rate. beanstalkd and NATS fsync serially and plateau by eight connections. At 256 connections durable sepp reaches 71% of its own buffered enqueue rate.
Job counts scale with concurrency in these runs (5,000 at one connection, 40,000 at eight, 200,000 above) to keep wall times comparable. Throughput is a rate, so rows remain comparable.
Environment
- AMD Ryzen 5 7600X (6 cores, 12 threads), 32 GB RAM
- Kingston Fury Renegade NVMe (7,300 MB/s read, 6,000 MB/s write)
- Docker Desktop on WSL2, Windows 11
Each broker container is capped at 8 CPUs and 8 GB of memory; drivers run unconstrained so the broker is always the bottleneck under test. Versions: sepp 0.1.0, beanstalkd 1.13, Faktory 1.9.4, NATS 2.10, Redis 7 (BullMQ Rust client 1.0.0).
Durable throughput tracks the disk's fsync latency above all else. Absolute numbers will differ on enterprise drives with power-loss protection or on cloud block storage, so rerun the suite on your own hardware before drawing fine-grained conclusions.