Designing webhooks your customers will not complain about
Webhooks are an API you deliver to someone else's unreliable server. Design for their failures, not just your success.
Move any work that is slow, unreliable or not needed for the response into a background queue: emails, third-party API calls, report generation, image processing and webhooks. Jobs must be idempotent because they will be retried, need exponential backoff with jitter, and need a dead letter queue plus alerting so failures are visible rather than silent.
The test: if this step fails, should the user's action fail? If a signup should still succeed when the welcome email provider is down, the email is a job.
Because every queue worth using guarantees at-least-once delivery, which means duplicates are normal operation, not an error condition.
A worker can process a job, crash before acknowledging, and have the job redelivered. Without idempotency that is a second charge, a second email, a second shipment. Exactly-once delivery is not something a distributed queue can offer you.
async function handleChargeJob(job: { orderId: string; idempotencyKey: string }) {
const existing = await db.charge.findUnique({ where: { key: job.idempotencyKey } });
if (existing) return existing; // already done — safe to re-run
const charge = await payments.charge(job.orderId, { key: job.idempotencyKey });
return db.charge.create({ data: { key: job.idempotencyKey, id: charge.id } });
}| Attempt | Delay (base 2s, with jitter) |
|---|---|
| 1 | immediate |
| 2 | ~2s |
| 3 | ~4s |
| 4 | ~8s |
| 5 | ~16s |
| 6 | ~32s, then dead letter |
| Option | Best for | Trade-off |
|---|---|---|
| Redis-backed (BullMQ) | Node apps already using Redis | Durability depends on Redis persistence |
| Postgres-backed (pg-boss, Graphile) | Small teams, transactional enqueue | Throughput limited by the database |
| Managed (SQS, Cloud Tasks) | Scale and durability without ops | Vendor-specific semantics |
| Kafka / streams | High-volume event pipelines | Operationally heavy for simple jobs |
For most product teams, enqueueing into the same database as the write that triggered it — inside the same transaction — removes an entire class of 'the row was created but the job never ran' bugs. That transactional outbox pattern is worth more than raw throughput at typical scale.
No. In-process timers are lost on deploy, crash or scale-down, and they do not retry. Anything that must happen needs durable storage.
Return a job id immediately, expose a status endpoint, and poll or push updates over a socket. Storing job state in a table you own gives you history and debuggability for free.
Writing the job row in the same database transaction as the business change, then having a separate process publish it. It guarantees the job exists if and only if the change committed.
Long enough to debug — typically 7 to 30 days for completed jobs and longer for failures. Purge on a schedule or the table becomes your largest one.
Harshal Patel
Founder & Lead Engineer, ROVQIX
Harshal leads engineering at ROVQIX, where he has shipped production Next.js, Node.js and PostgreSQL systems for startups, SaaS teams and ecommerce brands. He writes about the trade-offs behind architecture decisions rather than the framework of the week.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
Webhooks are an API you deliver to someone else's unreliable server. Design for their failures, not just your success.
Node is single-threaded for your code. Every millisecond you spend in a synchronous loop is a millisecond nobody else's request is being served.
Observability is not more dashboards. It is being able to answer a question you did not anticipate, without shipping new code.
No spam. Just the occasional case study and craft breakdown.