All WordPress HTML Templates Forms & Webhooks AI & Tools
Forms & Webhooks

Queues for Webhook Processing: Why You Should Not Do Work Inline

Why inline webhook handlers break under load, and how to build a webhook queue with Postgres SKIP LOCKED, idempotency keys, backoff and dead letter handling.

A real working moment illustrating the theme of an article about webhook queue. Wide 16:9 banner, one strong focal point, magazine editorial quality, authentic and unstaged.

At 04:12 on a Tuesday we got paged for a client’s checkout flow. Stripe had marked their webhook endpoint as failing, and the reason was almost funny: their invoice.payment_succeeded handler was generating a PDF, uploading it to S3, then calling a transactional email API before returning a 200. On a normal day that took 1.9 seconds. That morning the email provider was degraded and every request hung for 30 seconds, so Stripe timed out, retried, and each retry started another PDF render on a box already saturated with hanging PHP-FPM workers.

Nothing was lost in the end, because Stripe retries for up to three days. But 400-odd duplicate invoices went out, and the fix was the one nobody had wanted to build during the sprint: a webhook queue. Receive, verify, store, respond. Do the actual work somewhere else.

This is the piece I wish existed when we started doing this properly around 2019. It’s the version with the numbers, the SQL, and the failure modes we’ve actually hit on production client work.

Key Takeaways

  • Provider timeouts are tighter than people assume: GitHub cuts deliveries off at 10 seconds, Shopify at 5, Slack interactivity at 3. Any handler doing network I/O inline is gambling against someone else’s uptime.
  • Your inline handler should do four things only: verify the signature over the raw body, insert the event with a unique constraint on the provider’s event ID, enqueue a job, return 202. Everything else is a worker’s problem.
  • You don’t need Kafka. Postgres with FOR UPDATE SKIP LOCKED handles thousands of events per minute on a single small instance and gives you transactional enqueue for free.
  • Webhooks are at-least-once, never exactly-once. Deduplication belongs in your database as a unique index, not in application logic that checks “does this exist?” before writing.
  • Retries need exponential backoff with jitter plus a dead letter table you actually look at. An unmonitored dead letter queue is a data loss mechanism with extra steps.

The numbers you’re actually working against

Every provider publishes a delivery timeout and a retry schedule. Almost nobody reads them until something breaks. The ones we deal with most:

  • GitHub: 10 second timeout per delivery. Slow endpoints get flagged and, historically, repeatedly failing webhooks get disabled.
  • Shopify: 5 second timeout, 19 retry attempts spread over roughly 48 hours. Keep failing and Shopify removes the subscription entirely, which means you stop getting orders and nobody tells your customer.
  • Stripe: retries with exponential backoff for up to three days in live mode. Generous, which is exactly why teams get away with slow handlers for months before a dependency wobbles.
  • Slack: 3 seconds to acknowledge an interaction payload, then it shows the user an error. There is no version of “render a report” that fits in 3 seconds.

Compare that to what an inline handler typically does: one database write (2ms), a Stripe API call to fetch the full customer (180ms on a good day), an email send (300ms to 4s), a CRM push (anything from 200ms to a hard 30s hang), maybe a PDF render. The p50 looks fine. The p99 is a cliff, and providers measure you on the cliff.

There’s a second, nastier property: retries amplify load. If your handler is slow because your database is struggling, the provider’s retries add more concurrent work, which makes the database struggle more. We’ve watched that loop take a site from 300ms responses to a hard 502 in under four minutes.

What belongs inline, and nothing more

The HTTP request is a receipt, not a transaction. Your endpoint’s only job is to take custody of the event durably and say so.

Four steps:

  1. Verify the signature against the raw request body. This must happen inline, before anything else. If you can’t verify it, 400 immediately and don’t store it.
  2. Persist or enqueue with the provider’s event ID as the idempotency key. Durably. In-memory queues lose events on deploy.
  3. Return 202 Accepted. 200 is fine too; providers don’t care. 202 documents intent for the next developer.
  4. Process in a worker. Separate process, separate failure domain, separate scaling.

The raw-body detail bites people constantly. Stripe, Shopify and GitHub all compute HMACs over the exact bytes they sent. Run a JSON parser first, re-serialise, and your signature check fails on any payload with unicode escapes or key ordering you didn’t preserve.

import express from 'express';
import Stripe from 'stripe';
import { Queue } from 'bullmq';

const app = express();
const stripe = new Stripe(process.env.STRIPE_SECRET);
const queue = new Queue('stripe-events', { connection: { host: '127.0.0.1', port: 6379 } });

// express.raw() keeps the untouched bytes; a global express.json() would break the HMAC
app.post('/webhooks/stripe', express.raw({ type: 'application/json' }), async (req, res) => {
  let event;
  try {
    event = stripe.webhooks.constructEvent(
      req.body,
      req.headers['stripe-signature'],
      process.env.STRIPEWEBHOOKSECRET
    );
  } catch (err) {
    return res.status(400).send(signature check failed: ${err.message});
  }

  await queue.add(event.type, { payload: event }, {
    jobId: event.id,        // BullMQ treats jobId as unique, so a replayed delivery collapses
    attempts: 8,
    backoff: { type: 'exponential', delay: 2000 },
    removeOnComplete: 1000,
    removeOnFail: false,    // keep failures for inspection
  });

  res.status(202).end();    // acknowledged before any business logic exists
});

That handler’s p99 is dominated by one Redis round trip. Single-digit milliseconds on a local Redis, maybe 15ms across an availability zone. You are now structurally immune to a slow CRM.

Choosing the queue: Postgres, Redis or managed

The default advice is Redis, and for Node shops running BullMQ that’s reasonable. But if you already have Postgres, start there.

FOR UPDATE SKIP LOCKED has been in Postgres since 9.5 and it turns a table into a perfectly good work queue with concurrent consumers and zero extra infrastructure. More importantly, it gives you transactional enqueue: the job row and the business rows commit together, or neither does. With Redis you have two systems and a window where one succeeded and the other didn’t. That window is where the bugs live.

create table webhook_events (
  id           bigserial primary key,
  provider     text        not null,
  third-party_id  text        not null,
  event_type   text        not null,
  payload      jsonb       not null,
  status       text        not null default 'pending',
  attempts     int         not null default 0,
  run_after    timestamptz not null default now(),
  last_error   text,
  received_at  timestamptz not null default now(),
  -- the whole idempotency story, in one line
  constraint webhookeventsdedupe unique (provider, third-party_id)
);

-- partial index keeps the hot path small even with millions of processed rows
create index webhookeventsclaim_idx
  on webhookevents (runafter)
  where status = 'pending';

Claiming work, batch of 10, safe across any number of workers:

with claimed as (
  select id
  from webhook_events
  where status = 'pending' and run_after <= now()
  order by run_after
  for update skip locked          -- other workers silently skip these rows
  limit 10
)
update webhook_events e
set status = 'processing', attempts = attempts + 1
from claimed
where e.id = claimed.id
returning e.id, e.provider, e.event_type, e.payload, e.attempts;

We have this pattern running on a 2 vCPU managed Postgres handling roughly 40,000 events a day with four worker processes. Claim latency sits under 3ms. It is boring, and boring is the point.

When to reach for something else:

  • Redis / BullMQ: you need delayed jobs, rate limiting per queue, priorities and a decent dashboard without building one. Accept that Redis persistence is weaker than Postgres and that appendonly yes is not optional.
  • SQS: you want someone else to own durability. Watch the defaults: 30 second visibility timeout is too short for most webhook work, and standard queues do not preserve order. Max retention is 14 days, max message size 256KB, so store big payloads in S3 and pass a pointer.
  • Cloudflare Queues or Upstash QStash: you’re on edge runtimes with no long-running process to host a worker. QStash in particular is the pragmatic choice for serverless setups, since it does the retry scheduling for you.
  • Kafka: you need replayable ordered logs across teams. If you’re asking whether you need it for webhooks, you don’t.

At-least-once is a promise, not a bug

Every webhook provider worth using guarantees at-least-once delivery. You will get duplicates. Network partitions, your own 500s, provider replays, someone hitting “resend” in a dashboard. Design for it or get bitten by it.

Deduplicate in the database, not in code

The pattern that keeps failing is a SELECT to check existence followed by an INSERT. Two concurrent deliveries both see nothing and both insert. Use the unique constraint and let the database arbitrate:

insert into webhookevents (provider, third-partyid, event_type, payload)
values ('stripe', $1, $2, $3)
on conflict (provider, third-party_id) do nothing
returning id;

No row returned means it’s a duplicate. Return 202 anyway. The provider gets its acknowledgement and stops retrying, which is exactly what you want.

Make the work itself idempotent too

Storage dedupe protects against duplicate deliveries. It doesn’t protect against a worker crashing halfway through a job that already sent an email. So:

  • Pass an idempotency key to every downstream API that supports one. Stripe, SendGrid and most modern APIs do.
  • Prefer upserts over inserts and absolute assignments over increments. set balance = 500 is replay-safe. balance = balance + 100 is not.
  • Record side effects in the same transaction as the state change when you can. If the email send can’t be transactional, do it last, and store a notified_at timestamp you check first.

Ordering is usually a lie you can ignore

Webhooks arrive out of order. A subscription.updated can land before the subscription.created that caused it. Don’t try to enforce global ordering. Use the version or timestamp in the payload and refuse to apply stale data:

update subscriptions
set status = $2, providerupdatedat = $3
where provider_id = $1
  and providerupdatedat < $3;   -- late-arriving older event becomes a no-op

Where you genuinely need per-entity ordering, partition by key: SQS FIFO message groups, or in Redis, a queue per tenant. Global FIFO across all events is a throughput ceiling you’ll regret.

Backoff, jitter and a dead letter queue you read

A webhook processor needs a retry policy with three properties: exponential backoff, jitter, and a visible give-up point.

// 2s, 4s, 8s ... capped at 1 hour, with +/- 20% jitter so retries don't synchronise
function nextRunAfter(attempts) {
  const base = Math.min(2 ** attempts * 1000, 3600000);
  const jitter = base  (Math.random()  0.4 - 0.2);
  return new Date(Date.now() + base + jitter);
}

Jitter matters more than people expect. If a downstream API goes down for two minutes and 800 queued jobs all fail at once, a pure exponential schedule sends all 800 back simultaneously, repeatedly. Add jitter and the load spreads.

Then distinguish error classes. A 422 from a downstream API means the payload is wrong and no amount of retrying fixes it: fail fast to the dead letter table. A 503 or a timeout is transient: retry. Treating every error the same means poison messages burn your entire retry budget while real failures wait behind them.

Set a cap. Eight attempts over roughly two hours is our default for webhook work. After that the row moves to status = 'failed', an alert fires with the event ID and the last error, and a human decides. We run a daily digest that counts failed rows per provider. If that number is ever non-zero and nobody noticed, the queue is not a safety net. It’s a quiet hole.

What to measure

Four metrics tell you almost everything:

  • Queue depth by provider. Rising steadily means workers can’t keep up. Alert on sustained growth, not absolute value.
  • Oldest pending age. Better than depth, because 50,000 events processed in 20 seconds is fine and 12 events stuck for an hour is not.
  • Endpoint p99 response time. Should be flat and tiny. If it moves, something crept back inline.
  • Attempts distribution. A sudden shift from mostly-1 to mostly-3 is your early warning that a downstream dependency is degrading.

Log the provider event ID on every line related to that event. When a client asks why order 8812 has no confirmation email, you want one grep to give you the delivery, every attempt, and the final error. That one convention has saved entire afternoons.

The WordPress and shared hosting version

Plenty of client sites don’t have a worker process. You can still do this properly.

Do not use wpschedulesingle_event for webhook work. WP-Cron only fires when someone visits the site, so a low-traffic site can sit on a “background” job for 40 minutes, and under load it spawns overlapping loopback requests. Use Action Scheduler instead. It ships with WooCommerce, has its own tables, processes in batches (25 actions per batch by default), handles claiming properly and survives concurrency.

addaction( 'restapi_init', function () {
    registerrestroute( 'app/v1', '/webhooks/provider', [
        'methods'             => 'POST',
        'permission_callback' => '__return_true', // we verify by HMAC, not by WP auth
        'callback'            => function ( WPRESTRequest $request ) {
            $raw = $request->get_body();
            $expected = hashhmac( 'sha256', $raw, PROVIDERWEBHOOK_SECRET );

            if ( ! hashequals( $expected, $request->getheader( 'x-provider-signature' ) ?? '' ) ) {
                return new WPRESTResponse( [ 'error' => 'bad signature' ], 400 );
            }

            $event = json_decode( $raw, true );

            // uniquescheduledaction prevents a replay creating a second job
            asenqueueasync_action(
                'appprocessprovider_event',
                [ 'event_id' => $event['id'], 'payload' => $raw ],
                'webhooks',
                true
            );

            return new WPRESTResponse( [ 'queued' => true ], 202 );
        },
    ] );
} );

Two caveats from experience. Action Scheduler’s tables grow fast and its default retention is 30 days, which on a busy store is millions of rows. If you’re already fighting database bloat from revisions and transients, budget for this too. And on genuinely cheap shared hosting, PHP workers are so constrained that async processing helps less than you’d hope. That’s a hosting choice, not an architecture one.

If the site is static or you’re only capturing form submissions and fanning them out to a few destinations, you may not need to own any of this. A hosted endpoint like WebForms already does the receive-and-deliver part with retries behind it, which is the same argument we make about one form feeding many destinations: don’t build a queue to solve a problem a managed endpoint already solved.

Frequently Asked Questions

Should I ever return a non-2xx to a webhook provider?

Yes, in exactly one case: the signature failed verification, where a 400 is correct and you don’t want a retry. For everything else, once you’ve durably stored the event, return 202 even if processing later fails. Returning 500 to trigger a provider retry outsources your retry logic to someone whose schedule and give-up point you don’t control.

Is Redis or Postgres the better backing store for a webhook queue?

Postgres if you already run it, because transactional enqueue eliminates the class of bugs where the job exists but the data doesn’t (or vice versa). Redis with BullMQ once you need delayed jobs, per-queue rate limits and a dashboard you didn’t write. Latency difference at webhook volumes is irrelevant: both claim work in single-digit milliseconds.

How do I stop duplicate emails when the same event is delivered twice?

Store the provider’s event ID with a unique constraint and use on conflict do nothing, so the second delivery never creates a second job. Then make the job itself replay-safe by recording a notified_at timestamp before or alongside the send, and passing an idempotency key to your email provider. Belt and braces, because the worker can crash mid-job.

What about serverless? I have no long-running process for workers.

Use a managed scheduler that calls you back: Upstash QStash, Cloudflare Queues with a consumer Worker, or SQS triggering a Lambda. The endpoint verifies and publishes, the platform handles retry timing and the dead letter queue. Just remember SQS’s default 30 second visibility timeout is shorter than many jobs need, and a job that outlives it gets delivered again while still running.

Do I need ordering guarantees for async processing of webhooks?

Almost never globally. Include the payload’s own version or updated-at value in your write conditions so older events become no-ops, which handles out-of-order arrival without any queue-level machinery. If a specific entity truly needs sequential handling, partition by that entity key using FIFO message groups or a per-tenant queue rather than serialising everything.

If you’re currently doing work inline, the migration is smaller than it looks. Add an events table with a unique constraint, move the existing handler body into a worker function unchanged, and have the endpoint insert and return 202. One afternoon of work. Then spend the next day on error classification and the dead letter alert, because that’s the part that determines whether this is a queue or just a slower way to lose data.

Measure your endpoint’s p99 before and after. If it isn’t under 50ms afterwards, something is still running inline that shouldn’t be.

action scheduler wordpress webhooks background job webhook processing dead letter queue exponential backoff jitter postgres for update skip locked queue stripe webhook timeout retries webhook idempotency key deduplication webhook queue webhook queue architecture