Webhook retry logic is a reliability feature with business consequences. Cost depends on event volume, destination variety, timing promise, duplicate protection, replay controls, observability, and the team's recovery process.

Basic retry scope

A simple sender needs event ID, destination, timeout, retry count, backoff, response classification, and final state. Define what happens when the destination may have processed the request but did not respond.

Queue and idempotency work

Cost increases with durable queue, tenant isolation, rate control, jitter, idempotency key, ordering, dead-letter, replay, and manual approval. The delivery attempt history checklist helps define the evidence behind each retry decision.

Destination health

Budget for endpoint configuration, health status, pause, disable, circuit breaker, notification, and support view. Keep one slow destination from affecting unrelated event streams.

Security and privacy

Include signature, secret storage, payload minimization, access control, retention, redaction, tenant scope, and safe alert content. Do not log full sensitive payloads simply to diagnose retries.

Testing and operations

Test timeout after side effect, invalid response, provider outage, duplicate event, out-of-order event, poison payload, queue full, replay, and recovery. Plan monitoring, runbooks, and owner training.

Maintenance cost

Allow for new event types, endpoint changes, retry policy tuning, storage, provider updates, dead-letter review, and incident support. Ask for assumptions, limits, and ongoing ownership.

Webhook retries growing without safe visibility? Ask Vertinus to scope delivery reliability, replay, and operations together.