Product
Webhook ingress that behaves like infrastructure.
Marevans sits between payment processors, SaaS vendors, and your services. Every event is verified, stored, routed, retried, and visible - before your application has to care.
Architecture
One control plane, many destinations.
Providers hit a single HTTPS endpoint in your network. Marevans acknowledges only after persistence, then delivers to each destination on its own schedule.
Edge security
Signature verification at the edge.
Forged payloads never enter storage. Each source uses the provider’s native scheme or HMAC you configure.
Stripe, GitHub, Shopify, and dozens of others sign the raw body with a shared secret. Marevans recomputes the signature, checks clock skew, and returns 401 on mismatch - the same response the provider would see if they hit your app directly with bad credentials.
# sources/stripe.yamlpath: /in/stripeverify: scheme: stripe secret: ${ssm:/webhooks/stripe/signing-secret} tolerance: 300s # reject replays older than 5 minutesCustom partners can use HMAC-SHA256 over the body, a static bearer token, mTLS, or an IP allowlist. Verification runs before any queue write, so rejected traffic never consumes retention or pollutes your event log.
Durability
Persist, then acknowledge.
Providers only get a 2xx after the event is on durable storage in your account.
The classic webhook failure mode is acknowledging too early. Your handler returns 200, the provider stops retrying, and then your deploy rolls back. Marevans separates acceptance from delivery: acceptance means the event is safe on disk; delivery is a background concern with its own retry policy.
Storage uses encrypted volumes you control (cloud KMS or your key management). Metadata - event id, source, type, attempt history - is kept longer than raw payloads if you configure retention that way.
Delivery
Retries with backoff and jitter.
Per-destination timeouts, attempt budgets, and schedules - no hand-rolled Sidekiq jobs.
When your service returns 5xx, times out, or is unreachable, Marevans schedules another attempt with exponential backoff and jitter. You set the ceiling: typical configs start at 10s and cap at an hour, with 8–12 attempts before dead-lettering.
destinations: billing-svc: url: http://billing.internal:8080/webhooks/stripe timeout: 10s retry: max_attempts: 12 backoff: exponential initial_interval: 10s max_interval: 1h on_exhausted: dead_letter4xx responses can be treated as permanent failures (skip retry) or retried depending on your policy - useful when a bad deploy returns 404 for an hour and you want the queue to keep trying.
Idempotency
Deduplication that survives provider retries.
Map a field from the payload or a header to an idempotency key. Duplicates are collapsed before delivery.
Webhooks are at-least-once. Providers retry on timeouts; load balancers replay; your own replays can overlap. Marevans extracts an idempotency key (for example Stripe’s evt_… id) and suppresses duplicate deliveries within a window you define - often 72 hours.
Your handlers still receive X-Marevans-Idempotency-Key so application code can guard side effects even if a duplicate slips through after the window expires.
Fan-out
Routing and fan-out.
Match on source, event type, or JSON path. One inbound event can fan out to many internal services.
Route rules are declarative. Send invoice.paid to billing and analytics, send customer.subscription.updated only to entitlements, and drop health-check noise at the edge.
routes: - name: stripe-invoice-paid match: source: stripe type: invoice.paid deliver: - billing-svc - warehouse-svc - notify-svcEach destination gets an independent delivery record in the event log - success rate and latency per service, not blended averages that hide a broken downstream.
Operations
Dead-letter queue with full context.
When attempts are exhausted, the event moves to a DLQ you can browse - not a log line you grep for later.
Every dead-lettered event retains the original headers and body (subject to redaction rules), each delivery attempt, HTTP status, response snippet, and timing. Operators can see whether failures cluster on one destination or one event type after a deploy.
DLQ retention is configurable separately from successful deliveries - often 90 days - so post-incident review does not require log archaeology in CloudWatch.
Recovery
Replay one event or ten thousand.
Fix the bug, then re-drive delivery without asking providers to resend.
Replays are explicit actions in the event log: by event id, by filter (source + type + time range), or for everything currently in the dead-letter queue. Replayed deliveries are labeled in the UI so you can distinguish first pass from recovery traffic.
Replays respect deduplication settings by default, with an override when you intentionally need to force a second delivery to a fixed handler.
Visibility
Observability built for incidents.
Search, filter, and alert on the metrics that matter when webhooks fail at 2 a.m.
- Searchable event log with payload preview (redacted fields masked)
- Success rate, p95 delivery latency, and queue depth per destination
- Failure breakdown by HTTP code and timeout
- Export to your existing stack via OpenTelemetry metrics and structured logs
Alerts can fire when success rate drops below a threshold or DLQ depth grows - the same signals your on-call already uses for service health, applied to webhook plumbing.
Privacy
Retention and redaction.
Drop card numbers before storage, purge bodies on a schedule, keep metadata for audits.
Redaction rules run before storage. Fields you drop never hit disk; masked fields store a placeholder; hashed fields allow joins without storing raw PII.
redaction: apply: before_storage rules: - match: { source: "*" } fields: - path: $..card.number action: drop - path: $..email action: maskPayload retention defaults are measured in days; metadata can live longer for compliance reporting. See security and privacy for data-flow detail.
Deploy the gateway once, point every provider at it.
Start with one source - usually Stripe or Shopify - and expand routes as teams adopt the shared event log.