The Webhook Bug That Gives Cancelled Users Free Access
A customer pays for a subscription. A few seconds later they change their mind and cancel. Your payment provider sends two webhooks: "payment succeeded", then "subscription cancelled". Your server receives them in the opposite order. The cancellation is applied first, the late payment event re-activates the account, and the customer now has a cancelled subscription with full access. Nobody pays for it, and nobody notices.
I have found this bug, or one of its close relatives, in almost every subscription system I have reviewed. It does not throw exceptions, it does not show up in error tracking, and it passes every test that sends events in a tidy sequence. It shows up months later as a gap between what the payment provider says and what the product database says - or as a support ticket from a paying customer who was locked out, which is the same bug pointing the other way.
Webhooks are notifications, not commands. Treat them as "something changed, go look" - never as "apply this change now, in this order".
Why webhooks arrive out of order
A webhook is not a stream. Each event is an independent HTTP request, sent and retried on its own schedule. Stripe states this explicitly in its documentation: delivery order is not guaranteed. The same is true for PayPal, Paddle, Mollie, Adyen and practically every other provider.
The usual reasons in production:
- Retries. Your endpoint returned a
500or timed out once - a deploy, a database hiccup, a slow query. The provider retries that event minutes or hours later, long after newer events have been delivered successfully. Stripe keeps retrying for up to three days. - Parallel delivery. Events are dispatched concurrently. Two requests leave the provider a few milliseconds apart and are handled by two different PHP-FPM workers, or two different servers behind a load balancer.
- Your own queue. If you hand webhooks to background workers, which you should, several workers consume in parallel and the order is lost again on your side.
In other words: out-of-order delivery is not an edge case you might hit under heavy load. It is the normal operating mode, and a system that depends on order is correct only by luck.
The naive handler
This is the shape I see most often. It is short, readable and wrong:
public function __invoke(Request $request): Response
{
$event = json_decode($request->getContent(), true);
$customerId = $event['data']['object']['customer'];
match ($event['type']) {
'invoice.paid' => $this->users->setActive($customerId, true),
'customer.subscription.deleted' => $this->users->setActive($customerId, false),
default => null,
};
return new Response('', 200);
}
Every event overwrites the state with whatever it implies, so the last event to arrive wins - not the last event that happened. Swap the arrival order and you hand out free access. There are four more problems hiding in these ten lines:
- No signature verification. Anyone who knows the URL can activate any account.
- No deduplication. Providers deliver at-least-once, so the same event can be processed twice.
- All work happens inside the request. If it is slow, the provider times out and retries, which creates more duplicates and more reordering.
- A boolean
activeflag throws away the information you need to make a correct decision later.
First fix: only apply newer events
Every event carries the time it was created at the provider. The obvious improvement is to store the timestamp of the last applied event and reject anything older. Most write-ups stop with a version like this:
$subscription = $this->subscriptions->find($subscriptionId);
if ($eventCreatedAt <= $subscription->lastEventAt) {
return; // stale event, ignore it
}
$this->subscriptions->update($subscriptionId, $newStatus, $eventCreatedAt);
This is a real improvement, and in a code review I would still not approve it. Two things are broken.
Check-then-write is a race condition
Two events for the same subscription are processed at the same moment by two workers. Both read the same lastEventAt, both pass the check, both write. The one that commits last wins - and we are back to arrival order. The check and the write must be one atomic operation, which the database gives you for free with a conditional update:
$applied = $this->connection->executeStatement(
'UPDATE subscriptions
SET status = :status, last_event_at = :eventAt
WHERE provider_subscription_id = :id
AND last_event_at < :eventAt',
['status' => $newStatus, 'eventAt' => $eventCreatedAt, 'id' => $subscriptionId],
);
if (0 === $applied) {
$this->logger->info('Stale webhook ignored', ['subscription' => $subscriptionId]);
}
The row-level lock MySQL takes for the UPDATE serializes concurrent writers, and the WHERE clause re-evaluates the condition against the committed value. No application-level locking needed.
Timestamps have one-second resolution
Stripe's created field is a Unix timestamp in seconds. A checkout flow that creates a subscription, pays the first invoice and updates the subscription can emit several events within the same second. With <= you drop legitimate events; with < you are back to arrival order for everything inside that second. And event creation time is not even the same thing as the order in which the provider changed its internal state. The timestamp guard narrows the window. It does not close it.
The robust approach: webhooks as triggers
The design I implement for clients, and the one Stripe itself recommends for state-sensitive flows, removes the ordering problem instead of managing it. The payment provider is the source of truth for subscription state. A webhook only tells you that the state of a specific subscription changed. The handler then fetches the current state from the provider's API and stores that.
With this model, the order of events no longer matters. A late "payment succeeded" event triggers a fetch, the API returns canceled, and canceled is what gets stored. Whichever event is processed last, the result is the latest truth.
The full pipeline has three stages, and each stage has exactly one job.
Stage 1: verify, store, acknowledge
The endpoint does as little as possible. It verifies the signature, writes the raw event to the database and returns 200. Storing the raw payload first gives you an audit trail, deduplication by event ID and the ability to replay events after a bug fix.
CREATE TABLE webhook_events (
id INT NOT NULL AUTO_INCREMENT PRIMARY KEY,
provider VARCHAR(32) NOT NULL,
event_id VARCHAR(255) NOT NULL,
event_type VARCHAR(128) NOT NULL,
object_id VARCHAR(255) NOT NULL,
created_at DATETIME NOT NULL,
received_at DATETIME(6) NOT NULL,
processed_at DATETIME(6) NULL,
outcome VARCHAR(32) NULL,
payload JSON NOT NULL,
UNIQUE KEY uniq_provider_event (provider, event_id),
KEY idx_object (provider, object_id)
);
#[Route('/webhooks/stripe', methods: ['POST'])]
public function stripe(Request $request): Response
{
try {
$event = Webhook::constructEvent(
$request->getContent(),
(string) $request->headers->get('Stripe-Signature'),
$this->stripeWebhookSecret,
);
} catch (SignatureVerificationException|\UnexpectedValueException) {
return new Response('', Response::HTTP_BAD_REQUEST);
}
try {
$this->webhookEvents->store('stripe', $event, $request->getContent());
} catch (UniqueConstraintViolationException) {
return new Response('', Response::HTTP_OK); // duplicate delivery, already stored
}
$this->bus->dispatch(new ProcessWebhookEvent('stripe', $event->id));
return new Response('', Response::HTTP_OK);
}
The unique key on (provider, event_id) handles at-least-once delivery at the database level. A duplicate is acknowledged and dropped before any business logic runs. Because nothing slow happens in the request, the endpoint answers in milliseconds, which reduces provider retries and therefore reordering in the first place.
One detail matters for durability: the event must be committed before you return 200. If you acknowledge first and store later, a crash in between loses the event for good - the provider has no reason to send it again.
Stage 2: lock, fetch, store current state
A worker picks up the message. It serializes work per subscription with a lock, fetches the current subscription from the provider and writes it to the local table. The event payload is used only to find out which subscription to refresh.
#[AsMessageHandler]
final readonly class ProcessWebhookEventHandler
{
public function __construct(
private WebhookEventRepository $events,
private SubscriptionRepository $subscriptions,
private StripeClient $stripe,
private LockFactory $locks,
) {
}
public function __invoke(ProcessWebhookEvent $message): void
{
$event = $this->events->get($message->provider, $message->eventId);
$subscriptionId = $event->subscriptionId();
if (null === $subscriptionId) {
$this->events->markProcessed($event, 'ignored');
return;
}
$lock = $this->locks->createLock('subscription-sync-' . $subscriptionId, ttl: 30);
$lock->acquire(blocking: true);
try {
$remote = $this->stripe->subscriptions->retrieve($subscriptionId);
$this->subscriptions->upsertFromProvider(
providerSubscriptionId: $remote->id,
customerId: $remote->customer,
status: $remote->status,
currentPeriodEnd: new \DateTimeImmutable('@' . $remote->items->data[0]->current_period_end),
cancelAtPeriodEnd: $remote->cancel_at_period_end,
);
$this->events->markProcessed($event, 'synced');
} finally {
$lock->release();
}
}
}
Why the lock, if we always fetch the latest state? Because "fetch, then write" is again two steps. Worker A fetches active, worker B fetches canceled, B writes, then A writes its older snapshot. The lock guarantees that only one sync per subscription runs at a time, so the last write is always based on the last fetch. The lock store must be shared across all workers - Redis or the database, never the local filesystem.
If the provider API is down, the handler throws, Symfony Messenger retries with backoff, and permanently failing messages end up in a failure transport where someone can inspect them. The raw event is safe in the database either way.
The same structure works one-to-one in Laravel: a queued job instead of a message handler, Cache::lock() instead of the Lock component, and a unique index on the events table for deduplication.
Stage 3: derive access, do not store it
The last change is the least technical and the most important. Do not store active = true. Store what the provider tells you - status, end of the paid period, whether it cancels at period end - and compute access from it in one place:
public function grantsAccess(\DateTimeImmutable $now): bool
{
return \in_array($this->status, ['active', 'trialing'], true)
&& $this->currentPeriodEnd > $now;
}
This makes business rules explicit instead of scattering them across webhook branches. Does a customer who cancels keep access until the end of the paid period? Do we grant a grace period for past_due? Those are product decisions, and they now live in a single, testable method instead of being an accidental side effect of which event arrived last.
The safety net: reconciliation
Even a correct pipeline runs in an imperfect world. Webhook endpoints get misconfigured during a migration, a signing secret is rotated in one environment but not the other, a deploy drops events for an hour. That is why every billing integration I build gets a scheduled reconciliation job: once a night, list subscriptions from the provider, compare them with the local table, fix drift and report every correction.
The report matters more than the fix. A reconciliation job that corrects zero records every night is proof that the system works. One that suddenly corrects forty is an alert that something upstream broke - and you learn about it the next morning instead of at the next financial audit.
Why storing raw events pays off
Teams sometimes see the webhook_events table as storage overhead. In practice it is the most valuable table in the billing domain:
- Replay. After fixing a bug in the handler, re-dispatch all affected events. No manual data correction, no SQL scripts against production.
- Audit. When a customer disputes a charge or claims they cancelled, you can show exactly what the provider sent and when you processed it.
- Debugging. "Why does this user have access?" becomes a query instead of an investigation.
- Monitoring. Rows with
processed_at IS NULLolder than a few minutes are a precise, cheap health check for the whole pipeline.
Keep the payload as received, set a retention period that fits your accounting and data protection requirements, and purge older rows with a scheduled job.
A checklist for decision makers
If your product earns money through subscriptions, these questions tell you within one meeting whether your billing integration is trustworthy. You do not need to read the code to ask them.
| Question | What a good answer sounds like |
|---|---|
| What happens if two webhooks for the same customer arrive in the wrong order? | "Nothing. We fetch the current state from the provider, so order does not matter." |
| What happens if the provider sends the same event twice? | "It is deduplicated by event ID at the database level." |
| How do we know our access data matches what customers actually pay for? | "A nightly reconciliation compares both and reports every difference." |
| If we find a bug in billing logic, how do we repair past data? | "We replay the stored raw events through the fixed handler." |
| Where is the rule that decides whether a customer has access? | "In one method, with tests. Not spread across webhook handlers." |
| Who can call our webhook endpoint? | "Anyone can call it, but only requests with a valid provider signature are accepted." |
If the answer to any of these is a pause followed by "I would have to check", the integration is worth a closer look. Revenue leakage from billing bugs is silent by nature: it never crashes, it only compounds.
Conclusion
The free-access bug is not caused by a missing if statement. It is caused by a wrong mental model: treating webhooks as an ordered sequence of commands. Once you treat them as unordered, possibly duplicated hints that something changed, the solution follows naturally - verify and store every event, acknowledge fast, process asynchronously, serialize per subscription, fetch the current state from the source of truth, derive access in one place, and reconcile regularly.
None of this is exotic. It is a few hundred lines of code and one or two tables. But it is the difference between a billing system you trust and one you hope is right.
Is your billing integration leaking revenue?
I help SaaS companies and platforms build payment integrations that stay correct under retries, outages and growth - with Stripe, PayPal and other providers, in PHP, Symfony and Laravel. As a lead engineer I do not stop at a recommendation: I work in your codebase with your team until the fix runs in production.
Billing Integration Audit
A focused review of your webhook handling, subscription state and access logic: ordering, duplicates, race conditions, security and reconciliation. Delivered in writing with a prioritized list of fixes, understandable for engineering and management.
Payment Integration Implementation
Design and implementation of new subscription and payment flows - webhook ingestion, asynchronous processing, entitlement logic, reconciliation and monitoring from day one.
Repair and Data Correction
Hands-on work on integrations that are already drifting: find affected customers, correct the data safely, close the root cause and put monitoring in place so it does not happen again.
A response usually comes the same day. Remote and on-site, Germany and Europe.