WhatsApp is down: what happens to your Shopify order notifications

When Meta's API fails, are Shopify order WhatsApps lost, delayed or sent twice? Shopify's webhook retries, Beacon Messenger's queue, retries and dedupe key.

Five years ago today, on 4 October 2021, WhatsApp, Facebook and Instagram disappeared from the internet for about six hours after a routing configuration change at Meta. Stores that had moved their order notifications onto WhatsApp spent an afternoon not knowing whether confirmations had gone out, would go out later, or would go out twice when everything came back. Outages that long are rare. Partial ones, where the Cloud API returns errors for twenty minutes or throttles a number that is suddenly sending more than usual, happen often enough that a notification app should have a clear answer to three questions: what gets lost, what gets delayed, and what gets duplicated.

This post gives that answer for Beacon Messenger and, because half of the chain is Shopify's, for Shopify's webhooks too. The short version is that a notification is a row in a queue with a key, not a request fired at Meta the moment an order lands, and most of what follows is the consequence of that one decision.

Three links that can each break

An order confirmation on WhatsApp passes through three systems. Shopify notices the order and sends a webhook to the app. The app builds the message and calls Meta's Cloud API. Meta delivers it to the customer's phone. Each link fails differently, and a failure in one does not look like a failure in another, which is why "WhatsApp is down" is less useful than it sounds.

  • Shopify to app: your app's server is restarting, overloaded or unreachable. The order exists, but nobody has been told about it yet.
  • App to Meta: the Cloud API returns an error, times out or throttles. The app knows about the order and has a message ready, but cannot hand it over.
  • Meta to phone: Meta accepted the message, but the customer's phone is off, out of coverage or has WhatsApp uninstalled. The message sits at sent and may become delivered hours later; this link is covered in the delivery verification post and is not an outage at all.

Shopify's side: webhooks are retried for two days

Shopify does not fire a webhook once and forget it. According to Shopify's webhooks documentation, a delivery has to be acknowledged with a 2xx response within five seconds; if it is not, Shopify retries the delivery 19 times over the following 48 hours, and only after 19 consecutive failures does it remove the subscription. So an app server that is down for an hour misses nothing: the orders placed during the hour arrive as webhooks once it is back, in the order Shopify retries them.

The five-second rule has a consequence for how a well-behaved app is built. If the app tried to call Meta inside the webhook request, a slow Meta response would push the app past five seconds, Shopify would count the delivery as failed and retry it, and the app would now have the same order twice. The right design is to acknowledge the webhook immediately, store what needs to happen, and do the sending somewhere else. That is what the queue is for.

Beacon Messenger's side: a queue, a lock and a key

When an order webhook arrives, Beacon Messenger verifies it, writes one row per enabled automation into a scheduled-message table and returns a 200 to Shopify. The row holds the trigger, the time it is due and a snapshot of the order details the template needs. Order confirmations are due immediately; cash-on-delivery confirmations are due after the delay you set; abandoned-checkout reminders are due after the reminder delay. A poller wakes every thirty seconds, claims due rows by stamping a lock on them, and sends each one.

The send has a fifteen-second timeout. If Meta answers with a server error (any 5xx status), with error code 4, which Meta documents as "API Too Many Calls", or with code 130429, "Rate limit hit", the row is released back to the queue with a later due time: two minutes after the first failure, four after the second. Meta's error code reference lists both. After the third failed attempt the row is written to the message log as failed, with Meta's error code and message attached, and it stops retrying on its own. Errors that are not worth retrying, a rejected template, a malformed phone number, a paused template, go straight to failed on the first attempt, because trying again would produce the same answer.

Rows that reach failed are not gone. The Messages page lists them with the error, and each has a Retry action that re-runs the send from the stored order snapshot, right now, regardless of how old the order is. After a real outage, filtering Messages to failed and retrying the handful of rows that exhausted their attempts is the whole recovery procedure.

Why nothing is sent twice

Duplicates are the failure merchants fear most, because a customer who receives the same confirmation twice assumes they were charged twice. Beacon Messenger prevents them with a dedupe key: one string per shop, trigger and object, so an order confirmation for order 1042 always has the same key however many times Shopify delivers the webhook. The key is also the id of the message log row. Before any send, the app checks whether a log row with that key already exists in a sent state, and if it does, the queue item is dropped without calling Meta.

Two other paths are closed by the same mechanism. A webhook that Shopify retries after the app had in fact processed it (the 200 was lost on the network) re-queues the same key, which simply moves the existing row rather than creating a second one. And when two copies of the app server run at once, each poller claims a row by updating it only if the lock has not changed since it was read, so a row is sent by whichever process won the claim and skipped by the other. Shipping updates work the same way with a key per fulfillment, which is why a later tracking edit does not resend the original message; the split shipments post has the details.

What the customer experiences after a long outage

A confirmation that could not be sent for ninety minutes arrives ninety minutes late, once, with the order details as they were when the order was placed. That is the intended behaviour; a late confirmation is still useful, and the alternative, dropping it, leaves the customer with nothing. Cash-on-delivery confirmations keep their original due time, so if the outage ends before the delay has elapsed nothing was ever late.

Two situations get a different answer. If the order was cancelled during the outage, Shopify's cancellation webhook removes the queued confirmation before it can be sent, so the customer does not receive a confirmation for an order that no longer exists. And an abandoned-checkout reminder that fell due during the outage is dropped if the customer completed the order in the meantime, because the order webhook cancels the reminder by its checkout key. What the outage cannot do is make a message skip the plan's monthly cap or the consent check; those run at send time, not at queue time, so a message held over an outage is judged against the same rules as one sent on time.

What to do while it is happening

  1. Check Meta's status page for the WhatsApp Business Platform before touching anything in your store. If Meta is reporting an incident, the problem is not your template, your number or your connection.
  2. Do not disconnect and reconnect the WhatsApp number, and do not re-enter the access token. The connection is not broken; the API behind it is. Reconnecting during an incident can fail in confusing ways and leave you debugging the wrong thing afterwards.
  3. Do not send confirmations by hand from a phone unless an order is urgent. The queue will send them when the API recovers, and a manual message plus a queued one is the only way to get a duplicate.
  4. When the incident is over, open Messages, filter to failed, and use Retry on anything that exhausted its attempts during the outage. Rows that were still retrying will have finished on their own.
  5. Remember that the email confirmation from Shopify went out as normal. WhatsApp is the convenient copy; email remains the complete record, and during an outage that is exactly the arrangement you want.

An outage is the one time the design of a notification app is visible from outside. Fire-and-forget designs lose messages when Meta is down and duplicate them when Shopify retries; a queue with a key does neither, at the cost of being a few minutes late. Beacon Messenger is built the second way, on a free plan with 50 automated messages a month, so you can watch the Messages page during the next small incident and see each row retry, succeed and stay single.

Frequently asked questions

If WhatsApp is down, are my Shopify order confirmations lost?

Not with a queue-based app. Beacon Messenger acknowledges Shopify's webhook immediately and stores the message in a queue. Sends that fail with a Meta server error or rate limit are retried after two and then four minutes, and anything that still fails is listed as failed in the Messages page with a Retry action. Shopify, for its part, retries any webhook the app did not acknowledge 19 times over 48 hours.

Can a customer receive the same WhatsApp order confirmation twice?

Not from Beacon Messenger. Every notification has a dedupe key made of the shop, the trigger and the order, and that key is the id of its message log row. A repeated webhook, a retried send or a second server process all resolve to the same key, and the app refuses to send when a sent row with that key already exists.

What do WhatsApp error codes 4 and 130429 mean?

Both are throttling. Code 4 is Meta's general "API Too Many Calls" limit and 130429 is "Rate limit hit" for the WhatsApp Business account's throughput. Neither means the message was wrong; they mean Meta wants it later. Beacon Messenger treats both as retryable and backs off before trying again.

Should I reconnect my WhatsApp number during an outage?

No. If Meta's status page reports an incident on the WhatsApp Business Platform, the connection is fine and the API behind it is not. Reconnecting adds a second variable to debug. Wait for the incident to clear, then retry any failed messages from the log.

More guides