# What should happen when an integration fails?

> Plan visible exceptions, safe retries, and reconciliation so a failed connection becomes a manageable task.

- Published: Sep 25, 2026
- Topic: Software decisions
- Author: RDR Digital
- Original: https://rdr.digital/blog/when-an-integration-fails

A connection can work during a demonstration and still leave staff unsure what to do when an update fails. The missing requirement is often operational: who sees the problem, what state is the record in, and how can someone recover it without creating another mistake?

Design those answers alongside the successful path.

## Distinguish the business state from the delivery state

Consider an illustrative paid order that must reach a warehouse. Payment confirmation, receipt of a notification, and creation of a warehouse order are separate milestones.

Ask your team to show each milestone explicitly. “Sent” is not enough if it only means a request left one system. Staff need to know whether the receiving system accepted the record and whether the next business action happened.

Keep a stable business reference across the connection so support can find the same order on both sides. Agree what information operators may see without exposing unnecessary customer details.

## Expect a message to arrive again

Stripe's [webhook documentation](https://docs.stripe.com/webhooks) states that event delivery order is not guaranteed and that an endpoint may receive the same event more than once. It recommends handling duplicate events. These are documented Stripe behaviors and a useful example of why an integration brief must specify more than the normal sequence.

Ask a developer how a repeated message will be recognized and how the receiving operation avoids an unintended second effect. The technical term often used is **idempotency**: repeating the same operation has the same intended effect as performing it once. Whether an operation offers that property needs to be checked.

A “retry” button should not be accepted until its effect is demonstrated with duplicate and partially completed work.

## Decide which failures need a person

Do not treat every rejection as temporary.

| Situation                                | Question for the recovery design                                        |
| ---------------------------------------- | ----------------------------------------------------------------------- |
| Receiving service is briefly unavailable | Can the operation be retried safely, and when does someone get alerted? |
| Required customer reference is missing   | Who corrects the record before it is tried again?                       |
| The same order already exists            | How does the connection confirm it is the same business operation?      |
| One of several steps completed           | How does recovery resume without repeating completed effects?           |

Set a boundary for automatic attempts. After that boundary, put the record somewhere an operator can find and resolve it. Agree who owns that queue during normal hours and outside them if the business requires coverage.

## Build a usable exception view

Include the business reference, first failure time, most recent attempt, current reason, and the next permitted action. Translate a technical error into an actionable explanation where possible, while preserving detail for the maintainer.

Keep an audit of corrections and recovery actions. Staff should be able to distinguish a record that resolved automatically from one someone changed manually.

In the illustrative order flow, the operator might need to correct an account mapping before resubmitting. They should not have to guess whether the warehouse already received the order.

## Reconcile after recovery

When a service returns, verify the business records rather than relying only on a green connection indicator. Compare the intended orders with the accepted warehouse orders and investigate unmatched references.

Rehearse this process before launch: interrupt a test connection, introduce an invalid record, repeat a notification, and recover the backlog. Record what the operator actually had to do.

Include the resulting runbook in your [maintenance arrangement](https://rdr.digital/blog/software-maintenance-after-launch). A connection is ready for routine use when its exceptions are understandable as well as its normal path.

## Sources

- [Stripe Documentation: Webhooks](https://docs.stripe.com/webhooks)
