Fundamentals

Why Integration Failures Go Undetected for Weeks

Integration failures hide because they are partial, slow and distributed across tools that do not talk to each other. Here are the five patterns that keep them invisible, and what changes detection.

Why Integration Failures Go Undetected for Weeks

Integration failures stay hidden because they are rarely total. A connection that fails completely gets noticed within minutes. A connection that fails for four percent of records, or only for one customer segment, or only after a token refresh, produces no alert and no complaint until the damage has compounded.

Key takeaways
  • Total failures are detected fast. Partial failures are the expensive ones.
  • The gap between first occurrence and first report is where almost all of the cost sits.
  • Detection improves when the signal is correlated across tools, not when more alerts are added.

The detection gap

Every integration failure has two timestamps: when it started, and when a human first understood it was happening. The distance between them is the detection gap, and for integrations it is routinely measured in weeks rather than minutes.

This is not a tooling failure in the usual sense. The signals almost always exist. They sit in different systems, owned by different people, none of whom sees enough of the picture to raise an alarm.

Five patterns that keep failures invisible

1. The failure is partial

A sync that drops four percent of records produces no error rate worth alerting on. The missing records surface later as a reconciliation discrepancy, a customer complaint, or a month-end number that does not tie out. By then the cause is weeks in the past and the logs have rotated.

2. Someone masked the symptom

An engineer notices intermittent failures and adds a retry. The retry works. The error rate falls to zero and the signal disappears, but the underlying fault is still there, now consuming three times the request budget. This is the single most common reason a recurring failure looks like a new one each time it resurfaces. We cover the mechanics in retry storms and idempotency.

3. The evidence is split across tools

The error spike lives in monitoring. The customer complaint lives in support. The vendor's acknowledgement lives in someone's inbox. The workaround lives in a pull request. No single system holds enough to constitute a detection, and no one is paid to assemble them.

4. It only fails on a boundary condition

Token refresh, pagination past a certain depth, a leap second, month end, a record with a non-Latin character. These fail on a schedule that does not match anyone's on-call rotation, and they recover on their own before the investigation starts.

5. Nobody owns the integration

Services have owners. Integrations frequently do not. The connection between your billing system and your CRM was built by someone who has since moved teams, and it is nobody's explicit responsibility until it breaks loudly enough.

What this costs

Industry research consistently puts the share of developer time spent on integration work rather than product work above a third, and the cost of a single hour of customer-facing downtime at six figures for a meaningful share of enterprises. Those numbers are averages across large samples, so treat them as direction rather than measurement. Undetected integration failure is expensive long before it becomes an outage.

What actually closes the gap

Adding alerts does not help; it usually makes things worse by raising the noise floor until real signals are ignored. See alert fatigue in integration monitoring for why threshold-per-metric approaches degrade.

Three things reliably shorten the gap:

  • Correlation over collection. Tie every signal that mentions the same integration to one record, regardless of which tool it came from.
  • Recurrence detection. Treat the third appearance of a signature as a different event from the first. A fault that returns is a different class of problem from a fault that happens once, and it justifies a different response.
  • A named owner per integration. Not a team: a person, recorded in an inventory.

This is the problem Traxivo was built around: watching the record your teams already produce, recognising when scattered signals describe the same underlying fault, and raising it once it is worth a human's attention rather than at every threshold crossing.

How long is too long

There is no portable benchmark for an acceptable detection gap, because it depends entirely on what accumulates while the failure runs. A useful substitute is to ask, per integration, what becomes unrecoverable with time:

  • Bounded by a replay window. If the provider retains events for 24 hours, your detection budget is well under 24 hours. This is a hard constraint, not a target.
  • Bounded by a billing or reporting cycle. If the discrepancy must be caught before close, the budget is the cycle length minus the time to repair and backfill.
  • Bounded by customer tolerance. If customers see it, the budget is however long they will absorb it without opening a ticket, which is usually shorter than teams assume.

Written down this way, the detection target for each integration falls out of its own constraints rather than from an arbitrary service level. It also makes the trade-off explicit: an integration with a 24-hour replay window and no reconciliation is accepting permanent data loss as a matter of policy, whether or not anyone has said so out loud.

Frequently asked questions

How long do integration failures typically go unnoticed?

It varies with how total the failure is. Complete outages are usually caught within minutes. Partial failures (a percentage of dropped records, a single affected segment, a condition that only triggers on token refresh) commonly run for days or weeks before anyone connects the symptoms to a cause.

Why do alerts not catch partial integration failures?

Alerts fire on thresholds applied to a single metric in a single system. A four percent failure rate rarely crosses a threshold, and a retry added to mask the symptom removes the metric entirely while leaving the fault in place.

Who should own an integration?

A named individual, recorded alongside the integration in an inventory. Team-level ownership tends to mean nobody notices the slow failures, because no one person sees enough occurrences to recognise a pattern.

Stop rediscovering the same integration failure

Traxivo correlates the signals your tools already produce into one incident timeline, recognises a recurrence as a recurrence, and drafts the follow-up with the evidence attached. Nothing is sent without a named approver.

See how Traxivo works Browse use cases

Related reading