Context
The operating situation
A trading data platform where missing events had to be found and replayed by hand. Every hour of manual recovery extended customer exposure and pressed against contractual service levels.
Challenge
The operating problem
Missing-event recovery relied on manual source identification and replay, delaying restoration and increasing operational exposure.
Approach
Decisions that shaped it
- Detected sequence gaps through writer services and Prometheus telemetry.
- Automated source identification, replay coordination, and recovery-state management.
- Ran post-recovery validation and escalated remaining exceptions for human action.
Solution
What was delivered
An automated replay and validation pipeline: sequence gaps detected in real time, sources identified and replayed without hand-holding, recovery state tracked, and integrity validated after the fact with anything unresolved escalated to a person.
Outcome
Documented outcome
Routine recovery dropped from 48-72 hours to under 30 minutes within the documented platform workflow.
The 48-72 hours to under 30 minutes improvement is bounded to the documented platform recovery workflow. It is not an organisation-wide availability, recovery, or uptime claim, and the former employer stays anonymized.
Controls
What kept it safe to operate
- Post-recovery integrity validation, so a fast recovery is also a correct one.
- Exceptions escalate to a human instead of being silently absorbed.
- Recovery state is explicit and observable rather than inferred from logs.
