Reliability
The system keeps working when nobody is watching.
A process can report itself healthy while failing most of its work. That is why every system here checks the data, not only the daemon.
- Last check
- 34 sec ago
- Automated tests
- —
- Failed jobs, 24 h
- —
- Alerting
- Armed
Check the data, not the daemon.
A process can report itself healthy while failing most of its work. We have watched that happen for thirty-one hours before anything noticed. Every system checks the resulting data rather than trusting process status alone.
Measure before setting thresholds.
An alert that fires constantly becomes background noise, and an ignored alert is the same as no alert. So the normal gap between records is measured first and the limit derived from it, rather than guessing a number that feels about right.
Separate transient failures from real ones.
A dropped connection should be retried. A missing file should not — retrying it wastes time and hides the cause. The distinction is written down in each system, because the wrong choice in either direction is invisible until it matters.
Document why, not only what.
What a system does and why each check exists, so whoever inherits it can operate it without us. You should not depend on a vendor through not understanding your own system. If that means you eventually stop needing us, that is the correct outcome.
Reporting
Numbers with their freshness attached.
A figure without a timestamp is a guess. Every panel states when its data last arrived, so a stale dashboard announces itself instead of misleading you.
| Report | Due | State |
|---|---|---|
| Daily reconciliation | 06:00 | Sent |
| Supplier price delta | 07:30 | Sent |
| Weekly summary | Mon 08:00 | Queued |
| Stale-data sweep | Hourly | 1 warning |