Insights

Multi-site monitoring is sized on the on-call rota, not on the number of probes

A tool that reports everything produces noise nobody acts on. What decides the value of monitoring is what happens after the alert, and who takes it.

Schematic technical motif, generated from the subject of the publication. It depicts no installation.

Monitoring is deployed in a few weeks. It stops being read within a few months. The mechanism is always the same: the tool reports everything it knows how to report, the volume exceeds what the team can handle, and the team stops looking. From then on, the alert that mattered is buried among those that asked for nothing.

An alert is not monitoring

Monitoring is not measuring. Measuring is the easy part, and the only part the tool does by itself. Monitoring is deciding, before the incident, which deviations call for action, who takes it, within what time, and how anyone knows it was taken.

An alert nobody acts on no longer signals anything. Worse, it hides the one that counted, and it gives the board a sense of coverage that does not exist.

Three questions that come before the tool

Who receives it, and when. An organisation with no on-call rota does not need monitoring that wakes people at night; it needs a state it can read in the morning. An organisation with a rota must know who is reachable, on which channel, and what they are authorised to do alone at three in the morning.

What justifies waking someone. The list is short, and it is written before installation. On most sites it comes to five or eight situations. Everything else is data — available on request, not an alert.

What happens if nobody answers. An alert chain with no second level and no escalation delay is not a chain, it is a notification. The second level is named, and told that it is.

Thresholds come from the action, never from the measurement

A threshold is set by working back from the intervention. The question is not the value at which the measurement becomes abnormal, but the value at which somebody has to travel. The two answers are rarely the same, and the second is the useful one.

A fuel tank is the ordinary illustration. A low level is of no interest in itself. What matters is the point at which remaining autonomy falls below the time it takes to resupply that site — and that time is not the same for an urban building and for a station four hours away by road.

What has to come back from a remote site

On a distant site, monitoring carries an extra load: it is often the only way to know that anything is happening at all. Two points are consistently handled badly.

The first is monitoring the alert chain itself. A site reporting nothing may be a site with no incident, or a site whose link has failed. Without a periodic heartbeat the two look alike, and one of them is an emergency.

The second is powering the monitoring. A measuring device fed from the board it watches stops measuring at the exact moment its measurement becomes interesting.

Taking over existing monitoring

Monitoring that is no longer read is rarely recovered by changing tools. It is recovered by rebuilding the inventory of what is genuinely watched, removing every alert no action follows, and assigning each of those that remain to a person and a deadline.

The exercise needs no licence. It needs one decision per alert, which is why it is so rarely carried out.

Consultations · Pre-qualifications · Partnerships

Let’s discuss your project.

Describe your priority, its operating context and the intended outcome. We will route the enquiry to the right person.