Expertise

Monitoring and managed services

Monitoring is judged on alerts handled, not on alerts produced.

Parent division
Software and systems

The problem

A console producing three hundred alerts a day stops being read within a month.

Monitoring always decays the same way. A tool is installed, the supplied templates are enabled, and the console produces hundreds of events a day. Nobody can handle them, the team learns to ignore the colour red, and the real failure passes unnoticed among ordinary alerts.

Useful monitoring watches few things, but things whose degradation is certain: disk space, backups, certificates, carrier links, temperature, power, queue depth. Every alert must correspond to a known written action, and to someone who performs it.

Managed services raise a separate question. Delegating operations is legitimate; losing the ability to take them back is not. What separates the two is written in the contract: traceability of the provider’s access, named accounts, maintained documentation, and exit terms.

What separates an alert from an indicator, schematic

01Collection
Systems, network, backups, power, temperature. What degrades with certainty comes first.
02Justified threshold
A threshold with no justification produces an alert nobody will know how to read.
03Alert
Rare enough to be surprising. Repeated without effect, it is not an alert but a badly placed indicator.
04Action and person
Every alert is tied to a written action and to someone who performs it. Otherwise it is tied to nothing.
05Long retention
Long enough to tell a slow drift from a one-off incident.
06Review
The periodic review corrects thresholds and removes from the console what no longer belongs there.
Schematic. The accented path is what makes an alert useful; the return path is what keeps it from decaying.Typical figures. The alert delay is a sizing target; long retention covers a year-on-year comparison.

Scope

What the service covers

Metrics
Collection of indicators from systems, network, backups and power; retention long enough to tell a slow drift from a one-off incident.
Alerts and procedures
Few rules, each tied to a justified threshold, a written action and a person. Defined escalation, and planned silences during interventions.
Day-to-day operations
Scheduled updates, backup management and restoration checks, account management, and tracking of support contracts and end-of-maintenance dates.
Service commitments
Acknowledgement and restoration times by severity, hours covered, on-call arrangements, named contacts, and how an incident is recorded.
Traceability and security
Named accounts for the provider’s staff, logging of remote access, periodic review of rights, and separation of environments.
Reversibility
Operating documentation kept current, inventory, backups accessible to the client, and a handback procedure and schedule at contract end.

Situations

The most frequent situations

Organisation with no operations team
Two or three people hold IT and everything else. Managed services bring the continuity and the on-call cover the headcount cannot.
Estate spread across several sites
Travel is expensive and the real state of remote sites is unknown. Monitoring replaces declared checks with verifiable information.
Critical infrastructure
Power, cooling, links and security must be watched together. Monitoring limited to servers ignores the most frequent causes of outage.
Taking over an existing operation
The previous provider is leaving. Takeover starts with inventory, access, backups and documentation, before any response commitment is made.

Requirements

What to require, of us as of anyone

These requirements hold whichever supplier is appointed. Written into a tender, they filter out the responses that will not hold.

  • A maintained inventory of the systems operated, without which no response commitment means anything.
  • A limited number of alerts, each tied to a written action and to a person.
  • Response commitments by severity, with how the incident is recorded and timestamped.
  • Named accounts for every member of the provider’s staff, and logging of their access.
  • Periodic restoration checks, timed and recorded.
  • Operating documentation maintained by the provider and handed to the client at regular intervals.
  • A reversibility clause stating what is handed back, in what format and within what time.

Pitfalls

Common mistakes, and what they cost

Enabling every supplied alert template
The volume makes the console unreadable within weeks. Monitoring nobody opens is more dangerous than none, because it creates a sense of coverage.
Monitoring the servers and not the power
Most observed outages come from power, cooling or a carrier link. Monitoring that ignores them reports the effect and never the cause.
Letting the provider work under a shared account
A shared account makes every action unattributable. After an incident, reconstruction becomes impossible and responsibility indeterminable.
Signing a contract with no exit clause
Dependency is only noticed at the moment of leaving, when it is too late to negotiate. Handback, its format and its deadline are written on day one.

Questions

Questions asked before consulting

What should be monitored first?

What degrades with certainty and takes effect immediately: disk space, backup success, certificate expiry, carrier link status, temperature in technical rooms, and the state of backed-up power. Those few indicators cover most of the outages actually observed, for a modest setup effort.

How many alerts can a team handle?

Far fewer than assumed. A good rule is that an alert must be rare enough to be surprising. If a team sees the same alert every week without acting, it is not an alert but a badly placed indicator, and it belongs out of the console.

What response commitments should be required?

An acknowledgement time and a restoration time, kept distinct, by severity, with the hours covered and how the incident is timestamped. A single global commitment cannot be verified. It must also be clear what happens when it is missed, otherwise it commits nobody.

How is control retained under managed services?

Through four requirements: named accounts for the provider’s staff, logging of their access, operating documentation handed over periodically, and client access to its own backups. They cost a serious provider nothing and make leaving possible.

Is on-call cover necessary?

It is justified when interrupting a service outside working hours carries a real cost. It presupposes named people, a call procedure that works at night, and remote intervention means that have been exercised. On-call cover without working remote access is only a telephone number.

What should handback at contract end contain?

A current inventory, operating documentation, configurations, backups in a format readable without the provider’s tools, administration accounts and a transfer schedule. Without a schedule, handback stretches out to the date when nobody is available to perform it.

Evidence

Where we have applied it

  • Energy and critical infrastructure2024

    Mining site — server room and backup power

    Construction of a secure technical room and its power chain, on an isolated site with no reliable public grid.

    Weeks of measurement before sizing
    4
    Load tests per year
    12

    Reference KP-2024-008

  • Telecommunications2023

    Regional operator — core network modernisation

    Replacement of the core network and backup power at fourteen points of presence, with no commercial service outage.

    Points of presence migrated
    14
    Longest outage window observed
    4 min

    Reference KP-2023-021

  • Enterprise and sensitive-site security2025

    Industrial group — security supervision centre

    A single control room for six sites whose security systems did not talk to each other.

    Sites connected
    6
    Threat scenarios rehearsed before commissioning
    3

    Reference KP-2025-002

All projects

Consultations · Pre-qualifications · Partnerships

Let us discuss the actual case.

Describe the site, the dominant constraint and the deadline. We will say what requires a preliminary study and what can be committed directly.