DEV Community

Abhishek Singh
Abhishek Singh

Posted on AI-assisted

Cutting Alert Noise in Healthcare Operations: How We Took the Pager From Hundreds a Week to Dozens

Every on-call engineer in healthcare IT knows the 3 a.m. page that turns out to be nothing. A disk at 81% on a server that has sat at 81% for six months. A CPU spike on a batch job that spikes every night by design. A "service down" alert for an endpoint that was down for forty seconds during a scheduled restart.

None of these are incidents. All of them wake someone up. And in a hospital environment, where the same team is also carrying pages that matter — an interface engine backing up lab results, a clinical application server refusing logins at shift change — the noise does something worse than cost sleep. It teaches people to ignore alerts. The team that has been paged 400 times in a week for nothing is slower on the one page that is a patient-safety issue.

This is a write-up of how our operations team took a fleet of enterprise healthcare clients from several hundred alerts a week to a few dozen, without turning off a single monitor that mattered. The tools were AppDynamics and LogicMonitor feeding ServiceNow. The principles apply to any stack — OpenTelemetry, Prometheus, Datadog, SigNoz — because the problem was never the tool. It was the rules.

Step 1: Count before you tune

The first thing we did was not touch a threshold. We exported four weeks of alerts from ServiceNow and grouped them three ways: by rule, by configuration item, and by outcome (did a human take an action, or close it as noise?).

The result was uncomfortable and clarifying:

  • About a dozen rules produced the large majority of all alerts.
  • Several rules had never once resulted in a human action.
  • One "disk space warning" rule alone fired more than a thousand times a month across the estate.
  • The alerts that did lead to real incidents — interface queue depth, database log growth, application login failures — were a small fraction of the total and were buried in the rest.

If you take one thing from this article: do the count. Every ops team believes it knows which alerts are noisy. The export always disagrees with the belief.

Step 2: Every alert must name an action

We adopted one rule for the whole estate: an alert is allowed to exist only if it names the action a human should take when it fires. Not "disk at 80%". Instead: "disk at 80% and growing more than 2% per hour — extend the volume or purge the log directory."

That single test killed a third of the rules outright. Static thresholds on metrics that never change had no action. Informational alerts that duplicated a dashboard had no action. Alerts on components with a healthy redundant pair had no action, because the failover already handled it.

For the alerts that survived, writing the action down did two things. It forced us to set the threshold at the point where the action is actually needed, and it gave us the first line of the runbook for free.

Step 3: Replace static thresholds with rate and duration

Most of the remaining noise came from alerts that fired on a single sample. A server at 95% CPU for five seconds is not a problem; a server at 95% for fifteen minutes usually is.

We changed the rule shape in LogicMonitor from "value above X" to "value above X for N consecutive polls", and where the metric had a natural rate — disk growth, queue depth, error count — we alerted on the rate rather than the level. In AppDynamics the equivalent was moving from static baselines to dynamic baselines with a minimum deviation window, so that the nightly batch job's spike was learned as normal rather than paged as abnormal.

Two rules where this mattered most in a hospital context:

  • Interface engine queue depth. A queue of 200 messages is normal at 7 a.m. and alarming at 2 a.m. We alerted on queue growth over ten minutes, not on the count.
  • Database transaction log growth. A log at 60% is fine; a log that has gone from 40% to 60% in an hour will be full before anyone finishes their coffee. Rate, not level.

This step took the estate down by roughly two-thirds.

Step 4: Maintenance windows are not optional

A surprising share of the remaining alerts were self-inflicted: patching windows, planned restarts, backup jobs. Every one of these was known in advance and every one of them paged someone.

We made maintenance windows a hard part of the change process. A change request in ServiceNow that touched a monitored configuration item automatically created a suppression window in the monitoring tool for the change's duration, via a small integration. No window, no change approval.

This is boring engineering and it removed a further meaningful slice of the weekly volume. It also stopped the worst habit on the team, which was manually silencing a whole host "for the patch" and forgetting to turn it back on.

Step 5: Correlate before you page

By this point the remaining alerts were mostly real, but they still arrived one at a time. A network switch failing produced twelve alerts from twelve downstream servers, each opening its own ServiceNow incident, each paging.

We used the event management layer to correlate on topology: if a parent configuration item is down, child alerts are attached to the parent incident rather than opening their own. Where topology was not modelled, we correlated on time and host group — five alerts from the same cluster within two minutes become one incident with five child events.

The on-call engineer now got one page that said "switch X down, 12 services affected" instead of twelve pages that each said "service Y unreachable". Same information, one interruption, and the root cause is in the title.

Step 6: Review the survivors monthly

Alert rules rot. Applications change, capacity gets added, a threshold that was right in March is wrong in September. We put a thirty-minute monthly review on the calendar with one agenda: the ten noisiest rules from the last month, and the question "did any of these lead to an action?"

Rules that did not are tuned or deleted. Rules that did are left alone. The review has caught rules that would otherwise have crept back to the old volume within a quarter.

What changed for the people

The numbers are the easy part: several hundred alerts a week to a few dozen, and the mean time to acknowledge a real incident dropped from over twenty minutes to a handful. The harder-to-measure change is that the on-call engineer now trusts the pager. When it goes off, something is wrong, and they move.

In a hospital that is the outcome that matters. The interface engine backing up at 2 a.m. gets a human on it in minutes instead of being the fourteenth alert in a queue of noise.

A checklist you can run this week

  1. Export four weeks of alerts. Group by rule and by outcome.
  2. Delete every rule that has never led to a human action.
  3. For every rule that survives, write the action in the alert text.
  4. Convert static thresholds to rate-or-duration conditions.
  5. Tie maintenance windows to your change process, not to memory.
  6. Correlate on topology or host group so one cause is one page.
  7. Put a monthly thirty-minute noise review on the calendar.

None of this needs a new tool. It needs someone to do the count and be willing to delete rules that feel safe and are not.


Abhishek Singh leads incident management and monitoring for enterprise healthcare clients on Azure, with seven years in enterprise IT operations including five at Acquia as a Senior Support Engineer.

Top comments (0)