Opsgenie shuts down April 2027 - migrate to Pagerly in one click
PagerlyPagerly
← All postsEngineering

Six Ways to Reach the On-Call Engineer During Downtime

Slack channel, Slack DM, email, mobile push, desktop and a phone call. Six delivery paths for one incident — and the deduplication rules that stop six channels from becoming sixty notifications.

Every on-call tool eventually arrives at the same conclusion: one notification channel is not enough. Slack is where the work happens, but Slack is muted at 3am. Push notifications wake people, but only if the phone is charged and in the room. Email is reliable and completely useless as an interrupt. A phone call breaks through everything, which is exactly why you cannot use it for a P3.

So you add channels. And the moment you add channels, you create a much harder problem than the one you solved. Six delivery paths multiplied by a flapping alert that fires forty times in an hour is not redundancy. It is a denial-of-service attack on your own on-call engineer.

This post covers the six channels Pagerly delivers an incident through, what each one is actually good for, and — the part that matters more — the deduplication that has to sit underneath them.

The six channels

These are not six copies of the same message. Each one exists because the previous one has a specific failure mode.

NOTIFICATION LADDER Six ways to reach the on-call engineer Bar length is what the channel costs the person receiving it. record / interrupt Slack channel Incident opens — broadcast to the war room Email Assign + every status change — the record Slack DM New assignee — the working-hours workhorse Desktop Ownership change — native OS notification Mobile push Ownership change — this is the actual page Phone call Ownership change — defeats DND, opt-in Push, desktop and voice fire only when someone newly becomes responsible — never on a status change.
The two channels above the line carry the incident's state. The four below it interrupt a person, and cost more the further down you go.

1. The Slack incident channel

Broadcast, not interrupt. When an incident opens, Pagerly can spin up a dedicated channel and post the alert into it: severity, source, the link back to the originating monitor, and the current assignee. This is the coordination surface — where the timeline accumulates, where someone posts the graph, where the "I'm looking at it" lands.

Its failure mode is obvious: nobody is watching a channel that was created ninety seconds ago. A channel post is where an incident lives. It is not how anyone finds out about it.

2. The Slack direct message

Targeted, same medium, entirely different behaviour. The DM goes to the person who is actually responsible right now — the on-call resolved from the rotation, plus anyone the escalation policy pulls in as a fixed member of that layer. Slack DMs bypass channel mute settings and produce a badge on a surface most engineers already have open during working hours.

For daytime incidents this is the channel that does most of the real work. It fails the same way Slack always fails: it assumes the person is at a machine with Slack running.

3. Email

The slowest channel, and the only one designed to be read later. Pagerly sends transactional email on assignment and on status change, which makes it the one path that carries the incident's state transitions rather than just its birth.

Email is not a paging mechanism and should never be sold as one. It is a durable record that survives Slack retention, works when someone is off the corporate network, and gives the person picking up the handover something to scroll through. Treat it as an audit trail with a notification badge attached.

4. Mobile push — the actual page

This is the one that wakes people up. A push notification to the Pagerly mobile app is the channel that works when the laptop is shut, the VPN is off and it is 3:40am. It carries enough of the incident to make a triage decision from the lock screen — title, severity, service — so the engineer knows whether this is a "get up" or a "morning" before they open anything.

Its failure modes are the phone's: dead battery, silent mode, Do Not Disturb, a notification permission someone declined during onboarding eight months ago. Every one of those is invisible until the night it matters. This is why the ladder does not end here.

5. Desktop notification

The Pagerly desktop app holds an open connection and fires a native OS notification when an incident is assigned. It fills a gap the other five leave open: the engineer who is at their desk with Slack closed, or in a full-screen editor, or in a meeting with notifications suppressed in the browser but not at the OS level.

It is also the lowest-friction acknowledgement path during working hours. The notification is already on the machine where the fix will be written.

6. The phone call

The escape hatch. An automated voice call reads out the incident and does the one thing no other channel on this list can do: it defeats silent mode and Do Not Disturb on essentially every phone, because that is what phones are for.

This is deliberately gated behind a per-org setting rather than on by default. A phone call is the most expensive interrupt you can spend on a human being. Spend it on a P1 at 4am, not on a disk usage warning.

The ladder, not the list

Six channels firing simultaneously is not six times the reliability. It is one alert and six notifications, which trains people to ignore all six.

The useful framing is a ladder ordered by interrupt cost. Slack channel and DM are cheap and go out first. Email rides along as the record. Push and desktop fire when someone newly becomes responsible — a new incident, a reassignment, a team change — because those are the moments where a human needs to change what they are doing. The phone call sits at the top, off by default, reserved for the severities where waking someone is correct.

What matters is the distinction that ladder encodes: a status update is not a page. When an incident moves from open to acknowledged to mitigated, that is information — it belongs in the channel and in email. It does not belong on a lock screen. Pagerly deliberately does not fire push or desktop on status transitions, only on ownership changes. That single rule removes most of the notification volume people associate with on-call tooling.

Deduplication is the actual product

Everything above is delivery. The reason multi-channel alerting is bearable at all is what happens before delivery: deciding whether this alert is a new incident or the same one shouting again.

Every inbound alert resolves to a dedup key. If an open incident already carries that key, the alert updates it. If not, a new incident opens. One incident, one page, however many times the monitor fires.

DEDUPLICATION Forty-seven firings, one incident, one page Every inbound alert resolves to a dedup key before any channel fires. resolve dedup key aws-CheckoutLatencyHigh KEY MATCHES · 46 times Updates the incident already open. Nobody is paged again. NO MATCH · once Opens a new incident. Pages the on-call. Once.
The dedup key is resolved before any channel fires. Without it the same condition would page six ways, forty-seven times.

The built-in keys, and why they differ per source

The naive implementation keys on the payload's message ID. It is also completely wrong, and it is wrong differently for every monitoring tool. So Pagerly ships a built-in key per integration:

  • Datadog — alert_id, falling back to alert_title
  • AWS CloudWatch — AlarmName
  • Prometheus, Grafana, CubeAPM — the alert fingerprint
  • Sentry — the issue ID, so every occurrence of one error folds into one incident
  • New Relic — the issue id, which stays stable across activate and close
  • Dynatrace — ProblemID; Elastic — alert_uuid; Bugsnag — error.errorId
  • Pingdom, Checkly — check_id; StatusCake — TestID
  • Splunk — search_name, not sid
  • Generic AWS SNS — Subject, not MessageId

Those last two are the instructive ones. Splunk's sid changes every single time the saved search runs, and SNS mints a fresh MessageId per delivery. Key on either and you get perfect, well-formed deduplication that deduplicates nothing — a new incident and a fresh page for every firing. The name of the saved search and the subject of the SNS message are the things that stay constant for the same underlying condition, so those are the keys.

CHOOSING THE KEY The same four firings, keyed two ways Splunk mints a new sid every time the saved search runs. The name does not change. KEYED ON sid New value on every run sid=1725_441021_7 sid=1725_441622_3 sid=1725_442223_9 sid=1725_442824_1 4 incidents · 4 pages KEYED ON search_name Constant for the same condition search_name=checkout_latency search_name=checkout_latency search_name=checkout_latency search_name=checkout_latency 1 incident · 1 page Same trap on generic AWS SNS: MessageId is unique per delivery, so the built-in key uses Subject.
Both sides deduplicate correctly. The left one just has nothing to deduplicate on, because the key it was given changes every run.

When the built-in key is wrong for you

Built-in keys are per-tool defaults, not per-team truth. A team running one Grafana alert across thirty hosts may want one incident per host. A team with three noisy checks on the same service may want them collapsed into one. So the key is configurable per integration:

  • raw — key on a single path, e.g. alerts.0.labels.alertname
  • composite — join several paths, e.g. service plus environment, so the same alert in staging and production stays separate
  • template — interpolate a readable key like {{commonLabels.alertname}} on {{commonLabels.instance}}
  • static — collapse an entire noisy integration into one rolling incident
  • none — opt out, every alert opens its own incident

A prefix can namespace the result so keys from different integrations cannot collide.

The failure mode worth designing for

There is one subtle trap here, and it is where most homegrown dedup quietly breaks. What should happen when a configured rule points at a field that a particular payload does not contain?

The tempting answer is to generate a random key so the alert still gets through. That is the worst possible behaviour: dedup silently stops working, every alert becomes a new incident, and nobody finds out until on-call is buried. The right answer is to fall through — to an explicitly configured fallback value, and failing that, to the built-in key for the integration. An alert that cannot be keyed should still deduplicate on something, and the thing it deduplicates on should be predictable.

What to take from this

If you are evaluating on-call tooling, or building the notification layer yourself, the channel count is the least interesting number on the page. Six channels is easy. The questions that actually determine whether your on-call rotation is survivable are narrower:

  • Which events fire an interrupt, and which only update a record?
  • What is the dedup key for each integration, and is it stable across repeat firings of the same condition?
  • Can you change that key per integration without waiting on a vendor?
  • What happens when the key cannot be resolved?

Get those right and six channels feels like coverage. Get them wrong and one channel already feels like too many.

Pagerly runs on-call schedules, escalation and incident response inside Slack, with mobile, desktop, email and voice as the paths out. If you want to see the dedup configuration against your own alert payloads, start here.