24/7 On-Call Coverage With a Small Team
How to run 24/7 on-call coverage with four or five engineers: the real math, four rotation patterns, and when you are too small to try.

Running 24/7 on-call coverage with a small team is first of all an arithmetic problem, and most teams avoid doing the arithmetic because the answer is uncomfortable. There are 168 hours in a week. Your engineers work roughly 40 of them. Somebody has to be reachable for the other 128, and if you have four engineers, that somebody is each of them, one week in four, thirteen weeks a year.
That can be sustainable or it can quietly destroy a team, and the difference has very little to do with the rotation design. It has to do with how often the phone actually rings at night. This guide covers the honest math, how to decide what genuinely needs overnight coverage, four rotation patterns that work at small headcounts, and what to do when the truthful answer is that you are too small for 24/7 at all.
Start With The Honest Math
What One Week In Four Actually Costs
With four engineers on a weekly rotation, each person is on call for about a quarter of the year. Stated that way it sounds heavy. Stated as "one week in four" it sounds routine. Both describe the same thing, and the second framing is how teams talk themselves into arrangements they later resent.
The number that matters is not how often someone is on call. It is how often being on call costs them something. A week carrying a phone that never rings is a mild background tax. A week with four night pages is a week of degraded sleep, which has measurable effects on judgement and mood that persist for days afterwards. The same rotation produces both outcomes depending entirely on alert volume.
Measure The Interrupt Rate Before Redesigning The Rotation
Before changing anything about who is on call when, measure two things over a representative month:
- Pages per shift, split by whether they arrived during working hours or outside them
- Actionability rate, meaning the proportion of pages where the responder did something other than acknowledge and go back to sleep
A widely used rule of thumb is that more than two pages per on call shift is unsustainable, and anything under about a fifty percent actionability rate means you have an alerting problem rather than a staffing problem. If half your night pages did not require action, no rotation design will fix the fatigue, and hiring will not either. You would just be distributing the same noise across more people.
Separate Being On Call From Being At Work
Write down explicitly what on call means at your company, because ambiguity here is corrosive. Is the expectation that someone is reachable within fifteen minutes, or that they are actively watching dashboards? Can they go to dinner, a cinema, their child's school event? Can they drink a beer?
The reasonable answer for most teams is that on call means reachable and able to be at a laptop within a defined window, not that your evening is forfeited. If your actual expectation is closer to the second, you are asking for a lot more than the rotation implies and should be compensating accordingly.
Decide What Genuinely Needs A Night Page
This is the highest leverage section of this guide. Small teams cannot afford to page for everything, and the good news is that they do not need to.
The Can It Wait Until Morning Test
For every alert that currently pages overnight, ask one question: if this fires at 03:00 and nobody looks until 08:00, what is different at 08:00? Three honest answers are possible.
Something irreversible happens, such as data loss, money moving incorrectly, a security breach in progress, or a regulatory deadline missed. That pages. Something degrades further but recovers, such as a queue backing up that will drain once capacity returns. That usually waits, unless the degradation compounds. Nothing meaningfully changes, which describes the majority of alerts that currently page overnight at most companies. That should never have paged.
Tier Explicitly, Then Enforce It
A workable three tier scheme for a small team:
- Page immediately, any hour: customer facing total outage, data loss or corruption in progress, payment processing down, security incident, anything with a contractual response deadline.
- Page during extended hours only, roughly 07:00 to 23:00: significant degradation affecting a subset of customers, a failed deploy that is not yet customer visible, capacity approaching a limit.
- Ticket for the next working day: everything else, including most disk space warnings, most retryable job failures, and every alert that has ever been resolved by waiting.
The enforcement matters more than the scheme. Review every overnight page in your weekly team meeting and ask whether it belonged in tier one. Demote aggressively. A single engineer with authority to demote a noisy alert on the spot is worth more than a quarterly alert review nobody schedules.
Four Patterns That Work At Small Headcount
Pattern 1: Single Primary, Weekly (Three To Five People)
One person holds the pager for a week. No secondary. The simplest possible arrangement, and viable only if your alert hygiene is genuinely good and the tier one list is genuinely short.
The advantage is clarity: there is never any question who is responsible. The weakness is obvious, which is that there is no backup if the primary is unreachable, asleep through a page, or in a tunnel. Mitigate with an escalation path to a team lead or a manager after a defined non acknowledgement window. That person is not on call in any meaningful sense, they are a safety net that gets used a handful of times a year.
Pattern 2: Primary Plus Secondary (Five Or More)
A primary takes all pages, and a secondary is called if the primary does not acknowledge within a set period, commonly ten to fifteen minutes. The secondary shift runs offset from the primary, so nobody holds both roles simultaneously and the secondary is usually someone who held primary recently and has current context.
This needs at least five people to avoid everyone being on one rotation or the other almost permanently. With five engineers you get roughly one week primary and one week secondary out of every five, which is tolerable if the secondary is genuinely rarely called. If your secondary is being called regularly, that is a signal your primary is overloaded, not that the pattern is working.
Pattern 3: Weekday And Weekend Split
Two independent rotations: one covering Monday to Friday, another covering Friday evening to Monday morning. Weekends are the part people mind most, and splitting them separately means the weekend burden is distributed fairly rather than landing on whoever happened to draw that week.
Offset the two rotations so the same person rarely holds a weekday shift adjacent to their weekend shift. Otherwise you have accidentally created an eleven day stretch, which is worse than either rotation alone. This pattern works well from about five people upward, and it is frequently the single change that most improves how a rotation feels without changing how much coverage it provides.
Pattern 4: Follow The Sun (Two Genuine Regions)
If you have engineers in sufficiently separated time zones, follow the sun eliminates night pages entirely. Each region covers its own daylight hours and hands off at the boundary with a short overlap.
This is the best outcome available and it has a hard prerequisite: real depth in both regions. Two engineers in London and one in San Francisco is not follow the sun. It is one person in San Francisco with no backup, and it will fail the first time they take a holiday. Be honest about whether you have the headcount, because the label is appealing enough that teams adopt it before they qualify.
Making Nights Survivable
Alert Hygiene Is The Entire Game
Everything else in this guide is secondary to this. A four person rotation with two actionable pages a month is fine indefinitely. A twelve person rotation with nightly noise will still burn people out. If you do one thing, spend a month driving down the overnight page count rather than redesigning the schedule.
A Runbook For Every Paging Alert, No Exceptions
Make this a hard rule: if an alert can page someone at night, it has a runbook. If it does not have a runbook, it does not page. This single policy does two useful things at once. It makes nights survivable for whoever gets woken, because they are not debugging from first principles at three in the morning. And it creates natural resistance to adding paging alerts casually, because adding one now costs you a document.
The runbook does not need to be long. What the alert means, what to check first, the three most common causes, what a safe remediation looks like, and when to escalate. One page.
Compensate, And Give The Time Back
On call is work. Being reachable constrains your evening, your weekend and your plans whether or not the phone rings, and treating that as a free extension of the role is how teams lose their most experienced people. Pay a stipend, or give time off in lieu, or both. If you want to see what your current rotation is actually costing in those terms, our on-call pay calculator works it out per year.
Separately from compensation, institute a recovery norm: anyone paged after midnight starts late the next day, or takes the day, depending on severity. Make it automatic rather than something people have to ask for, because the people who most need it are the least likely to ask.
Rotate The Improvement Work, Not Just The Pager
A pattern worth adopting: whoever comes off an on call shift spends part of the following week fixing whatever annoyed them most during it. This closes the loop between experiencing pain and having authority to remove it, and it steadily reduces alert volume instead of letting it accumulate. Without it, everyone suffers the same noisy alert for a year and nobody ever owns removing it.
Covering Absence Without A Single Point Of Failure
Small teams have a specific structural risk: one person who knows the payment system, or the deployment pipeline, or the legacy service nobody else has touched. When that person is on holiday, coverage exists on paper and not in practice.
Address it in two ways. Make overrides trivially easy, so swapping a shift is a thirty second action rather than a negotiation, and people actually use the system instead of arranging informal cover that the alerting does not know about. And deliberately spread knowledge: pair on incidents, rotate who writes the postmortem, and have the person with exclusive knowledge write the runbook for their system before their next holiday rather than after.
When You Are Too Small For 24/7
Some teams should not be running 24/7 coverage, and saying so plainly is more useful than helping them build a rotation that will fail. With two or three engineers, a genuine always on rotation means each person is on call a third to a half of the year. That is not sustainable for longer than a few months, and pretending otherwise usually ends with resignations.
The honest alternatives:
- Define a real support window and be explicit with customers about it. Extended hours coverage that you actually honour is better than nominal 24/7 that nobody answers.
- Page for sev one only, genuinely, with a very short list of what qualifies. Everything else waits for morning by design rather than by accident.
- Buy down the risk technically. Auto remediation, redundancy, graceful degradation and automatic rollback all reduce the number of situations that require a human overnight. Engineering effort here is often cheaper than the human cost of the rotation.
- Be clear in hiring. Tell candidates what the rotation is. People who join understanding it cope far better than people who discover it in month three.
Metrics Worth Watching
- Pages per shift, split by in hours and out of hours
- Actionability rate, the proportion of pages needing real action
- Night pages per person per month, watching for uneven distribution
- Time to acknowledge, which climbs when people are fatigued or alerts are not trusted
- Repeat alerts, the same alert firing across multiple shifts without anyone fixing the cause
- Shift swap frequency, a decent proxy for how much people dread specific slots
Running It Without A Second Console
For a small team, the operational overhead of the on call system itself matters. Nobody has a dedicated ops person to maintain schedules, and any process that requires opening a separate tool to swap a shift will be routed around with informal arrangements the alerting does not know about.
Keeping it in Slack is largely the point of Pagerly's rotations: the schedule, swaps and overrides are handled with a few commands where the team already is, and Slack user group sync keeps a group like @oncall pointing at whoever currently holds the shift, so the rest of the company can reach the right person without asking. Paging escalates through Slack, email, SMS, phone call and the mobile apps, which covers the single primary pattern's main weakness: if the primary sleeps through it, it moves on to someone else rather than sitting unacknowledged.
The Takeaway
Small team 24/7 coverage is viable, and the rotation pattern is the least important part of making it work. Four engineers with quiet nights do fine on the simplest possible schedule. Twelve engineers with noisy nights struggle regardless of how cleverly the shifts are arranged.
So do the measurement before the redesign. Count the pages, count how many needed action, and cut the overnight alert list to the things that genuinely cannot wait until morning. Then pick the simplest pattern your headcount supports, put a runbook behind every page, pay people for the constraint, and give the time back when a night gets broken. If the arithmetic still does not work, the answer is a narrower support commitment rather than a rotation that quietly consumes your team.
Want to run a rotation without adding another console? Pagerly handles schedules, swaps, overrides and escalation inside Slack, so a small team can cover 24/7 without a dedicated tool to maintain. Get started free.
