Engineering On-Call: Rotation Design, Compensation

On-call is how software teams keep production running outside business hours. Designed well it's a shared responsibility and a source of ownership.

On-Call Rotation: Building a Sustainable Rotation Engineers Don't Quit Over

On-call is the rotation where engineers are responsible for responding to production alerts outside business hours. It's a foundational practice for any team running software customers depend on — the team that ships the code should be the team accountable when it breaks. Poor on-call design (too-frequent rotations, too-noisy alerts, no compensation) is one of the top reasons senior engineers leave.

Rotation structure

Standard structures: (a) Primary + Secondary — primary responds first, secondary escalates if primary doesn't ack within 15 minutes. (b) Weekly rotation — Mon-Mon shifts, typical for teams of 4-8 engineers giving each engineer 1 week every 4-8 weeks. (c) Follow-the-sun — larger teams split coverage across geographies so nobody is on-call overnight. Follow-the-sun is the gold standard where headcount and locations allow.

Alert quality is the leverage point

The difference between a healthy on-call and a burnout machine is alert quality. Every alert should be (a) actionable — the on-call can do something, (b) customer-impacting or leading indicator of same, (c) rare — pages that fire multiple times per shift retrain the on-call to ignore them. Quarterly alert reviews to delete or tune noisy alerts is the single highest-leverage on-call improvement.

Compensation

Options: (a) On-call stipend per shift ($200-$500/week is common for primary), (b) compensatory time off after heavy shifts, (c) neither — on-call is expected as part of the role, no extra pay. All three are defensible. Not compensating combined with a noisy rotation is the fastest path to attrition. Publishing the on-call comp policy transparently reduces resentment.

Escalation paths

Every on-call needs a clear escalation: primary → secondary → engineering manager → VP Eng → CTO. Escalation is not failure — it's the system working. Culturally reward early escalation on genuine incidents rather than expecting the primary to solo-fight everything. Runbooks for common alert types (linked from the alert itself) dramatically reduce time-to-resolution and escalation frequency.

What breaks on-call

Warning signs: (a) pages during typical sleep hours >2x per shift on average, (b) the same alert fires repeatedly and never gets fixed, (c) engineers negotiating shift swaps constantly, (d) attrition correlated with on-call load, (e) on-call always falls on the same 2-3 people because others are 'too senior' or 'too new.' Any of these signals a rotation that needs redesign before it drives out the team.

Frequently asked questions

Should managers be on-call?
Engineering managers typically aren't in the primary rotation but are the first escalation. VPs and above escalate to for major incidents. Manager-in-rotation works for very small teams; breaks at scale.
How many people do we need for on-call?
Minimum 4 for a sustainable weekly rotation (1-in-4). Below 4, individual burnout is high. Below 2, on-call is unsustainable and the team needs to hire before productizing anything that requires 24/7 uptime.
Do product-focused engineers need on-call?
Yes — the team that ships owns the pager. Splitting product engineers from a dedicated on-call team ('SRE handles it') creates a quality gap; product engineers write code without ops empathy, incidents rise.

Related fundraising guides (40)

Investor directory · Fundraising library · Articles A–Z · Company funding database