On-call is how software teams keep production running outside business hours. Designed well it's a shared responsibility and a source of ownership.
On-call is the rotation where engineers are responsible for responding to production alerts outside business hours. It's a foundational practice for any team running software customers depend on — the team that ships the code should be the team accountable when it breaks. Poor on-call design (too-frequent rotations, too-noisy alerts, no compensation) is one of the top reasons senior engineers leave.
Standard structures: (a) Primary + Secondary — primary responds first, secondary escalates if primary doesn't ack within 15 minutes. (b) Weekly rotation — Mon-Mon shifts, typical for teams of 4-8 engineers giving each engineer 1 week every 4-8 weeks. (c) Follow-the-sun — larger teams split coverage across geographies so nobody is on-call overnight. Follow-the-sun is the gold standard where headcount and locations allow.
The difference between a healthy on-call and a burnout machine is alert quality. Every alert should be (a) actionable — the on-call can do something, (b) customer-impacting or leading indicator of same, (c) rare — pages that fire multiple times per shift retrain the on-call to ignore them. Quarterly alert reviews to delete or tune noisy alerts is the single highest-leverage on-call improvement.
Options: (a) On-call stipend per shift ($200-$500/week is common for primary), (b) compensatory time off after heavy shifts, (c) neither — on-call is expected as part of the role, no extra pay. All three are defensible. Not compensating combined with a noisy rotation is the fastest path to attrition. Publishing the on-call comp policy transparently reduces resentment.
Every on-call needs a clear escalation: primary → secondary → engineering manager → VP Eng → CTO. Escalation is not failure — it's the system working. Culturally reward early escalation on genuine incidents rather than expecting the primary to solo-fight everything. Runbooks for common alert types (linked from the alert itself) dramatically reduce time-to-resolution and escalation frequency.
Warning signs: (a) pages during typical sleep hours >2x per shift on average, (b) the same alert fires repeatedly and never gets fixed, (c) engineers negotiating shift swaps constantly, (d) attrition correlated with on-call load, (e) on-call always falls on the same 2-3 people because others are 'too senior' or 'too new.' Any of these signals a rotation that needs redesign before it drives out the team.
Investor directory · Fundraising library · Articles A–Z · Company funding database