Developer Practices & Culture

Why New Backend Devs Should Reject Always-On Work

A backend team can report good morale on Friday and spend Sunday recovering from avoidable pages. That is why I would not make a wellness survey, mindfulness session, or dashboard the default response to developer strain: none can stop a broken alert from waking someone up. For most teams, software wellness should begin with limits on operational work that engineers can enforce.

The popular wellness default leaves the source of strain untouched

Developer-wellness initiatives often start with a survey, a meeting, and an invitation to adopt healthier habits. Those may help people describe a problem, but they are a weak first intervention for a backend team whose work is controlled by deploy frequency, dependencies, and an on-call rota. If the same alert pages someone every night, asking that person to manage stress better shifts responsibility away from the system that caused it.

Software Wellness Philosophy for Healthy Dev Teams names a worthwhile ambition, but I would make the first question less aspirational: what work can reach an engineer outside the plan, and who has authority to stop it? That question is useful because an engineer cannot decline a production incident in the way they can decline an optional meeting.

The distinction matters when you are new to this area. Backend developers tend to look for something they can instrument, so a team-health dashboard can feel like progress. Yet a dashboard showing survey sentiment beside DORA deployment frequency does not tell you whether the person on call handled a customer outage or acknowledged an alert that cleared itself. The figures describe different things, so combining them into a single wellness score conceals the decision the team needs to make.

I would not rank individual engineers by pages handled, pull requests merged, or after-hours activity, because those counts reflect rota assignments and service conditions as well as personal choices. SPACE is a useful reminder that developer productivity has several dimensions; it is not a specification for scoring people. Record burdens at the team and service level first, then use private conversations to learn what the records miss.

A narrower position is more actionable: protect recovery time by reducing interrupt load and making that load visible in planning. Google’s published SRE guidance uses a ceiling of 50% operational work to preserve time for engineering. That figure is a reference from a particular operating model, not a universal quota, but it exposes the trade-off a wellness campaign can avoid naming: interrupts consume the same hours promised to feature work.

A page is a better starting signal than a sentiment score

Start with events the team can verify: pages, incident duration, after-hours acknowledgments, and the work that followed each incident. Prometheus can supply service metrics, Alertmanager can route alerts, PagerDuty can retain escalation history, and OpenTelemetry can connect a failing request to a trace. These tools do not measure well-being; they help identify operational events that may be causing avoidable strain.

For a first pass, export paging events to a CSV file with a fired_at column containing UTC ISO 8601 timestamps. This Python 3 script counts events in a rolling window; run it as python3 pages.py incidents.csv. It deliberately counts pages rather than people, because the first decision is whether the service is generating too many interruptions.

import csv
import sys
from datetime import datetime, timedelta, timezone

cutoff = datetime.now(timezone.utc) - timedelta(days=7)
with open(sys.argv[1], newline="") as handle:
    rows = csv.DictReader(handle)
    pages = sum(
        datetime.fromisoformat(row["fired_at"].replace("Z", "+00:00")) >= cutoff
        for row in rows
    )
print(f"Pages in last 7 days: {pages}")

The 7-day window is a starting value to tune to your rota, not a standard: a longer rotation may need a longer view. Suppose the export shows 18 pages in that window. That would be a count measured from your incident record, not evidence that the team is unwell; you still need to inspect which alerts required action, which repeated, and whether one outage produced many notifications.

Review a sample with the engineer who was on call. If a PostgreSQL connection-pool alert repeatedly fires during a planned Kubernetes rollout and resolves without intervention, the alert threshold or routing may be wrong. If requests are returning HTTP 503 and a person must restore service, the page is doing its job. This classification matters because “reduce pages” is a dangerous target if it rewards silencing real failures.

Service level objectives give that review a boundary. For an availability SLO of 99.9% over 30 days, the calculated error budget is about 43 minutes; that arithmetic describes the permitted unavailability, not a permitted number of sleepless nights. Pair the SLO with the on-call record, because a service can meet its availability target while its alerts still interrupt engineers unnecessarily.

Operational limits beat wellness programs when interrupts are the problem

Consider two choices. A wellness program can win when people need confidential support or a forum to discuss problems that operational logs cannot reveal; it costs facilitation time and cannot repair a noisy alert by itself. An operational workload limit wins when repeated pages and unplanned fixes are the strain; it costs planning flexibility because the team must delay some feature work to remove the source. Most backend teams should try the second choice first when incident records show recurring interruptions, because it acts on the cause they can already see.

A limit must trigger a decision rather than decorate a dashboard. As a trial setting, treat more than 2 non-actionable pages in an on-call shift as a reason to review the alert before accepting new discretionary work on that service. That is a threshold to adjust locally: a quiet service and a service undergoing a risky migration should not inherit the same number without discussion. Do not automatically disable alerts at the threshold, because the next notification may report a real customer-facing failure.

Give the team a short set of responses it can actually choose: change an Alertmanager route, revise a Prometheus alert expression, fix a retry loop, or schedule a reliability task ahead of the next feature. Grafana can show whether an alert fires again after the change, while an incident note can record why the change was made. The point is not to buy more observability software; it is to connect evidence to authority over the work queue.

Software Wellness Philosophy for Healthy Dev Teams is a useful label for the desired outcome, but a label cannot settle a scheduling conflict. If a product deadline and an alert repair compete for the same engineer, someone must explicitly decide which work moves. A team that says it values wellness while leaving every deadline untouched has made the default decision to spend engineers’ recovery time instead.

This is also why I would resist a blanket “no pages after hours” rule. Backend services sometimes fail outside office hours, so removing escalation without changing the failure mode can transfer harm to users and the next shift. Aim instead for pages that demand a human response, with ownership and a runbook clear enough that the response does not depend on guessing.

The policy only works when planning absorbs the repair work

At the next planning meeting, bring the on-call record beside the proposed feature work. If incident follow-ups are merely added to an already full sprint, the policy has no capacity behind it. Reserve a share of planned time for fixes and revise that share against actual interruptions; 25% of team capacity is a trial allocation to negotiate, not a benchmark every team ought to meet.

Make the trade-off visible in the same system used for ordinary delivery. A GitHub Issues ticket can link the incident, the responsible service, the alert rule, and the proposed fix; a GitHub Actions check can then verify that a changed rule parses before it is deployed. That workflow helps because the repair has an owner and acceptance criteria, rather than surviving as a promise made during a retrospective.

Keep the policy small enough to survive a busy week. One service owner should review recurring pages with the current on-call engineer, propose a change, and check the next rotation for recurrence. If the alert still fires, reopen the decision rather than declaring success because a ticket closed. If it stops firing, verify that the underlying failure remains observable through another signal, since silence alone is not reliability.

Managers also need a rule for exceptions. A launch may justify temporarily accepting more operational load, but the exception should have an end date and named follow-up work because temporary burdens otherwise become the new rota. Conversely, a serious incident should override a feature freeze or paging target, because restoring service is more urgent than preserving a metric. These exceptions make the limit credible: engineers can see what it protects and when it yields.

A private check-in still belongs in the process, because logs cannot show fatigue, anxiety, or the cost of being repeatedly interrupted during family time. It should complement the operational record rather than validate it: a quiet PagerDuty history does not prove someone is fine. Ask what changed in the work and whether recovery time was available, then keep personal answers out of service dashboards.

Start by removing one avoidable interruption

Export the last on-call rotation’s pages and review them with the engineer who carried the phone. Pick one recurring notification, decide whether it required action, and either repair its cause or change its routing. Put that work ahead of one planned feature task. That small scheduling decision is a more credible wellness policy than a new score, because it gives the next person on call a chance to sleep.