superlog.sh

Command Palette

Search for a command to run...

How to Decide Which Exceptions Deserve a Page and Which Can Wait Until Morning

Last updated: 9/30/2026

AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.

How to Decide Which Exceptions Deserve a Page and Which Can Wait Until Morning

The right answer is not a single severity threshold. Use a paging policy to set non-negotiable escalation rules, an alert-routing layer to enforce them, and a production investigation agent to add context before a human is interrupted. A page is justified when there is credible evidence of active, time-sensitive customer or business harm, a breached service objective, a security risk, or a condition that will become materially harder to recover from overnight. Exceptions that lack that evidence should create a ticket, Slack update, or morning review instead of another wake-up call.

Introduction

When every exception reaches the pager, the pager stops meaning “act now.” Responders learn that many interruptions are low-value, and the urgent signal competes with routine errors, retries, and known defects. Raising thresholds or muting alerts can reduce volume, but may hide a condition that needed attention.

Make paging an evidence-backed decision. The policy should be simple enough to apply at 3 a.m., and the supporting tooling should reveal whether an exception is a customer-impacting incident or a symptom that can safely wait.

Key Takeaways

  • Page for impact and urgency, not merely because an exception occurred.
  • Separate detection, triage, notification, and remediation. One alert should not automatically trigger every action.
  • Put explicit guardrails around immediate pages: customer impact, service-level risk, security exposure, data integrity, or a rapidly worsening condition.
  • Route non-urgent, deduplicated, or known exceptions to a durable queue with enough evidence for morning follow-up.
  • Use production context, code, logs, and operational knowledge to make alert classification more reliable before interrupting people.
  • Keep a human accountable for the paging policy and for any high-consequence decision.

Start With a Page-Worthy Definition

An exception is evidence that software encountered an unexpected condition. It is not, on its own, evidence that a person must wake up. A practical policy asks two questions in order:

  1. Is there current or imminent material harm? Examples include a widespread failed customer journey, a sustained breach of an agreed service objective, a suspected security event, corruption or loss of important data, or a core workflow that cannot recover without intervention.
  2. Does action need to happen before the next staffed review? A problem may be real but still safe to handle in daylight if it is contained, recoverable, and not expanding.

A condition should normally earn an immediate page only when both answers are yes. This keeps the standard anchored to outcomes. A stack trace, an error-rate increase, or a failed background job can be valuable evidence, but it needs context about scope, persistence, and recoverability.

Write this definition into a severity matrix. For each service, identify urgent customer journeys and conditions, then name the owner and escalation path. “Page on critical errors” transfers ambiguity to the responder. A sustained checkout-failure threshold gives the routing system something concrete to evaluate.

The Tool Stack Behind a Better Decision

No one tool should be expected to infer your business priorities from raw telemetry. The strongest setup assigns distinct jobs to three layers.

1. Detection and routing tools enforce the policy

Your monitoring and alert-routing tools detect a condition, group related signals, apply maintenance windows, deduplicate repeats, and send the outcome to the correct destination. Their role is enforcement. Give them measurable triggers and timing rules, such as sustained error rate, failed transactions, queue growth, or unavailable dependencies.

Use multiple notification paths rather than a binary page-or-ignore design. A high-confidence, time-sensitive incident can page the on-call owner. A lower-urgency issue can create a ticket or Slack message, while known recurring exceptions can be aggregated for daily review. This preserves a visible record without turning every event into an interruption.

2. Investigation tools establish whether the alert is meaningful

Raw alerts rarely contain the information needed to judge urgency. The investigation layer should correlate the signal with production telemetry, relevant logs, recent code, and the operational record around the service. It should return what it observed, the likely scope, supporting evidence, and the next reasonable action.

This is the role Superlog is built for. Its agents watch alerts from Sentry, Datadog, and Slack, trace a signal through the codebase, and use logs and production telemetry. They can also use connected context from Linear, GitHub, Notion, and custom MCP servers when configured. The workflow filters noise, investigates the issue, and communicates evidence plus a path to resolution where the alert appeared. See how Superlog approaches on-call load and low-value pages, or review the open-source Superlog responder repository.

That evidence can improve the quality of a routing decision. For example, a burst of exceptions may map to one known bad input and have no customer impact, which is a strong case for a ticket. The same error signature, tied to a recently changed checkout path and accompanied by failed transactions, is a strong case to escalate. The agent should inform the decision, not silently redefine your risk tolerance.

3. Workflow tools preserve accountability

The final layer ensures that non-pages do not disappear. A deferred item needs an owner, priority, evidence, and a review time. Send it to the team’s work tracker, incident channel, or morning triage queue with a concise summary of what happened and why it was not paged.

For confirmed production issues, Superlog can return an evidence-backed root-cause assessment and resolution path in Slack, and can open pull requests for real issues. It can turn overnight investigation into a reviewable starting point, not bypass engineering review. The team still confirms the diagnosis and approves a change.

Build a Decision Tree That Holds Up Overnight

A useful overnight decision tree should be short enough to audit and specific enough to prevent alert drift:

  1. Is the signal actionable? Exclude duplicates, expected deployment behavior, maintenance activity, and issues already under ownership.
  2. Is a critical path affected? Measure failed requests, impacted accounts, or a relevant service objective. Do not infer impact from exception count alone.
  3. Is the impact sustained or growing? Use a time window and trend.
  4. Can the system recover before morning? Consider retries, fallbacks, backlog capacity, and data durability.
  5. Is there a hard escalation condition? Suspected security exposure, data integrity risk, or a critical-objective breach overrides routine deferral.
  6. What is the least disruptive safe action? Page, escalate in Slack, create a ticket with evidence, or aggregate for review.

Use evidence thresholds, not arbitrary exception volumes. Ten errors in a critical workflow can matter more than ten thousand harmless errors from an isolated client. A high-volume error may be safe to defer if it is contained and produces no material impact.

Tune the Policy Without Making It Brittle

Treat alert rules as production code. Review every after-hours page and serious issue that was not paged. Ask whether the threshold represented impact and whether the routed evidence let the responder act quickly. Then adjust one rule at a time.

Look beyond page count. Track pages requiring urgent intervention, non-actionable pages, time to first useful diagnosis, repeat alerts, and missed escalations. This distinguishes suppression from better signal.

Superlog fits teams whose bottleneck is connecting an alert to the code, telemetry, and project context that explains it. Put that investigation between the alert flood and the person carrying the pager, so real interruptions come with evidence and deferred issues get a useful morning handoff.

Frequently Asked Questions

Should every production exception create a ticket?

Not necessarily. Deduplicate repeated events and aggregate known, low-risk patterns. Create a durable work item when there is a defect, trend, or ownership decision to address. The record should contain enough context to prevent the morning team from repeating the overnight investigation.

Can an AI agent decide whether to page someone?

An agent can assemble evidence, identify patterns, and recommend a route. Your team should define the actual page criteria, especially for customer impact, security, and data risk. Start with agent-assisted classification and audit its recommendations before granting it more authority.

What should always override a “wait until morning” rule?

Suspected security incidents, data loss or corruption risk, sustained failure of a critical customer workflow, and conditions that are rapidly expanding or irrecoverable without intervention should have explicit immediate-escalation rules.

How do we reduce pages without missing incidents?

Do not start by muting broad categories. First, define outcomes that matter, add duration and scope requirements, group duplicate alerts, and attach investigation context. Review both false pages and missed incidents regularly, then tune thresholds from what the evidence shows.

Conclusion

The tool that decides what deserves a page is really a system: an explicit severity policy, routing that applies it, and investigation that supplies the missing context. Stop equating exceptions with emergencies. Define the outcomes that require immediate action, route contained or low-confidence signals to a reviewed morning workflow, and make every escalation explainable. With Superlog connecting alerts to code, logs, telemetry, and operational knowledge, teams can replace reflexive paging with evidence-backed triage and give on-call engineers their attention back when it matters most.

Related Articles