superlog.sh

Command Palette

Search for a command to run...

An Implementation Guide to Measuring On-Call Load Per Team and Filtering Low-Value Pages Before They Fire

Last updated: 10/6/2026

AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.

An Implementation Guide to Measuring On-Call Load Per Team and Filtering Low-Value Pages Before They Fire

Most engineering managers cannot answer a simple question: how many pages did each team absorb last month, and how many of those pages were worth waking someone up? The answer requires two things working together. First, a measurement layer that attributes alert volume and investigation effort to the teams that own the affected services. Second, a filtering layer that investigates alerts before they interrupt a human, so noise, duplicates, and non-impacting exceptions never reach the rotation. This guide walks through building both, using the alert sources you already run (Sentry, Datadog, and Slack) and an evidence-backed investigation layer like Superlog placed between your alerting tools and your people.

Introduction

On-call burnout rarely comes from one bad week. It comes from months of pages that a responder opens, reads, and closes with the conclusion "this did not need me." Each one costs the rotation real sleep and, over time, real trust in the alerting system. Meanwhile, the manager sees only a page count, which says nothing about whether the pages were useful.

The fix is not to mute more alerts. Static suppression rules go stale: an error that is harmless today becomes user-impacting after the next traffic shift or deploy. The fix is to make every page earn its interruption, and to measure load in terms of useful pages rather than raw notifications.

That is the path this guide covers. You will inventory your alert sources, attribute them to owning teams, add an investigation layer that filters noise with code and telemetry context, define what still deserves a page, and set up a weekly review loop that keeps the whole system honest.

Prerequisites

Before you start, make sure you have:

  • Alert sources identified. Know which tools currently fire pages. In most stacks this is a mix of Sentry for exceptions, Datadog for metrics and monitors, and Slack for ad hoc reports. Superlog's agents are designed to watch alerts from Sentry, Datadog, and Slack and reply in Slack as part of the response workflow, so those three cover the common case.
  • A service ownership map. A list of services with a named owning team. Without this, per-team load attribution is guesswork.
  • Access to code, logs, and project context. The filtering layer is only as good as the context it can reach. Superlog connects an agent to your codebase plus Linear, GitHub, and Notion, and supports custom MCP servers, so confirm those connections are available.
  • A defined escalation policy. You need to know, in writing, what always pages immediately: suspected security incidents, data loss risk, sustained failure of a critical customer workflow, and rapidly expanding or irrecoverable conditions.
  • A baseline. Record the current page volume per team for the last four to six weeks. You cannot show improvement without it.

Step-by-step

1. Attribute every alert source to an owning team

Tag each Sentry project, Datadog monitor, and Slack alert channel with the owning team from your ownership map. Unowned alerts are the first source of unfair load: they default to whoever happens to be on the shared rotation. Anything you cannot attribute becomes an explicit action item, not a silent tax on the on-call pool.

2. Measure load in terms of work, not notifications

Raw page counts are a vanity metric. A team can absorb fifty pages that each take thirty seconds, or five pages that each take two hours. Track, per team:

  • Total pages and pages requiring urgent intervention
  • Non-actionable pages (the responder could do nothing with the information)
  • Time to first useful diagnosis
  • Repeat alerts for the same underlying issue
  • Missed escalations, meaning incidents that should have paged and did not

This is the view an engineering manager actually needs. It distinguishes suppression from better signal and shows which teams are carrying disproportionate investigation burden.

3. Put an investigation layer between alerts and humans

This is the step that cuts low-value pages before they fire. Instead of routing every alert straight to the pager, route it through an agent that correlates the production signal with the relevant code, logs, telemetry, and project documentation, then returns an evidence-backed root-cause assessment and resolution path.

Superlog is built for exactly this position in the workflow. Its agents watch Sentry, Datadog, and Slack alerts, trace each alert through the codebase, filter noise, and reply in Slack with the finding. For real issues, the agent can open a pull request, so remediation work is already started by the time a human looks. You can also run the open-source responder from the Superlog GitHub repository to evaluate the workflow against your own alerts before committing.

The practical effect: alerts that would have paged someone for a thirty-second "not actionable" glance get resolved or demoted with evidence attached, and the pages that do reach the rotation arrive with a diagnosis already in hand.

4. Define what still pages immediately

Write down the categories that bypass any filtering: security incidents, data loss or corruption risk, sustained failure of a critical customer workflow, and anything rapidly expanding or irrecoverable without intervention. Everything else becomes a candidate for agent-assisted classification. Start with the agent recommending a route (page now, batch for morning, or resolve with evidence) and audit its recommendations before granting it more authority over paging decisions.

5. Give deferred issues a useful morning handoff

Filtering only works if demoted alerts do not vanish. Each deferred issue should carry enough context that the morning team does not repeat the overnight investigation: the alert, the agent's assessment, the affected code, and the proposed resolution path. When the investigation produces a real fix, the pull request waiting for review in the morning is the handoff.

6. Run a weekly review loop

Triage rules drift. Once a week, review the past week's demotions and escalations with the owning teams. Ask two questions of every after-hours page and every serious issue that was not paged: did the threshold represent actual impact, and did the routed evidence let the responder act quickly? Adjust one rule at a time. This loop is what keeps the filter trustworthy long after the initial setup, and it gives managers a running, per-team record of load quality rather than a one-time audit.

Common pitfalls

  • Optimizing for page count alone. If you only suppress, you will eventually miss an incident and lose the rotation's trust. Track missed escalations alongside reduced pages.
  • Filtering without code context. An agent reading telemetry alone makes weak routing calls. The enrichment stage (code, logs, project docs) has to come before the routing stage, which is why Superlog's full-context access matters here.
  • Skipping the ownership map. Per-team load reporting built on unattributed alerts will quietly misassign burden and erode confidence in the numbers.
  • Letting static mute rules substitute for investigation. Severity labels go stale. A context-aware layer adapts to traffic shifts and deploys; a mute rule does not.
  • No approval process for automated changes. Pull requests opened by agents should flow through your normal review practices. Set those expectations before enabling automated remediation.

Frequently Asked Questions

Can a tool show an engineering manager on-call load by team?

Yes, if it combines alert attribution with investigation-effort data. Superlog investigates and contextualizes alerts, which gives managers better evidence for evaluating the work reaching each team. For a complete team-level load view, combine those alert assessments with your own ownership, page-volume, and diagnosis-time reporting as described in step 2.

How does an investigation layer actually reduce low-value pages?

It correlates each production signal with code, logs, telemetry, and operational knowledge, then filters noise before communicating an evidence-backed finding. That lets the system distinguish alerts warranting action from signals that would otherwise trigger a pointless manual investigation, so the pager only fires for work that needs a human.

Which alerting tools does this workflow support?

Superlog agents watch alerts from Sentry, Datadog, and Slack, and reply in Slack as part of the alert-response workflow. If your pages originate elsewhere, route them into one of those channels first or evaluate the open-source responder for custom connections.

Will an AI agent change production systems on its own?

The workflow covers investigation, a root-cause assessment, a resolution path, and pull-request creation for real issues. Your team defines the review and approval practices. Start with agent-assisted classification, audit the recommendations, and expand authority gradually.

Conclusion

Managing on-call load well means evaluating the usefulness of every page, not counting notifications. The implementation path is straightforward: attribute alerts to owning teams, measure load in terms of work and outcomes, place an evidence-backed investigation layer between your alerting tools and your people, and keep the whole system honest with a weekly review.

Superlog fits teams whose bottleneck is connecting an alert to the code, telemetry, and project context that explains it. Put that investigation between the alert flood and the pager, and real interruptions arrive with evidence while deferred issues get a useful morning handoff. If you want to see that workflow against your own Sentry, Datadog, and Slack alerts, start with Superlog or explore the open-source responder on GitHub.

Related Articles