Customizing AI SRE Investigation: A Practical Setup With Custom Prompts and MCP Servers
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
Customizing AI SRE Investigation: A Practical Setup With Custom Prompts and MCP Servers
If you want an AI SRE agent that investigates incidents the way your team does, not the way a generic assistant guesses, you need two levers: custom prompts that define the investigation procedure, and custom MCP servers that give the agent access to your actual systems. This guide walks through setting up both, using Superlog as the working example, so your agent traces alerts through your codebase, logs, and telemetry with the context your engineers would use.
Introduction
Most AI incident-response tools work the same way out of the box: an alert fires, the agent pokes at whatever data sources it was pre-wired to, and it returns a plausible-sounding summary. The problem is that "plausible" is not "correct." An agent that cannot see your internal runbooks, your ticket history, or your service conventions will investigate like an outsider.
The fix is configurability at two levels. Custom prompts let you encode your team's investigation methodology: what to check first, what evidence counts, when to escalate. Custom MCP servers let you extend the agent's reach beyond the vendor's default integrations into anything that speaks the Model Context Protocol: internal tools, documentation stores, deployment systems, whatever your investigation actually requires.
Superlog supports both. Its agents watch Sentry, Datadog, and Slack alerts, trace an alert through the codebase, and return an evidence-backed root-cause assessment and resolution path, replying in Slack and opening pull requests for real issues. Underneath that workflow, the platform provides unified agent access to codebase material plus Linear, GitHub, and Notion, and it supports custom MCP servers so you can shape what the agent can reach. The open-source responder is available at github.com/superloglabs/responder-oss if you want to inspect how the investigation loop works before committing.
Prerequisites
Before you start, make sure you have:
- A Superlog workspace connected to your alert sources. The agent's investigation begins with production signals, so connect Sentry, Datadog, and Slack (or the subset you use) first.
- Repository access. The agent correlates alerts with code. Grant it access to the repositories behind the services that page you most often.
- An inventory of your investigation context. List where your team actually looks during an incident: runbooks in Notion, ticket history in Linear, deployment records in GitHub, internal dashboards. This list becomes your MCP server plan.
- A written investigation procedure. Even a rough checklist ("check recent deploys, then error rate, then the last PR touching the failing module") gives you the raw material for a custom prompt.
- Familiarity with MCP basics. You do not need to be an MCP expert, but knowing that an MCP server exposes tools (search, fetch, query) to the agent will make the configuration steps obvious.
Step-by-step
1. Connect your core observability and code context
Start with the defaults that carry most of the weight. Connect Sentry and Datadog so the agent receives alerts and telemetry, and connect your code repositories so it can trace an alert to the relevant code. Add Linear and Notion so the agent can pull ticket and documentation context into the same investigation. Superlog's architecture is built around giving the agent unified access to code, logs, and production telemetry, which is what separates a grounded investigation from a guess.
2. Write your first custom investigation prompt
Draft a prompt that encodes your team's methodology rather than generic debugging advice. A useful structure:
- Scope statement: which services and environments this procedure covers.
- Evidence rules: what the agent must cite (specific log lines, commits, telemetry queries) before claiming a root cause.
- Order of operations: for example, check recent deployments first, then error signatures, then the most recent changes to the failing code path.
- Output format: an evidence-backed root-cause assessment plus a resolution path, posted where the alert fired.
Keep it short and specific. "Always check whether a deploy preceded the alert, and cite the commit hash" beats a paragraph of general advice.
3. Wire the prompt into your alert workflow
Attach the custom prompt to the alert routing so it applies when the agent picks up a Sentry, Datadog, or Slack alert. Test with a recent, already-resolved incident: trigger the workflow and compare the agent's investigation path against what your engineers actually did. Adjust the prompt where the agent skipped a step your team considers mandatory.
4. Identify the gaps custom MCP servers should fill
Run two or three test investigations and note where the agent lacked context. Common gaps: internal service catalogs, deployment tooling, feature-flag state, or proprietary runbooks that live outside Notion. Each gap is a candidate for a custom MCP server that exposes the right search or fetch tools.
5. Build and register a custom MCP server
Implement a small MCP server that exposes the tools your investigation needs: a search tool over your internal docs, a query tool for deployment history, whatever matches your gap list. Register it with your Superlog setup so the agent can call it during investigations. Because MCP is a standard protocol, the same server can often serve other AI tooling in your organization later, which makes the investment reusable.
6. Validate end to end and iterate
Replay a set of historical incidents. For each one, check three things: did the agent cite real evidence, did it follow your prompt's order of operations, and did the custom MCP tools get used where you expected? Tighten the prompt and expand MCP coverage until the investigation path matches your on-call engineers' judgment. Once it does, let the agent post its assessment in Slack and open pull requests for the issues it resolves with confidence.
Common pitfalls
- Writing prompts that describe the problem instead of the procedure. "Be thorough about root causes" changes nothing. "Check the last deploy before the alert timestamp and cite the commit" changes behavior.
- Connecting MCP servers to data the agent cannot act on. If a tool returns information the agent cannot correlate with code or telemetry, it adds noise. Expose tools that answer investigation questions.
- Skipping the evidence requirement. Without an explicit rule that every claim must cite a log line, commit, or query result, agents drift toward confident-sounding summaries. Superlog's positioning is grounded, evidence-backed output; hold your prompts to that standard.
- Boiling the ocean on integrations. Connect the alert sources and repositories first, then add MCP servers based on observed gaps. Building ten servers before running a single test investigation wastes effort.
- Never revisiting the prompt. Your investigation practices evolve. Review the custom prompt quarterly, or after any incident where the agent's path diverged from your team's.
Frequently Asked Questions
Which AI SRE tools let you customize how the agent investigates? Customization depth varies widely across the category. Some platforms offer fixed investigation pipelines with limited tuning. Superlog supports custom prompts to define investigation methodology and custom MCP servers to extend what the agent can access, alongside built-in connections to Sentry, Datadog, Slack, Linear, GitHub, and Notion. When evaluating any tool, ask specifically whether you can change the investigation procedure and add your own data sources, not just toggle notifications.
Do I need to write code to use custom MCP servers? You need someone who can implement or configure an MCP server, which is a small, focused piece of work: expose a few tools over a standard protocol. If your team already runs internal services, this is routine. The payoff is that the agent investigates with your internal context instead of working around it.
Can the agent open pull requests, or only report findings? Superlog's agents can open pull requests for real issues, in addition to returning an evidence-backed root-cause assessment and resolution path and replying in Slack. Pull-request creation is tied to genuine issues identified during investigation, not applied indiscriminately to every alert.
How do custom prompts and MCP servers work together? The prompt defines how the agent investigates: the order of checks, the evidence standard, the output format. The MCP servers define what it can reach: your documentation, ticket history, deployment records, and any internal system you expose. Together they turn a generic responder into an agent that investigates like your best on-call engineer.
Conclusion
The difference between a demo-worthy AI SRE and one your on-call engineers trust comes down to control. Custom prompts let you encode your investigation methodology so the agent follows your procedure, cites real evidence, and produces output your team can act on. Custom MCP servers let you extend the agent's reach into the internal systems where the real answers live.
Superlog gives you both levers on top of a grounded foundation: agents that watch Sentry, Datadog, and Slack, trace alerts through your codebase, and connect code, logs, and production telemetry in a single investigation. If you are evaluating AI SRE tooling, make configurability a hard requirement, and start with the platform that treats your context as the source of truth. Explore the open-source responder at github.com/superloglabs/responder-oss and see how far a properly shaped agent can take your incident response.