Stack · Operations

On-call desk

Answers the page: what fired, what changed, what the runbook says, and what to tell the channel.

Built for: The engineer holding the pager who has to be useful in the first ten minutes, at two in the morning.

Install all 10 parts

The button opens the checkout, where 6 servers and 4 skills are listed one by one with what each does to the bill — free, already yours, monthly or a one-off licence. Nothing is charged until you confirm it there, in Stripe’s own card frame on that page rather than a redirect, and each paid member keeps its own budget cap.

$25/mo
6 servers at their publishers’ prices, with one-off licences spread over 12 months
$0
the free-tier version keeps 5 servers and 1 skill — 0 of the 2 tools
$0
4 skills with no monthly cost, plus $127 paid once
15.2k
tokens of context the skills add to every session
01 · Outcome

What your agent can do with this

The reason to buy a stack rather than five listings: each line below needs more than one member connected at the same time.

  1. 01

    Read the alert and the metric behind it in Datadog or Prometheus, and say what the series looked like before it fired.

  2. 02

    Inspect pods, deployments and services on the cluster and name the one that restarted.

  3. 03

    Search LogScale for the window around the alert instead of scrolling a console by hand.

  4. 04

    Post the state of play into the Discord channel that is already asking, in the format your team writes updates in.

02 · Assembly

The assembly, part by part

What each part contributes, and why it was picked over the obvious alternative. Prices and permissions are read from the listings, so nothing here can disagree with the catalogue.

6 servers4 skills2 tools
DA01

The Datadog API where the alert and its dashboards live, so the first question — what is on fire — is answered from the source.

read/write split not recordedruns on your machine A v1.8.0 credential not declared · runs locally
Free
$0 of the monthly total
PR02

Queries the series directly, which is what you need when the dashboard shows a spike and you want the raw numbers behind it.

read/write split not recordedruns on your machine A v1.1.3 credential not declared · runs locally
Free
$0 of the monthly total
KU03

Inspects and manages pods, deployments and services — the layer where most pages turn out to have started.

read/write split not recordedruns on your machine A v4.1.4 credential not declared · runs locally
Free
$0 of the monthly total
SE04

Ties the page to the release: errors, events and the deploy that preceded them, which is the fastest correlation available in an incident.

read/write split not recordedruns on your machine A v1.0.0 credential not declared · runs locally
Free
$0 of the monthly total
DI05

The channel the incident is already being discussed in, so the update lands where people are looking rather than in a document nobody opens.

read/write split not recordedruns on your machine A v2.1.1 credential not declared · runs locally
Free
$0 of the monthly total
LO06

Queries CrowdStrike LogScale for the window around the alert, which beats scrolling a log console at two in the morning.

2 readread-onlyruns on your machine A v0.1.8 credential not declared · runs locally
from $25/mo
$25 of the monthly total
and the instructions that drive them A skill is a prompt file, not a server: it adds context and a procedure, never a tool or a permission of its own.
EN07
Engineering Runbook Agent skill Workflow by nexu-io

Keeps the runbook in one shape — service, alert table, dashboards, copy-pasteable procedures — so the page is answered from a document rather than from memory.

3.9k tokens of context never asks for a write 2 files · Apache-2.0 needs no server of its own
$89
$7.42 of the monthly total
IN08
Internal Comms Agent skill Output format by anthropics

Writes the incident update in the format your company already uses, which is what stops the status post from becoming its own argument.

5.6k tokens of context never asks for a write 6 files · Complete terms in LICENSE.txt needs no server of its own
Free
no monthly cost
PY09
Python Observability Agent skill Output format by wshobson

The other half of on call: after the incident, the structured logging and metrics that would have made this page shorter.

3k tokens of context never asks for a write 2 files · MIT needs no server of its own
$19
$1.58 of the monthly total
DE10
Debugging and Error Recovery Agent skill Guardrail by addyosmani

Keeps the first ten minutes systematic — evidence, hypothesis, test — instead of restarting things until the graph looks better.

2.7k tokens of context never asks for a write 1 files · MIT needs no server of its own
$19
$1.58 of the monthly total
03 · Cost

What it costs, and on what assumption

Every member is a subscription or a licence bought once, so the monthly figure is a price rather than an estimate: what moves it is adding or dropping a member, not how hard the stack is worked. The one assumption is that a one-off licence is spread over a year so it can sit in the same column as a subscription.

PartWhat you are paying forMonthly, as quoted
Datadog Free
Prometheus Free
Kubernetes Free
Sentry Free
Discord Free
Logscale from $25/mo $25/mo
Skills
Engineering Runbook $89 · $7.42/mo over 12 months $7.42/mo
Internal Comms Free · context cost only
Python Observability $19 · $1.58/mo over 12 months $1.58/mo
Debugging and Error Recovery $19 · $1.58/mo over 12 months $1.58/mo
Everything above $25 of servers plus $11 of skills, the same in a quiet month and a busy one $36/mo

Subscriptions at their monthly plan price; one-off licences spread over 12 months. One-off purchases in this stack total $127 — Engineering Runbook $89, Python Observability $19, Debugging and Error Recovery $19 — paid once and spread here so they sit in the same column as a subscription. Everything arrives on one mcprush invoice, taken by Stripe from the card on your account, not one per publisher — mcprush.com is the merchant of record and each publisher is paid out of it.

The free-tier version$0/mo

Install only these and the bill is nothing: 5 servers and 1 skill, 0 of the 2 tools.

Left out, and what goes with it:

  • Logscale · from $25/moQueries CrowdStrike LogScale for the window around the alert, which beats scrolling a log console at two in the morning.
  • Engineering Runbook · $89Keeps the runbook in one shape — service, alert table, dashboards, copy-pasteable procedures — so the page is answered from a document rather than from memory.
  • Python Observability · $19The other half of on call: after the incident, the structured logging and metrics that would have made this page shorter.
  • Debugging and Error Recovery · $19Keeps the first ten minutes systematic — evidence, hypothesis, test — instead of restarting things until the graph looks better.
What moves the bill
Flat every month$25 · 1 subscription
Carries a call allowance1 member
Paid once$127
Traffic assumed5k calls / month

Logscale at from $25/mo. Each of those plans states the calls it includes in a month, and running past one never arrives as a larger invoice: the gateway refuses the call over the allowance and returns an MCP error naming the plan. The figure above is what the stack costs in a busy month as well as a quiet one — what a heavy month changes is which plan you need, not what this one bills.

Budget caps are set per install and enforced at the gateway, so a retry loop is refused at the cap rather than left to run through an allowance overnight.

04 · Setup

Setting it up, in order

One step per part, in the order they are useful: connect what the work reads before what it writes, and install the skills that decide how the work is done last. Each step is a command you can read before you run it.

1
Two read keys, not one
Datadog and Prometheus each need their own credential. Give both read-only scopes; nothing in an incident's first ten minutes needs a write key.
2
Scope the cluster access
A kubeconfig context limited to the namespaces you are on call for. Cluster-admin in a pager stack is how a bad night becomes a worse one.
3
Write the runbook first
The runbook skill formats and keeps one — it does not invent your alerts. Fill in the service, the alert table and the escalation path before the first page.
4
Pick the channel
The Discord server carries 95+ tools across guilds; point it at the incident channel and leave the rest of the workspace out of the token's reach.
One command10 parts
npx mcprush@latest stack add on-call

Nothing in this stack installs from one command today: 10 members are either paid, run from its own source, or a skill with its own command — the steps above name each one. Nothing is connected until you approve it.

Before you start
  • 6 members have not declared what credential they need — check each one’s own page before you start.
  • What this stack can write is not recorded — 5 members of 6 have no imported tool surface. Section 05 says what is known before you approve anything.
  • 6 members can run on your own machine instead of ours, if you would rather they did.
05 · Permissions

What the whole stack can reach

Installed together, these tool surfaces add up. It is the first thing a security reviewer asks for, so what has been counted — and what nobody has counted yet — is on the page rather than in a PDF.

tools that only read
Which tools read and which write is counted from the surface a publisher imports, and 5 members of 6 have no imported tool surface.
tools that can change something
Not recorded, and not estimated: a total added up from the members that have a surface would be read as the whole stack’s. The table below is what is known, member by member.
0
tools that reach the network
2 of the 2 never leave your machine — they belong to the local members.
Who grants the write accessnot recorded
MemberTool surfaceWrite tools
Datadog not imported not recorded
Discord not imported not recorded
Kubernetes not imported not recorded
Prometheus not imported not recorded
Sentry not imported not recorded
Reading the number

A stack's blast radius is the union of its members, not the worst of them. That union cannot be taken here, because 5 members of 6 have no imported tool surface — so the figure a review asks for is missing rather than low, and a member marked not imported is one nobody has counted rather than one that cannot write.

Tools in total2
Write sharenot recorded
Members that can writenot recorded
Surface imported1 of 6 members
06 · Swaps

Sensible swaps

A stack is a default, not a verdict. These are the substitutions the maintainer would make, and what each one costs or saves.

A shop whose dashboards and logs are in Grafana queries them directly, and the desk stops depending on a paid monitoring account.

same money at 5k calls the same 0 tools both grade A

The same update, into the workspace where your team actually gets paged, over a user-token Slack server.

same money at 5k calls the same 0 tools both grade A
07 · Limits

Where this stack stops

Written by the maintainer, kept on the page rather than in a support thread.

  • It cannot roll back or scale anything on its own account. Kubernetes here inspects and manages what you grant it, and a rollback is a decision the human on call makes.

  • It does not page anyone. Routing, escalation and acknowledgement stay in the tool that owns your rotation.

  • Metrics answer what happened, never why. The runbook and the error tracker are in this stack because the graph alone has never closed an incident.

08 · Maintenance

Who keeps this current

A stack has an owner: whoever keeps it re-checks the combination when a member changes, and the members themselves are published by the people named on each row.

Maintainer
mcprush
Publishes this stack only
Version
null
Set by the curator
Member installs
none yet
Added up across the parts. Nothing counts installs of the set as one thing.
Member rating
No reviews yet
None of the parts has been reviewed, and neither has the stack.
Composition
6 + 4
6 servers, 4 skills
Nearby

Stacks that share parts with this one

All 20 stacks