Mentu

Introducing the Monitor Protocol

Most agents watch the world from inside their own session. When the session ends, the watching ends. We tried something different: we took the Monitor tool idea out of the session and made it durable, shareable and accountable. Today we are publishing it as an open protocol.

It turns out most AI agents watch the world the same way: from inside their own session.

Say an agent needs to know when a build fails. It polls. Or it starts a watcher, like the Claude Code Monitor tool, which runs a command in the background and turns each line of output into a message for the session. That is the right instinct. You point it at something, it watches, and it tells you when something happens.

But the watcher lives where the session lives. It stops when the session stops. It runs for thirty minutes at most. Only the session that started it can see it. And it remembers nothing. So every session starts its own watches again, and whatever happened in between is gone.

We tried something different. We took the Monitor idea out of the session and made it a service of its own. Durable, so it keeps watching when every session is closed. Shareable, so people, agents and other monitors can subscribe to it. Accountable, so every entry says who saw it and how much to trust it.

Today we are publishing it as the Monitor Protocol, version 0.1: a specification, JSON schemas, a conformance suite with 31 checks, and a reference server on npm, all under the Apache 2.0 license.

What's a monitor?

In the Monitor Protocol, a monitor is a definition: what is watched, how often, and with what authority. Once it exists, three things happen around it.

  • Producers publish observations to it. An observation is one thing seen, with where it came from, who caused it, and how much it is worth believing.
  • The monitor keeps a state. The state is a running summary computed from the observations, and it says out loud what it does not know.
  • Readers subscribe. Each one gets its own place in the stream, reads at its own pace, and acknowledges only after it has done the work.

A fifth part, the lease, lets two workers share the work without doing the same job twice. And configuration is not a sixth part: pausing a monitor, changing its filter or retiring it is recorded as an observation too, so the question "why did this change?" always has an answer.

Anatomy of an observation

Here is what a monitor stores when a CI probe reports a failed build. It is a CloudEvent, so any tool that reads CloudEvents can read it. The provenance is trimmed for space.

{
  "specversion": "1.0",
  "id": "4",
  "source": "monitor:ci",
  "type": "com.example.ci.run",
  "subject": "build-412",
  "time": "2026-09-22T17:10:11.188Z",
  "sequence": "00000000000000000004",
  "tier": "measured",
  "origin": "probe",
  "verified": "machine_verified",
  "horizon": "minute",
  "actor": "probe:ci",
  "data": {
    "status": "failed",
    "provenance": {
      "actor": "probe:ci",
      "on_behalf_of": null,
      "attested": false,
      "source_ref": "build-412"
    }
  }
}

Three fields do most of the work. sequence is the observation's place in the monitor's log, so a reader always knows where it is. tier says how much the claim can carry: measured means an instrument measured it. And data.provenance records who saw it, who it was for, and whether anyone independent vouched for the monitor's owner.

What's wrong with watching from inside a session?

Watching from inside a session has four problems, and they add up.

It forgets. When the session ends, whatever the watcher saw ends with it. The next session starts blind.

It is private. One session watches and one session hears. A teammate or a second agent that wants the same signal has to start its own watcher, with its own gaps.

It cannot say how much to trust what it saw. A line of output is a line of output. Nothing records whether a person checked it, a probe measured it, or an agent guessed.

Notifications get lost. Anything that pushes a message can fail to deliver it, and the MCP specification says as much: its notifications are best effort. If the notification is the only record, a lost notification is a lost event.

But the Monitor tool is still useful, because it is simple

We did not want to replace the part that works. A session that gets one line per event, as it happens, is a good way for an agent to follow something.

So we kept it. The protocol's command line tool has a watch command made for the Monitor tool. It follows a subscription, prints one line per observation, and acknowledges each batch once its lines are printed. When the Monitor's time runs out, you start it again, and it picks up exactly where it stopped. The session still gets its live feed. The difference is that the feed now sits on top of something that remembers.

Monitor(command: "npx -y @mentu/monitor-protocol watch --base http://localhost:8130 --subscription sub-8f2a91c0 --token $READER_TOKEN --catch-up")

OK, how does it work?

The loop has two halves.

On the writing side, a producer publishes an observation to a monitor. The monitor checks it, appends it to its log with the next sequence number, and updates its state. If the observation breaks a rule, it is refused, and the refusal itself is appended as an observation. Nothing disappears silently.

On the reading side, a subscriber pulls. A pull returns everything after the subscriber's cursor and never moves the cursor. The subscriber does its work, then acknowledges, which moves the cursor forward. If it crashes before acknowledging, the next pull returns the same observations, marked redelivered. That is at-least-once delivery, in Kafka's words, and it is the guarantee. Wake-ups, whether an MCP notification, a server-sent event or the Monitor tool, only tell a reader to look.

# define what is watched
curl -s localhost:8130/mp/v0/monitors \
  -d '{"id":"ci","name":"CI","horizon":"minute","capabilities":["observe"],"visibility":"public","types":["com.example.ci.run"]}'
 
# a producer records what it saw
curl -s localhost:8130/mp/v0/monitors/ci/observations -H "Authorization: Bearer $OWNER_TOKEN" \
  -d '{"type":"com.example.ci.run","subject":"build-412","actor":"probe:ci","tier":"measured","origin":"probe","data":{"status":"failed"}}'
 
# a reader subscribes, takes what it has not seen, and commits after handling it
curl -s localhost:8130/mp/v0/subscriptions -d '{"monitor":"ci","subscriber":"agent:claude","capabilities":["observe"]}'
curl -s "localhost:8130/mp/v0/subscriptions/$SUBSCRIPTION/pull?wait=25" -H "Authorization: Bearer $READER_TOKEN"
curl -s localhost:8130/mp/v0/subscriptions/$SUBSCRIPTION/ack -H "Authorization: Bearer $READER_TOKEN" -d '{"cursor": 5}'

A state that says what it does not know

Most dashboards have a quiet failure mode. When a number is missing, a program fills it in with zero, and the dashboard shows zero with full confidence. We learned this one the hard way: a missing measurement turned into a confident one, and no test failed.

The Monitor Protocol forbids it. A monitor's state carries a confidence, and the confidence lists its inputs as present or missing. While any input is missing, the value stays empty. The state also lists its gaps by name: independence_unknown until something checks that the sources are independent, no_subscribers when nobody is listening, single_actor when everything comes from one source, stale_source when the source has gone quiet, and unattested_origin when nobody independent vouched for the owner.

"confidence": {
  "value": null,
  "inputs": {
    "present": ["observations", "ages", "actors", "contradictions_open"],
    "missing": ["independence", "corroboration", "track_record"]
  },
  "gaps": ["independence_unknown", "unattested_origin"]
}

A source that did not answer is reported as not answering, never as zero. And a reader's count of waiting items covers only what that reader would actually receive, so it can reach zero.

A machine cannot rate itself, and your agent can act for you

Every observation carries a tier, from src (a person checked it against the source) through measured, derived and unverified, down to falsified (checked, and found untrue). The rule we would defend hardest is simple: a machine cannot give itself a person's level of certainty.

A probe can report measured. An agent's own claims start at unverified, and only a person promotes them. If a bot asserts src, the server refuses with TIER_NOT_ASSERTABLE and keeps the refusal in the log.

But agents also work for people, and a rule that locks out a person's own agent is not a good rule. So an agent can act at its owner's level by naming the owner in one field:

{
  "type": "com.example.ci.review",
  "actor": "agent:claude",
  "on_behalf_of": "human:you",
  "origin": "human",
  "tier": "src",
  "data": { "status": "approved" }
}

The record keeps both names: the agent that saw it, and the person it worked for. The agent can only name the person who owns the monitor, the person whose token it holds, so it cannot borrow anyone else's level. Naming a stranger is refused with PROVENANCE_CEILING.

We kept this light on purpose. Trust starts from relationships that already exist, and the one that exists here is the monitor itself. Where the protocol cannot check something, it says so instead of pretending: a server that does not verify who owns a monitor marks every state it serves with unattested_origin. The details are in Agents acting for people.

Built from parts that already work

We did not want to invent a transport, an envelope or a delivery model. Each of those already has a good answer.

Concern Taken from What the Monitor Protocol adds
Event envelope CloudEvents 1.0 attributes for sequence, tier, origin, verification, horizon and actor
AI client door MCP, as the extension ai.mentu/monitors tools, monitor:// resources and wake-ups
Filters Nostr filter objects keys for tier, origin and horizon, and tag matching
Delivery Kafka consumer offsets and Kubernetes leases a cursor that never moves backwards, and acknowledgement after processing
Provenance W3C PROV terms a tier ladder a machine cannot climb on its own

The protocol owns only the rules about evidence: who saw it, how much it can carry, and what is still unknown.

Tested like a protocol

A protocol is only as good as the test that says an implementation follows it. The conformance suite has 31 checks and two runners, one in TypeScript and one in Python, and you can point it at any server:

npx @mentu/monitor-protocol conform --base http://127.0.0.1:8124

We also tested the test. An independent audit broke the reference server's rules on purpose, one at a time, and found that the suite passed servers it should have failed, including one whose provenance check could be fooled from inside the request. We added the checks that catch those servers before fixing the server, so every fix is one the suite can see. Today the reference server passes all 31 checks, and a second implementation, the event log we run our own work on, serves the protocol and runs the same suite.

Try it now!

The quickest way is the playground. It loads the reference engine from the npm package straight into the page, so you can publish, subscribe, crash a reader and watch nothing get lost, without installing anything.

To run it yourself, start a server:

npx @mentu/monitor-protocol serve --port 8130

Then follow the quickstart, or connect Claude Code with Claude Code and MCP.

The Monitor Protocol is at version 0.1. The specification, schemas and conformance suite are on GitHub, and the package is on npm. If you build an implementation, run the suite against it and open an issue to tell us what it taught you.

© 2026 Mentu.