Conversation
324d341 to
62c0e78
Compare
| // events surface no spec-following client can discover, and SKIP would | ||
| // report it as a clean run. Distinguish the two by asking. | ||
| const probe = await eventsListPage(conn); | ||
| if ('error' in probe && probe.error.code === JSONRPC_METHOD_NOT_FOUND) { |
There was a problem hiding this comment.
Flagging this deliberately, because it is the one place this suite departs from how every other extension scenario in the repo behaves, and the departure is what turns the run against mcpkit red rather than green.
Everywhere else here, an optional capability the server never declared is a clean SKIP. That is correct when the server genuinely does not implement the feature. It is wrong when the server implements it and never says so: mcpkit answers events/list, serves all three delivery modes, and declares nothing under capabilities.events, so a client that reads capabilities to decide what to call never reaches any of it. A SKIP reports that as a clean run over a surface no client can use.
So the gate asks before it skips. One extra events/list separates "does not do events" (-32601, skip everything) from "does events off-menu" (returns a catalog, fail the declaration check and keep grading the rest).
The precedent is declaredSkillsCapability in src/scenarios/server/skills/helpers.ts, split out from skillsCapability for the same reason: a malformed declaration was being folded into "not declared" and skipping the whole suite, which "reads as a clean run against a server that is plainly wrong". This applies that argument to an absent declaration rather than a malformed one.
The cost is one extra request against servers that do not implement events at all. If you would rather pay nothing there and accept the false SKIP, say so, but the finding seems worth the request.
… phase 1 Adds the requirement-traceability yaml for MCP Events plus the first two of five server scenarios, scoring against the design sketch that merged on main of modelcontextprotocol/experimental-ext-triggers-events on 2026-09-08. The SEP number is a placeholder. Events has no PR in modelcontextprotocol/modelcontextprotocol, and both traceability gates are numeric (the filename regex at src/traceability/index.ts:174, the check-id regex at line 41), so an unnumbered file is dropped from the manifest without an error. 9999 is far from the live range; the rename is mechanical and the yaml header names the rename as a merge blocker rather than a follow-up: a merge under 9999 would publish it as a real SEP. The reservation question is open with the WG. src/seps/sep-9999.yaml declares 131 checks and 30 excluded rows against 144 RFC 2119 keyword occurrences. Pure MAY and OPTIONAL sentences get no check. Client, host, receiver and SDK-guidance obligations are excluded rather than declared, so the denominator holds only what a server-side run can observe. events-discovery covers the capability, events/list, the descriptor fields and the error-code contract. events-poll covers poll delivery, the EventOccurrence shape and cursor lifecycle, driving a real two-poll quiet-period loop so cursor advancement is gradeable against a server with no traffic. Together they emit 45 rows; the other 86 report untested until the push and webhook scenarios land. The capability gate asks before it skips. An optional capability a server never declared is not a defect, but a server that answers events/list while declaring nothing has a surface no spec-following client would reach, and a SKIP would report that as a clean run. mcpkit is in exactly that state. Against examples/events/kitchen-sink at mcpkit main: discovery 8/11, poll 18/28. Six divergences, five of them server-side defects not previously tracked, one confirming the known nextPollSeconds drift. The yaml header lists each. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017kyEZdgh1zjdhURgdcZ4nL
62c0e78 to
cb0f5af
Compare
|
Superseded by modelcontextprotocol/conformance#504, opened as a draft against upstream so the WG can review it there. Same branch ( |
What changes
Adds a requirement-traceability yaml for MCP Events and the first two of five server scenarios, scoring an implementation against the design sketch that merged on
mainofmodelcontextprotocol/experimental-ext-triggers-eventson 2026-09-08.events-discoverygrades the capability,events/listand the error-code contract;events-pollgrades poll delivery, theEventOccurrenceshape and cursor lifecycle. Run against mcpkit'sexamples/events/kitchen-sink, they report 26 pass / 13 fail / 6 warn, and the failures are six real divergences rather than harness noise.Prerequisite knowledge
AGENTS.md§ Scenario design and § Check conventions — why this is two scenarios with 45 checks instead of 45 scenarios, and why a missing prerequisite fails rather than skips.src/traceability/— how a yaml row becomestestedoruntestedin the manifest, which is what makes declaring all 131 rows up front the right move rather than an overreach.Reviewer's guide
A restaurant can cook a dish that is not on the menu. The kitchen has the ingredients, the line knows the recipe, and if you somehow know to ask for it you get it. But nobody reading the menu ever orders it, so from the outside the dish may as well not exist. mcpkit's events surface is that dish:
events/listanswers, the three delivery modes are implemented, and yet nothing appears undercapabilities.events, which is the menu a spec-following client reads before deciding what to ask for. This matters for how the suite is shaped, because the obvious design gets it wrong. Conformance harnesses normally treat an undeclared optional capability as "not applicable" and skip, which is right for a restaurant that genuinely does not serve the dish and wrong for one that serves it off-menu: the skip reports a clean run over a surface no client can reach. So the capability gate here asks before it skips. It callsevents/list, and only a server that both declares nothing and implements nothing gets the skip. The same instinct runs through the rest of the diff. Wherever a check cannot be exercised, the scenario says which prerequisite was missing and fails, because the one thing worse than a red conformance run is a green one that tested nothing.Read in this order:
EVENTS_EXTENSION_IDis a suite-selection key, not a path into the capability object; Events declares top-level, unlike every other extension in this repo.checks(). Everything after it is field-by-field grading of whatever the server serves.quietAdvanceChecks: two polls with the first response's cursor fed into the second, which is what makes cursor advancement gradeable against a server where nothing is happening.How it works
The
truncatedrows that need a stale cursor are reported untestable rather than skipped. Cursors are opaque by definition, so the harness cannot mint one old enough to fall outside a retention window, and saying so is more useful than a silent pass.Decision log
modelcontextprotocol/modelcontextprotocol, so no number exists to use. Both traceability gates are numeric (the filename regex atsrc/traceability/index.ts:174, the check-id regex at line 41), and a non-matching filename is dropped from the manifest silently rather than rejected, so there is no honest name that works. The alternative was blocking the extraction on a WG round trip. The rename is mechanical, the yaml header names it as a blocker before upstreaming, and the file stays on this fork branch until then soplan.modelcontextprotocol.ionever sees 9999.untestedin the manifest, which is the mechanism working rather than a gap in this file.excluded:rather than deleted, so the keyword sweep stays auditable against the source.Risk / blast radius
EXTENSION_IDS. No existing scenario, suite list or check id changes.sourceis{ extensionId }, somatchesSpecVersionreturns false for every--spec-versionand the scenarios never join a dated run. They are also in the pending list. Reaching them takes--scenario events-*or--suite all.src/scenarios/server/events/negative.test.ts, 29 cases. Full suite 651 pass across 48 files.src/seps/traceability.jsonwill drift once these check ids are emitted in a real run. PerAGENTS.mdthe traceability workflow refreshes it by PR and it is not a gate, so this PR leaves it alone.Before / after
There was no events suite before, so the comparison that matters is what an implementation looks like when scored. Against
examples/events/kitchen-sinkat mcpkitmain:The six divergences, each verified against the running server rather than inferred:
sep-9999-capability-events-objectevents/listwhile declaring nocapabilities.eventssep-9999-descriptor-input-schemainputSchemasep-9999-descriptor-delivery-subsetevents.topologysendsdelivery: nullwhere an array is requiredsep-9999-poll-events-arrayeventsinstead of returning an empty arraysep-9999-poll-mode-unsupportedevents/pollfor a type advertising onlypush/webhooksep-9999-poll-next-poll-msnextPollSeconds, the name spec commit197c32b4retiredOnly the last was previously tracked. The other five are new.
flowchart TD A[server under test] --> B{capabilities.events declared?} B -->|yes| G[grade everything] B -->|no| C{events/list answers?} C -->|"-32601"| D[SKIP all: does not do events] C -->|returns a catalog| E[FAIL the declaration check] E --> G G --> H{check exercisable?} H -->|yes| I[SUCCESS or FAILURE] H -->|no| J[untestable: FAIL and name the prerequisite]Out of scope
events-push,events-webhook,events-webhook-delivery— phases 2 and 3. The 86 rows they cover are declared and reportuntested.--headeroption on theserverrunner. Needed to score Peter'smetronome-mcp.fly.dev, which requires an OAuth bearer and whose AS offers onlyauthorization_code. Its own PR, ahead of phase 3.sep-9999-*once the WG reserves a number.🤖 Generated with Claude Code
https://claude.ai/code/session_017kyEZdgh1zjdhURgdcZ4nL