A conformance test suite for Agent Plugins 1.0.0 client loaders. It ships 133 real plugin directories on disk, each paired with the load report a conformant client has to produce for it, plus a runner that diffs your client's answer against the corpus.
Six things that have already gone wrong, all in the last few months:
- agentplugins/agent-plugins-spec#77.
The published
plugin.schema.jsonsetsadditionalProperties: false, but Β§5.2 says a client MUST report an unknown top-level field, ignore it, and keep loading. Β§8.1 says the same for a non-objectextensions. Validate-and-reject, the obvious implementation, is non-conformant in both cases and the schema cannot express the difference. - openai/codex#39895. A root
plugin.jsonrouted Codex through the Agent Plugins loader, which has no hook support, so every hook declared in.codex-plugin/plugin.jsonstopped running. No warning, no error, no log line. Two plugins were dead for a week. - can1357/oh-my-pi#8853. omp 17.3.5
routes any package declaring an
agent-plugins.org$schemato its strict provider, which drops anySKILL.mdcarrying a frontmatter key outside the Agent Skills set. - EveryInc/compound-engineering-plugin#1411.
The downstream half of that: 3 of 33 skills loading, every
/skill:command gone. The fix was to delete$schemafrom the root manifest, so conforming cost them the standard. - dotnet/skills#1087. First-party manifests
with no
$schemaand withskills,agentsandmcpServersas top-level fields. Kiro reportedplugin.json has an unsupported or missing $schemaand refused the package. Adding$schemawas necessary but not sufficient: it then loaded with every functional component excluded. - VS Code's own troubleshooting page.
The largest shipping client has no validation surface at all. It tells you to open
SKILL.mdand check thenamefield by hand, because "Invalid names cause the skill to be silently skipped."
Every one of these is a loader disagreeing with the specification, and every one of them was found by a user rather than by a test. That is the gap this fills.
This tests clients, not packages. If you want to check a plugin you are publishing, use @hiai-gg/agent-plugins-doctor or @booyaka/mcp-vet. They do the package-side job and this kit deliberately does not.
npm install --save-dev agent-plugins-conformance-kit
Node 22 or newer.
Write an adapter. It takes a plugin directory as argv[2] and prints one JSON load
report to stdout. That is the whole contract, and it is usually about twenty lines:
#!/usr/bin/env node
import { loadPlugin } from 'your-client';
const result = await loadPlugin(process.argv[2]);
console.log(JSON.stringify({
rejected: result.ok ? null : result.reason,
loaded: {
skills: result.skills?.map((s) => s.name) ?? [],
mcpServers: result.servers?.map((s) => s.name) ?? [],
},
skipped: result.dropped?.map((d) => ({ what: d.path })) ?? [],
reported: result.warnings?.map((w) => ({ field: w.field })) ?? [],
}));Then run the suite:
npx apconform run --adapter ./my-adapter.mjs
ADAPTERS.md has the full contract, including the skipped path vocabulary and how to
use a non-Node adapter.
Run against two independently written 1.0.0 clients. This is the first,
@kuralle-agents/plugins 0.25.0,
through adapters/kuralle.mjs. Both the image above and this block come from
npm run capture:images, which renders real output, so neither can drift from what the
tool prints:
$ node dist/cli.js run --adapter adapters/kuralle.mjs --quiet
apconform 1.0.0 133 fixtures adapter adapters\kuralle.mjs
FAIL AP-4.1-BOUNDARY-COMPONENT-LOCATION (spec 4.1) expected loaded.skills not to contain [alpha], got [alpha]
WARN AP-4.1-BOUNDARY-COMPONENT-LOCATION (spec 4.1) expected skipped to contain [skills], got []
SKIP AP-4.1-BOUNDARY-MANIFEST (spec 4.1) this fixture needs a symlink and the platform refused to create one (enable Developer Mode on Windows, or run the suite on Linux)
FAIL AP-4.1-BOUNDARY-SKILL (spec 4.1) expected loaded.skills not to contain [escaped], got [alpha, escaped]
WARN AP-4.1-BOUNDARY-SKILL (spec 4.1) expected skipped to contain [skills/escaped], got []
WARN AP-7.1-IMMEDIATE-CHILD (spec 7.1) expected skipped not to contain [skills/beta], got [skills/beta] (core/AP-7.1-IMMEDIATE-CHILD__skill-md-is-directory)
core 122 pass 2 fail 0 error 1 skipped
disputed 4 pass 0 fail 0 error 0 skipped
regressions 4 pass 0 fail 0 error 0 skipped
total 130 pass 2 fail 0 error 1 skipped 3 warnings
Both failures are the same defect at two levels. Β§4.1 lists five failure boundaries for a
path that resolves outside the plugin root. That loader enforces the first, for
plugin.json, and not the two for skills/, so a link under skills/ is followed and
the skill behind it loads. Reported as
kuralle/kuralle-agents#23.
The SKIP needs a file symlink, which Windows refuses without Developer Mode. It runs and
passes on Linux and macOS, so the totals there are 131 pass, 2 fail, 0 skipped.
A suite that only ever runs against one implementation cannot tell "the client is wrong"
from "the fixture is wrong". So the corpus is also run against
pi-agent-plugins 0.1.8, a client
written by someone else with its own conformance document:
core 125 pass 0 fail 0 error 0 skipped
disputed 4 pass 0 fail 0 error 0 skipped
regressions 4 pass 0 fail 0 error 0 skipped
total 132 pass 0 fail 0 error 1 skipped 10 warnings
It passes everything, including the two Β§4.1 fixtures the first client fails, which is what makes those two a defect in that client rather than an argument in this corpus.
Doing that found three problems in the corpus, all now fixed:
- Two fixtures required a remote MCP entry to be activated when the rule under test was
only that headers and urls are not expanded. A client that validates such an entry and
then declines to connect, for documented safety reasons, is not violating a MUST. One
fixture lost the header from its control server, the other became
partialwith the entry optional. AP-8.1-EXTENSIONS-MEMBER-OBJECTSwas incore. Both clients read Β§8.1's report-and-ignore exception as covering member values, against my stricter reading. Two independent implementations agreeing against me is not a defect in them, so it moved todisputedand now accepts either outcome, pending spec#77.
fixtures/
βββ core/ 125 fixtures. The specification states the required outcome directly.
βββ disputed/ 4 fixtures. More than one outcome is conformant. Recorded, not graded.
βββ regressions/ 4 fixtures. Ported from a real bug report, with the issue cited.
Each fixture is a directory holding a real plugin/ tree and a fixture.json:
{
"ruleId": "AP-5.2-UNKNOWN-FIELD",
"confidence": "core",
"spec": "5.2",
"title": "an unknown top-level manifest field is reported and ignored",
"rationale": "The published plugin.schema.json sets additionalProperties: false, so the obvious implementation (validate, reject on failure) is non-conformant here. This is the case spec issue #77 describes.",
"quote": "Clients MUST report and ignore each unknown field and MUST continue loading the plugin if the manifest otherwise satisfies this section.",
"issue": "https://github.com/agentplugins/agent-plugins-spec/issues/77",
"observability": "report",
"expect": {
"rejected": null,
"loaded": { "skills": [], "mcpServers": [] },
"skipped": [],
"reported": [{ "field": "skills", "ruleId": "AP-5.2-UNKNOWN-FIELD" }]
}
}rules.json holds all 89 rules behind those fixtures. Every one carries the normative
sentence verbatim, and npm run verify:sources proves each quote is still a byte-identical
substring of the published spec/1.0.0.md. apconform explain <rule-id> prints one:
$ npx apconform explain AP-7.1-DEPTH
AP-7.1-DEPTH
specification Β§7.1 of Agent Plugins 1.0.0
severity accept - The client must load this. Rejecting or skipping is the failure.
confidence core - The specification states the required outcome directly.
"Clients MUST NOT recursively search deeper descendants for additional skills."
https://raw.githubusercontent.com/agentplugins/agent-plugins-spec/main/spec/1.0.0.md
checklist.json maps every item on the
official non-normative client checklist
onto the rules that cover it, and states plainly where coverage is partial and why. It is
used to audit the rule table for gaps, never as a source of requirements: fixtures cite
specification sections.
A suite that grades everything equally is a suite people turn off. Three distinctions run through the whole corpus.
rejected is compared on rejectedness, not on the string. The specification defines no
rejection vocabulary, so your reason is echoed in failure output and never asserted.
Reporting a skipped component is a SHOULD. Β§7.1 and Β§7.2.2 say the client SHOULD report
an invalid skill or server entry, so a mismatch in skipped is a warning. Whether the
component actually loaded is a MUST, and that is checked against loaded. The two rules
that produce reported entries both say MUST report, so those are failures. Pass
--strict-reporting to promote the warnings.
Some outcomes are genuinely open. Where the specification permits more than one answer,
the fixture accepts both and records which one your client chose as a note. sse support
is OPTIONAL, Β§4.1 symlinks to in-root targets MAY resolve, and Agent Skills lists its
frontmatter fields without saying the set is closed. That last one is oh-my-pi#8853, and a
corpus that graded it would be picking a side in an argument the specification has not
settled.
Some rules are only partly observable through a load report. Whether PLUGIN_ROOT and
PLUGIN_DATA reach a subprocess environment, and what ${PLUGIN_DATA} expands to inside
args, are not visible in a report about what loaded. Those fixtures are marked
"observability": "partial" and carry a note saying exactly what is and is not asserted.
apconform list flags them. Nine fixtures are partial; the other 124 are fully asserted.
Nobody wires a new conformance suite into CI and gets a green run. Gate on change instead:
apconform run --adapter ./adapter.mjs --baseline conformance-baseline.json --update-baseline
Commit that file. From then on, --baseline conformance-baseline.json without
--update-baseline fails only on fixtures that regressed. Everything already failing is
reported as known and does not gate, fixtures you have since fixed print as FIXED, and
a fixture that is new to the corpus after a kit upgrade has to pass on its own, because
new coverage is exactly the thing you want to hear about.
A failure the baseline already records prints as KNOWN rather than as a red FAIL the
summary then contradicts:
- uses: Booyaka101/agent-plugins-conformance-kit@v1
with:
adapter: ./conformance/adapter.mjs
only: coreInputs: adapter (required), only, fixture, strict-reporting, timeout, baseline,
update-baseline, json, junit, version, fail-on-error. Outputs: report, passed, failed. It writes the summary
to the job summary and the full result to apconform-report.json.
Set fail-on-error: false to report without gating while you work through the list.
apconform run --adapter <path> [options]
apconform list [options]
apconform rules [--json]
apconform explain <rule-id>
apconform verify
apconform show <fixture-id>
| Option | Meaning |
|---|---|
--adapter <path> |
The executable to test. Required for run. |
--adapter-exec <cmd> |
Launch the adapter with this program instead of guessing from the extension. |
--only <groups> |
core, disputed, regressions or all. Comma-separated. |
--fixture <substring> |
Run only fixtures whose id contains this text. |
--json <file> |
Write the full machine-readable result here. |
--junit <file> |
Write JUnit XML, for CI systems that render test reports. |
--strict-reporting |
Treat SHOULD-report mismatches as failures. |
--timeout <ms> |
Per-fixture adapter timeout. Default 30000. |
--concurrency <n> |
Adapters to run at once. Default min(8, cpus). |
--fixtures <dir> |
Use a different corpus root. |
--baseline <file> |
Gate on change against this file instead of on the total. |
--update-baseline |
Write the current result to --baseline and exit 0. |
--quiet |
Only print the summary and anything that is not a clean pass. |
Exit codes: 0 all passed, 1 conformance failures or adapter errors, 2 the suite
could not run at all.
apconform show <fixture-id> prints a fixture: the rule it tests, the normative sentence,
the plugin tree on disk, the manifest and mcp.json, and the expected load report. It is
the fastest way to answer "what is this failing fixture actually doing".
apconform verify checks the corpus against itself: every fixture names a real rule, sits
in the folder its confidence implies, quotes that rule correctly, and expects only
component names that exist on disk with matching SKILL.md frontmatter.
This kit targets Agent Plugins 1.0.0, which is the version marked Published
and the only one whose canonical schema identifiers resolve.
A 1.1.0 working draft exists in the spec repository. As of writing it is a version
bump and nothing else: the prose is byte-identical apart from version strings and one
reworded conformance sentence, both schemas differ only in $id, description and the
$schema const, and https://agent-plugins.org/schemas/1.1.0/ still returns 404.
87 of the 89 rules here quote sentences that are unchanged in the draft. The two that
are not are AP-5.2-SCHEMA-CANONICAL-VALUE and AP-7.2.1-MCP-SCHEMA-CANONICAL-VALUE,
which embed the version in the canonical identifier and are expected to change.
Supporting a second version is a --spec flag over the same corpus rather than a new
corpus, so it can wait until 1.1.0 is published and its identifiers resolve.
- A load report describes what a client loaded. It cannot see subprocess environments, network access, or expanded placeholder values, so Β§9 is covered only where a violation changes what loads. The eight partial fixtures say so individually.
AP-5.2-NO-SCHEMA-FETCHneeds the network to be unavailable to bite. Run the suite offline to make it a real assertion.- One fixture needs a file symlink and reports
SKIPon Windows without Developer Mode. - The kit targets Agent Plugins 1.0.0 only. It has no compatibility-policy opinion about older or newer versions beyond rejecting identifiers it does not recognize.
- The corpus tests loading. It does not start MCP servers or execute skills.
npm install
npm test # builds, then runs 481 tests
npm run verify:sources # checks every quote against the live specification
npm test runs the diff engine against synthetic reports, asserts every rule quote is
verbatim in the published spec, walks the whole corpus for self-consistency, and drives
the runner end to end against a perfect-client echo adapter and against a deliberately
broken, hanging and non-JSON one.
To run the suite against the reference loader locally:
npm install --no-save @kuralle-agents/plugins @kuralle-agents/fs
node dist/cli.js run --adapter adapters/kuralle.mjs
examples/naive-adapter.mjs is a deliberately non-conformant loader, written the way an
implementer reaches for first. It is not a client and exists to show what the suite
catches:
$ node dist/cli.js run --adapter examples/naive-adapter.mjs --quiet
...
FAIL AP-5.2-UNKNOWN-FIELD (spec 5.2) expected rejected=null, got rejected="additional-properties"
...
total 68 pass 64 fail 0 error 1 skipped 52 warnings
npm run capture:images regenerates the README images from real output. The SVGs need
nothing; the PNGs need a Chrome on --remote-debugging-port=9222, because npm strips SVG
from rendered READMEs.
Add a directory under fixtures/core/<RULE-ID>/ (or <RULE-ID>__<variant> if the rule
already has one), put a real plugin tree in plugin/, write fixture.json, and run
apconform verify. If the rule is new, add it to rules.json with the normative sentence
copied byte for byte and run npm run verify:sources.
If a fixture encodes an outcome the specification does not actually require, it belongs in
fixtures/disputed/ with a note saying why, not in core/.
The highest-value thing to do with a new conformance suite is not to announce it, it is to use it. Comment on agentplugins/agent-plugins-spec#77 with the two fixtures that encode the disagreement it describes, since that issue is already asking for something executable and the maintainers are the exact audience. The findings in "Real output" above should be filed against kuralle/kuralle-agents the same way, one issue per boundary rule with the fixture id and the reproduction. A suite that has already found three real defects argues for itself better than a launch post does.
MIT

