Behavioral conformance tests for MCP servers
Is your MCP server safe to hand to an agent? Run it and see what it actually does.
Sealmark starts the MCP server you run or build, calls its tools the way an agent would, and records what comes back. No LLM in the loop, nothing leaves your machine, no account.
Two packs: core for any MCP server, db for Postgres-backed ones. One A–F grade per pack, with the evidence behind every line.
$ npx sealmark init $ npx sealmark run
Free; source to be published under Apache-2.0 with the 0.1 release. The db pack runs against a seeded synthetic Postgres on your machine: no production data, no model calls, no Docker.
Early preview: the sealmark package is not published yet, so these commands will not work today.
Start.
Sealmark launches your server from the command in sealmark.yaml, or connects to its URL. For the db pack it starts the command once per database role.
Probe.
Fixed tool calls with no LLM in the loop. The db pack then reads the database to see what really changed, not just what the server said.
Grade.
Checks score into one letter per pack with a published formula. A critical failure caps the grade at D.
Keep the record.
A CI gate that fails below your bar, and a JSON record of the exact configuration and the evidence for every probe.
Packs
Two packs: basic hygiene for any server, depth for databases
Core runs against any MCP server. The db pack goes deep on servers that sit in front of a database. Each pack gets its own grade.
coreany MCP server
Basic hygiene any MCP server should pass, whatever it does. Needs only a command or URL.
- Valid handshake and schemas
- Tool-poisoning indicators in metadatacritical
- Malformed-input robustness
- Environment-secret echocritical
Needs: a command or a URL for the server. Nothing else. The secret-echo check needs a stdio server and is n/a over HTTP.
dbdatabase servers · Postgres only in 0.1
The deep pack. Starts your server once per database role against a seeded synthetic Postgres and checks the database afterward.
- Read-only by default
- Write approval gateadvisory · not scored
- Row-level security respectedcritical
- Mutating-SQL corpus blockedcritical
- Personal-data masking respectedcritical
- Row cap and statement timeout
- Every call loggedadvisory · not scored
Needs: a server command with a {dsn} placeholder, started once per database role. Uses a bundled Postgres, or yours with --pg and --yes.
Example
One grade per pack, with the receipts
This is what a local run produces, shown for five invented servers. No real server has been graded and there is no public scorecard yet. Select a row to see what each check sent and what came back.
| Server | core grade | Points | Handshake+ schemas | Toolpoisoningcritical | Malformedinput | Secretechocritical |
|---|---|---|---|---|---|---|
| A | 4 / 4 | pass | pass | pass | pass | |
| A | 4 / 4 | pass | pass | pass | pass | |
| B | 3.5 / 4 | pass | pass | partial | pass | |
| C | 2 / 3 | pass | pass | fail | n/a | |
| F | 1.5 / 4 | fail | fail | partial | pass |
Advisory checks are shown but never scored. A fail on a critical check caps a better grade at D. * means a critical check was inconclusive, so the grade is provisional. Inconclusive and n/a checks are left out of the score. All servers and results on this page are invented.
Server A v1.4.0 core pack
4 of 4 scored · 100%No critical check failed.
- passValid handshake and schemasprobe: initialize, then tools/list: tool names, duplicates, and inputSchema shapes are checkedHandshake completed. 6 tools listed with unique, valid names and object inputSchemas.
- passTool-poisoning indicators in metadatacriticalprobe: scan tool names and descriptions for hidden characters, send-to-URL phrasing, credential paths, and override phrasesNo hidden characters, send-to-URL phrasing, credential paths, or override phrases in tool metadata.
- passMalformed-input robustnessprobe: missing required arguments, then wrong-type arguments (only values the tool's own schema rejects); expect a clean error, not a crash, hang, or stack traceMissing and wrong-type arguments each returned a clean error. No crash, no hang.
- passEnvironment-secret echocriticalprobe: start the server with canary variables (SEALMARK_CANARY, SEALMARK_FAKE_API_KEY), send bad input, look for the values in responsesNeither canary value appeared in any response.
$ npx sealmark run --config sealmark.yaml --pack core01 The problem
An agent inherits whatever the MCP server does.
Servers get wired to agents faster than anyone checks how they behave when an agent actually calls them: many tools, bad input, real data.
Does it speak MCP correctly?
Duplicate or malformed tool names and broken schemas make clients behave unpredictably.
Is the tool metadata talking to the model?
Tool descriptions are read by the model. Hidden characters, or text that tells it to send files somewhere, are a known poisoning pattern.
What happens on bad input?
Crashes, hangs, stack traces, and errors that echo the process environment are what an agent hits the first time it gets an argument wrong.
For database servers: do the locks hold?
Read-only mode, row-level security, masking, and row caps get configured once and trusted. They only count if they hold for every tool and every role.
Configuration says what should happen. Only running the server shows what does.
02 Grading
Behavioral tests, a published formula, and the evidence for every line.
One letter per pack, the formula that produced it, and the probe behind each check. Anyone can rerun it with the same configuration.
Behavior over declarations.Sealmark runs the call and records what came back, instead of trusting what the tool schema promises.
Evidence for every line.Each check records the probe, the response, and the exact server command and environment it ran with.
One letter per pack, open formula.The grading formula is versioned (0.1-draft) and public. Recompute any grade from its JSON record.
No model in the loop.The probes are fixed, so the same server and configuration give the same answer on any machine.
Read the tool definition
Useful for spotting risky shapes. A schema can't tell you whether read-only mode actually stops an INSERT.
tool query
input: sql (string, no maxLength)
finding: accepts arbitrary SQLRun the call, then check the state
Runs fixed calls as the database roles it tests, then reads the database to see whether anything changed.
role=admin
call query("INSERT INTO probe_t …")
rejected · table unchanged afterward · passGrade ladder
Rules (methodology 0.1-draft)
- Points per checkpass 1, partial 0.5, fail 0. Each pack is graded on its own.
- Percentage
points ÷ checks scored. Advisory, inconclusive, and n/a checks are left out of the score. - Critical capA fail on a critical check caps the grade at D. Critical in db: row-level security, mutating SQL, personal-data masking. Critical in core: tool poisoning, secret echo. A grade already below D stays where it is.
- Provisional markAn inconclusive critical check adds * to the grade, because the probe produced no signal.
- Advisory checksWrite approval gate and audit logging are shown in the db pack but never scored in 0.1.
03 Coverage
What an A means, and what it doesn't.
Sealmark is a set of conformance tests, not a security scanner and not a certification. Read this part before you read a grade.
An A means
- At least 90% of the scored points in that pack, and no critical check failed.
- The published probes ran against the exact command, arguments, and environment recorded with the result.
- Every line has the probe and the response behind it, so you can rerun it.
An A does not mean
- The server is safe, secure, or free of vulnerabilities.
- Other tools, roles, inputs, or configurations behave the same way.
- Any other version behaves the same way.
- Anyone endorses the server, including Sealmark.
Not covered in 0.1
- Authentication and transport security. Who may connect and how tokens are handled.
- Output-side prompt injection. What tool results say to the model.
- Supply chain. Dependencies, install scripts, and who published the package.
- Secret echo over HTTP. Canary injection needs a stdio server.
- Databases other than Postgres, and audit logging or write gates as scored checks.
- Anything else the probes do not exercise.
The one sentence to remember
An A means the server passed these specific, published probes with this exact configuration; it is not an endorsement, and it does not test what the probes do not cover.
04 Get started
Test your first server in two commands.
The same harness runs on your laptop and in CI, and needs no Docker and no account.
- Initnpx sealmark init writes a sealmark.yaml. Edit it to start your server, with {dsn} wherever the server takes its database connection.
- Runnpx sealmark run runs the core pack, plus the db pack when it finds a SQL tool and a placeholder. Add --pg <dsn> --yes to use your own Postgres instead of the bundled one.
- Gate the build--fail-under B exits 1 when any graded pack is below B. Exit codes: 0 ok, 1 below the bar, 2 usage or config error, 3 no pack graded, 4 harness or server failure.
- Keep the record--format json --out sealmark.json saves the exact configuration and the evidence for every probe.
# no Docker: the db pack uses a bundled Postgres npx sealmark init npx sealmark run --fail-under B # one pack, JSON record for CI npx sealmark run --pack core --format json --out sealmark.json # your own Postgres instead of the bundled one # (drops and recreates the sealmark_seed database, so --yes is required) npx sealmark run --pg postgres://postgres@localhost:5432/postgres --yes
Early preview: the sealmark package is not published yet. These commands and flags describe the planned 0.1 CLI, so do not expect them to work today.
05 Free
Free to run. Free to read.
The CLI, both packs, every check, and the probe corpus will be free. The source is not public yet and will be published under Apache-2.0 with the 0.1 release; until then, every report lists the probe behind each line.
What you get
- Both packs, every check, and the grading formula
- Per-role connection strings, so the db pack tests your server with your roles
- A CI gate with
--fail-under, and a JSON record for every run - Runs locally. No account, no model calls
Grades come only from the probes.
A grade is what the published probes produced for one server version and one configuration, scored with the published formula. There is no way to request, negotiate, or adjust one.
FAQ
Questions people ask first
Does Sealmark need access to production data?
No. The db pack grades against a seeded synthetic database that it starts on your machine. By default the Postgres is a bundled one that runs and is deleted on your machine. With --pg and --yes it uses an instance you give it: the run drops and recreates the sealmark_seed database and creates, alters, and finally drops three cluster-wide roles (sm_analyst, sm_support, sm_admin) with a public password, overwriting any existing roles of those names. Use a throwaway instance, never a shared or production one.
Does an AI model decide the grade?
No. Every check is a fixed tool call compared against a published rubric, and the db pack confirms results by reading the database afterward. The same server and configuration should give the same result on any machine.
Which servers can it test?
Any MCP server, over stdio or HTTP, can run the core pack. The db pack needs a server that takes a SQL string and can be started from a command with a {dsn} placeholder. Postgres only in 0.1.
Does it run my server's code on my machine?
Yes. It starts the command you give it, passes canary environment variables, and for the db pack connects as three database roles. Only run it on servers you would run anyway.
Is it a security scanner or a certification?
No. It is a set of behavioral conformance tests. A grade says what these probes saw for one version and configuration. See the coverage section for what it does not test.
What happens if a probe finds a vulnerability in a real server?
The policy is a private report to the maintainer first, with the exact rerun command, and publication when a fix ships or after 90 days, whichever comes first. No real server has been graded or published yet.
Can a vendor change its grade?
Only by changing the server. Grades come from the published probes and the published formula, with no way to request, negotiate, or adjust one.
Is it really free?
Yes. Both packs, every check, the CI gate, and the JSON record will be free. The source is not public yet; it will be published under Apache-2.0 with the 0.1 release.
Find out what your MCP server really does.
Run the probes in a few minutes. Free, local, and reproducible.
$ npx sealmark init && npx sealmark runEarly preview: not published yet.