PUBLIC SCORING RUBRIC · v1.0

Scanner methodology

Every /scan score is a weighted fraction of applicable checks. This page defines the full check catalog, the tier weights, and the rules that keep the measurement neutral and reproducible. The rubric is versioned; changes are additive and published here.

1. Passive only

Remote scans perform the MCP initialize and tools/list JSON-RPC methods over Streamable HTTP. No tool is ever invoked and no adversarial payload is sent to the target. If tools/list requires OAuth, tool checks leave the denominator instead of failing.

2. Applicability

A check the target cannot satisfy by design (no mutation tools, read-only server, no RAG surface) is excluded from the denominator rather than scored as a failure. Score = earned weight / applicable weight.

3. Product neutrality

No check rewards a specific product. zn-gate appears only as one of several remediation options, always listed next to an open-source alternative. Reachable grades A+ to F require no purchase.

Grade bands

GradeScoreMeaning
A+95 to 100Certified Resilient
A85 to 94Strong Baseline
B70 to 84Moderate Risk
C50 to 69Elevated Risk
Fbelow 50Critical Vulnerability

Essential tier

12 pts each

Basic security. Missing controls score as failures.

TLS 1.2+ transport encryptionRFC 8446tls_transport

Endpoints reachable over HTTPS with a valid certificate chain. Cleartext HTTP transport is rejected.

Fix: Terminate TLS at the edge with a managed certificate and redirect all cleartext requests.
Open alternative: Let's Encrypt + Caddy or Nginx certbot with HSTS.

Authenticated transport boundaryOWASP LLM07auth_boundary

MCP endpoints enforce an auth boundary (OAuth 2.1 with PKCE or signed bearer tokens) before serving tools.

Fix: Require OAuth 2.1 authorization-code flow with PKCE on every MCP transport endpoint.
Open alternative: Open-source identity providers (Keycloak, Ory Hydra, oauth2-proxy) expose the same flows.

Input bounds and payload limitsOWASP LLM02input_bounds

Declared parameters include max lengths, enums, or patterns so a single call cannot carry unbounded payloads.

Fix: Constrain string parameters with maxLength or pattern and reject oversize bodies at transport level.
Open alternative: JSON Schema maxLength plus framework-level body size guards (e.g. express body-parser limits).

Tool description hygieneOWASP LLM01tool_desc_hygiene

Tool descriptions and metadata scanned for adversarial imperatives buried in docstrings (indirect tool poisoning).

Fix: Rewrite descriptions as neutral capability statements; strip imperative clauses.
Open alternative: mcp-scan (Invariant Labs) and LlamaFirewall both flag poisoned descriptions.

No unconstrained shell primitivesOWASP LLM08shell_primitives

Tools exposing arbitrary shell, eval, or command execution on the host without sandboxing are critical by design.

Fix: Replace open shell tools with granular scoped operations or run them inside a sandbox (Docker, Firecracker).
Open alternative: Firecracker microVMs, gVisor, or Docker with read-only rootfs and dropped capabilities.

Filesystem mutation boundedOWASP LLM08fs_mutation

Write or delete tools must be scoped to a declared sandbox root. Unscoped paths fail.

Fix: Bind the server to a single workspace root with respect for hidden-file exclusions.
Open alternative: The reference MCP filesystem server scopes roots the same way.

No plaintext credential parametersOWASP LLM07credentials_exposure

Tool parameters must not accept API keys, passwords, or tokens as model-visible arguments.

Fix: Inject credentials at the gateway layer via environment or secret manager, never through tool arguments.
Open alternative: Vault, AWS Secrets Manager, or Doppler with env-scoped injection.

Recommended tier

6 pts each

Good practice for exposed servers. Failures subtract weight.

Strict parameter schema typingOWASP LLM08param_typing

Every tool parameter declares an explicit primitive type and required set. Open object types with no declared properties fail.

Fix: Declare types (string, integer, boolean) and required fields in every tool inputSchema.
Open alternative: Zod, Pydantic, or vanilla JSON Schema generate strict schemas the same way.

Human gate on high-stakes mutationsOWASP LLM08financial_human_gate

Irreversible fin actions (refund, transfer, charge, deploy) require an auditable human approval step.

Fix: Add a two-step confirmation flow with server-side approval tokens for mutation tools.
Open alternative: Any webhook-based approval queue (Slack bot, Linear flow) implements the same gate.

Tool-output filtering capacityOWASP LLM02output_filter_capacity

The target must show some mechanism to inspect or filter content returned by tools before it reaches the model, regardless of vendor.

Fix: Filter tool outputs inline (neural gate, regex policy, or framework middleware) before re-injecting into context.
Open alternative: LlamaFirewall, LLM Guard, or Azure Prompt Shield provide equivalent capacity.

Tool namespace isolationOWASP LLM08namespace_isolation

Context cannot redefine, shadow, or override registered tool names from untrusted documents.

Fix: Rebuild the tool registry each turn and reject shadowed names.
Open alternative: The reference MCP SDK keeps client-controlled tool names immutable.

Bonus tier

2 pts each

Extra credit. If absent by design, the check is excluded from the denominator: it never subtracts.

Zero data retention policy publishedOWASP LLM06zdr_policy

A published, linkable statement about what the server retains from tool calls or conversations.

Fix: Publish a data-retention statement on the site.
Open alternative: A public docs page or trust report satisfies this without any vendor.

Cryptographic audit trailOWASP LLM06audit_trail

Verdicts and mutations produce hash-chained or signed evidence records.

Fix: Log verdicts with hash chaining or append-only storage.
Open alternative: Hash-chained logs with OpenSSL or an append-only Merkle log implement this pattern.

Remote scan mechanics

  • DNS resolution is validated before every connection; requests pin one resolved address (no DNS rebinding window).
  • Private, loopback, link-local (including 169.254.169.254), CGNAT, multicast, and documentation ranges are rejected at DNS time and again per redirect hop.
  • Maximum 2 redirects, 4 second connect timeout, 512 KB response cap per request.
  • Rate limit: 10 scans per hour per source IP.
  • TLS certificates must validate against the scanned hostname (SNI, full chain).
  • Every report carries per-check evidence and a permanent URL (/scan/r/{id}/) with JSON, Markdown, and badge variants, so any score can be re-verified by running the same scan again.

Static mode (pasted artifacts)

Pasted system prompts and tool schemas are analyzed in the browser and never uploaded. The same lexicons and tier weights apply; transport checks are excluded from the denominator because pasted artifacts cannot carry them. This mode proves what the artifacts themselves declare, nothing more.

Normative references

  • TLS 1.3 transport: RFC 8446.
  • Model Context Protocol specification (2025-06-18): initialize handshake, tools/list, Streamable HTTP transport, authorization via OAuth 2.1.
  • OAuth 2.0 Protected Resource Metadata: RFC 9728.
  • OWASP Top 10 for LLM Applications (2025): LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM06 Excessive Agency, LLM07 System Prompt Leakage, LLM08 Vector and Embedding Weaknesses.
Rubric versions are additive. The current version is v1.0. Report an inaccuracy or request a check addition through the contact form.