# Galvanize-60M

The open-weights prompt-injection classifier behind the zn Cloud API, and how it was evaluated.

Source: https://usezn.com/docs/galvanize/

Galvanize-60M is zn's prompt-injection classifier for agent traffic. It is a 4-layer slice of `answerdotai/ModernBERT-base` (60M parameters) that keeps native rotary positions for inputs up to 8,192 tokens and runs in about 11.52 ms p50 on a standard CPU.

- Weights: [`usezn/Galvanize-60M`](https://huggingface.co/usezn/Galvanize-60M) on Hugging Face, Apache-2.0 (INT8 ONNX at `onnx/model_quantized.onnx`).
- Benchmark data: [`zn-prompt-injection-bench`](https://huggingface.co/datasets/usezn/zn-prompt-injection-bench) (23,699 rows, CC-BY-4.0).
- Background: [Introducing Galvanize-60M](https://usezn.com/blog/introducing-galvanize-60m/).

## Security pooling

Generic sentence embeddings confuse ordinary JSON, SQL and code in tool arguments with attacks. Galvanize-60M replaces CLS or mean pooling with four learned query vectors, concatenated into a 3,072-dimensional representation:

| Query | Focus |
| - | - |
| 0 | Instruction overrides and privilege escalation |
| 1 | Persona, role-play and hypothetical framing |
| 2 | Legitimate JSON, XML and Markdown versus delimiter escapes |
| 3 | Attempts to leak environment variables, memory or tokens |

## Benchmark

First-party evaluation published on 6 September 2026. It is reproducible with the public dataset, but it has not been run by an independent party.

| Metric | Galvanize-60M | Prompt-Guard-2-86M | Prompt-Guard-2-22M | ProtectAI DeBERTa-v3 |
| - | - | - | - | - |
| Tool false-positive rate | 1.00% (conservative headline; 0.67% measured at threshold 0.80) | 5.00% | 0.00% | 90.33% |
| Out-of-distribution recall (deepset) | 91.60% (calibrated) | 9.58% | 3.75% | 20.42% |
| Long-context needle recall | 77.00% to 97.00% | 7.00% | 0.00% | 1.00% |
| Adversarial robustness (attacks generated by our internal red-team model) | 94.00% blocked | 70.00% | 26.00% | 82.00% |
| CPU latency, p50 | 11.52 ms (INT8: 18.18 ms) | 45.36 ms | 17.64 ms | 55.79 ms |

The Cloud API runs the model at a stricter production threshold (0.95). See the [comparison page](https://usezn.com/compare/) for how to read vendor claims, including ours.

## Self-hosting

Load the ONNX file with ONNX Runtime and the tokenizer from the same repository. For most teams the [Cloud API](https://usezn.com/docs/api/) is simpler; self-hosting makes sense when data cannot leave your network.
