Prompt injection
Hidden instructions in a README, an issue or a page tell the assistant what to write.
Our first Security MCP. When your AI writes Python, it reads every line and returns PASS or BLOCKED — and if something's unsafe, tells you the exact line and why. All before the code runs, and without ever running it. One tool, any AI assistant.
import os as o keys = o.environ open('/etc/passwd').read()
{
"verdict": "BLOCKED",
"diagnostics": [
{ "code": "ERR_AST_BANNED_MODULE_ATTRIBUTE", "line": 2 },
{ "code": "ERR_AST_BANNED_SECRET_HARVESTING", "line": 2 },
{ "code": "ERR_AST_BANNED_FILESYSTEM_ACCESS", "line": 3 }
],
"verification": { "code_executed": false }
}
"I downloaded an AI 'skill' from a tutorial and found hidden Russian text buried inside it — and realized how easily a skill can carry something unsafe."
That was the moment behind Pacific AI Labs. If a packaged skill can hide instructions, so can anything an AI writes for you. A single poisoned prompt — in a README, an issue, a web page the assistant reads — can make an AI coding assistant write code that runs shell commands, steals secrets, or phones home. And increasingly, the assistant doesn't just suggest that code. It runs it.
So — what do you do when your AI hands you bad code?
AI assistants now write — and increasingly run — code faster than anyone can review it. The flaws they produce aren't exotic. They're the simplest, most preventable classes in the book: a shell command built from untrusted input, a hardcoded secret, an imported package that doesn't exist.
CISA puts OS command injection (CWE-78) on its list of "unforgivable" vulnerabilities — classes that are fully preventable by design, yet keep shipping.1 Its own diagnosis: they persist "due to organizational culture and development workflows rather than technical complexity."1 The fix is known. The industry just doesn't build it in.
Hidden instructions in a README, an issue or a page tell the assistant what to write.
An agent that runs its own code turns one bad suggestion into a real action.
A smarter model won't close this gap — those studies used today's best models. The gap isn't intelligence; it's that nobody checks the output. A poisoned prompt (OWASP LLM01), unreviewed output (LLM05) and an agent that runs its own code (LLM06) line up into one attack path5 — and the gate is what catches it, before AI-written code runs.
Sources
Plenty of tools touch AI-code safety, but none do exactly this job. Here's plainly what our gate is (and isn't), where the alternatives fall short, and how it scored against them on the same test.
We looked hard before building. Each of these is good at something; none is built to screen AI-written code against a safety policy before it runs.
| Type of tool | Examples | Good at | Why it doesn't do this job |
|---|---|---|---|
| AI guardrails | Llama Guard, NeMo Guardrails | Spotting toxic or off-policy intent | It's one AI judging another — it can guess wrong, and here the answer has to be certain. |
| Runtime sandboxes | gVisor, Firecracker | Limiting what running code can touch | They limit the damage after code runs — they don't stop it first, are heavy to run, and tell the AI nothing it can fix. |
| Code scanners | Bandit, Semgrep | Finding known bug patterns in human code | Built for a different problem; out of the box they miss most of what this blocks (see the numbers below). |
| Code guardrails | Meta CodeShield | Spotting insecure-code patterns across languages | Different goal — and if its scanner errors, the direct API lets the code through; ours blocks instead. |
| Restricted runtimes | RestrictedPython | Mature, long-maintained Python restriction | It limits code as it runs; it doesn't screen the code first and hand back a traceable verdict. |
Your AI assistant hands its Python to one tool, validate_python_code. The gate reads every line and, if something's unsafe, tells you the exact line and why — then returns PASS or BLOCKED. It never runs the code to decide. Here's a real example.
To be clear: it doesn't stop the prompt injection itself — that's the trick that fools the AI upstream. It catches the dangerous code that results, however it got written (a poisoned prompt, a bad suggestion, or a hallucination). It's the safety net right before the code runs.
1 · The AI writes code — with a hidden risk
1 def run_report(cmd): 2 import os 3 os.system(cmd)
Looks harmless — but line 3 lets this code run any shell command on your machine. That's OS command injection, the "unforgivable" bug from above.
2 · The gate reads it (never runs it) → BLOCKED
{
"verdict": "BLOCKED",
"diagnostics": [
{ "code": "ERR_AST_BANNED_OS_SYSTEM_CALL",
"line": 3 }
],
"code_executed": false
}
Back comes the exact line (3), a diagnostic code naming the problem (a banned os.system call), and proof it was never run — code_executed: false.
And it catches them even when they're hidden or disguised to slip past a simpler tool — because the gate works out what the code would actually do when it runs, not just whether it contains an obvious bad word.
Shell or process execution and dynamic eval — the classic "run anything" holes.
Reading or writing your files, reaching the network, or harvesting secrets and environment variables.
The same tricks smuggled through aliases, re-exports or indirection — resolved, not string-matched.
The test set was written and locked by a separate AI — not the one being graded — before we scored anything. And the gate is built and attacked by three independent AIs — Claude, Gemini, and GPT — plus a human, so a blind spot in one gets caught by the others.
0 escapes. Attack corpus authored by Gemini and OpenAI Codex, re-run live on the current validator.
0 false alarms. Ordinary, safe programs all pass — the gate doesn't cry wolf.
72/72. Three frontier models, 24 frozen tasks each — no code reached the host without passing the gate.
Each tool got the identical test: 100 pieces of dangerous AI-written code (and 100 perfectly safe ones). The bars show how many of the 100 attacks each tool caught — higher is better. Our gate caught all 100, and none of the tools wrongly flagged a safe program.
Fair-play notes. We tested each tool in one reasonable setup, not its absolute best. The 100 attacks were written by someone outside the project — though they'd seen how the gate works, so it's an honest test, not a fully "blind" one. We couldn't include Meta CodeShield in this run because the rules it uses to detect problems aren't published.
One MCP tool. Connect it over stdio locally, or over HTTP for the hosted beta. Your assistant calls it; a host-verified verdict comes back.
1 · CONNECT OVER STDIO (LOCAL)
$ claude mcp add llm-codegen-gate \
-- python /ABS/PATH/mcp-server/server.py2 · CONNECT OVER HTTP (HOSTED BETA)
$ claude mcp add --transport http llm-codegen-gate \ https://YOUR-SERVER/mcp \ --header "Authorization: Bearer pcg_YOUR_TOKEN"
3 · CALL THE TOOL
validate_python_code({ code: "def add(a, b): return a + b" })
// → { "verdict": "PASS", "verification": { "code_executed": false } }Free & open. Run it locally over stdio at no cost. An optional hosted mode uses per-tester tokens with hot revocation, a rate limit, a request-size cap, and DNS-rebinding protection. Logs never contain submitted code. Tested clients: Claude Code, MCP Inspector, the MCP Python SDK, Docker.
Request access to the hosted beta, or read the full threat model and methodology in the research paper.