What's shipped, and what's next.

A safety engineer sees the system, judges it, probes it, guards it, and watches it in production. One of those is shipped. Everything else below says upcoming and carries no date.

Shipped now

v1 — the detector

A linter for one real bug class: untrusted input reaching a model whose output reaches a dangerous capability.

SeeShipped

Find where a model touches a capability

Static taint from untrusted input, through the model, to exec, a shell, raw SQL, a model-chosen URL, or a handoff into a dangerous tool. All three hops, or nothing.

scan · map · baseline · fix · SARIF · GitHub Action · pre-commit
Offline by defaultShipped

No key, no account, no network

The core reads your code and never runs it. A test in the suite fails the build if a scan so much as imports the networking module.

python · javascript / typescript · MIT · no paid tier
Measured

Both numbers, always

A scanner that reports nothing has perfect precision for free, so neither figure is published without the other.

1.000
Precision · 0 false positives · 40,466 files · 50 repos
0.000
Held-out recall · 0 of 45 independent observations · 20 repos

Our ground truth is known-incomplete: one of our own audits missed two real paths the engine later found. So the denominators are smaller than the truth, and that recall figure is, if anything, an overestimate. Every miss is listed by file and line in the proof-scan report.

The boundary

It does not yet catch tool-calling agent frameworks

The shape the field is moving to. If we are quiet on your agent framework, that is a known limit, not a clean bill of health.

Why

These paths end where types stop helping

The capability is injected at a wiring site in another file, or the arguments are assembled in a container and splatted into the call. Three rule-shaped attempts each moved held-out recall by zero paths, which is how we know the answer is not another rule.

next capability · whole-program object tracking · follow values, not types
Planned next

The rest of the role

Nothing here is finished, and nothing here has a date.

JudgeUpcoming

Decide what a finding is worth

Ships today as audit and review, and deliberately not counted as done: advisory, on a preliminary 10-case seed corpus. It does not decide whether you are safe.

needs per-check calibration on a full benchmark corpus
ProbeUpcoming

Attack the system with its own map

Red-team synthesis and gated execution exist. The reproducible measurement that would make the results quotable does not.

attack benchmark · landed-attack rates worth publishing
GuardUpcoming

Turn a finding into a control

Deterministic fixes ship now, each with a regression test. The guardrail and safety-case generators are not built.

guardrail generator · safety-case export
WatchUpcoming

Notice it happening for real

Opt-in, self-hosted runtime monitoring and incident capture, feeding back into a new static check. Nothing here is built.

not built