Skip to content

Roadmap

How Palisade goes from a working v0.5.x to applied agentic-safety infrastructure a team puts in front of every PR. This sequences the work and argues why this order - then reports current status honestly against it.

The growth path: one loop, four verbs

We are building the AI safety engineer for a codebase. The complete role is the destination, and exactly one verb of it ships today: SEE, the static detector for prompt-injection-to-execution paths. JUDGE, PROBE, GUARD and WATCH are upcoming, with no dates, and the table below says which is which rather than implying the whole loop exists.

That ordering is not modesty. Three engine passes each moved held-out recall by zero paths, so a claim to cover the agent-framework shape would be a claim a first scan could embarrass. Each capability below maps onto a pillar of the long-horizon agenda, and carries its real status:

VerbStatusPalisade capabilityAgenda pillar it instantiates
SEEshipped (v1)scan / map - the action-boundary surface (input → model → exec/shell/SQL/payments/secrets)Agentic safety - the model→high-impact-action interface, which is the loss-of-control surface as autonomy scales
JUDGEupcoming - advisory today, calibration preliminaryaudit / review + posture score, grounded in verified static factsSafety cases - a structured, evidence-backed argument about a system’s safety posture
PROBEupcoming - synthesis and gated execution exist, the calibrated measurement does notredteam synthesis + gated execution + scoring, on a pinned corpusEvals - a grounded harness for a verifiable failure class
GUARDupcoming - deterministic fix ships; the guardrail and safety-case generators are not builtfix, with a regression test per findingOversight - a human at the high-impact boundary, machine doing the labor
WATCHupcoming - nothing builtopt-in self-hosted runtime SDK: monitoring, circuit-breaking, incident captureGovernance - makes safety practice enforceable as an org requirement
(cross-cutting)shippedadvisory + approval gates; SARIF, CI gates, CWE/OWASP-LLM mapping, disclosure/provenanceOversight + governance at the pipeline boundary

No dates on any upcoming row, deliberately. The next capability that unblocks SEE on the agent-framework shape is whole-program object tracking - following values rather than types - specified from documented train misses only. The cheaper abstract-dispatch shortcut is recorded as rejected for now: it would raise recall by spending the precision the shipped product rests on.

The scope is deliberate and honest: this is engineering infrastructure at the deployment layer, not frontier alignment research. Its catastrophic-risk relevance is anticipatory - the failure it hardens today is the same shape that scales as agents gain capability and autonomy. It is layered on the deterministic core, never replacing it. It is all MIT and free; the split is keyless-and-offline versus bring-your-own-endpoint.

Shipped:

  • map - offline inventory of the AI surface (LLM calls, prompts, tools, agents, retrieval, dangerous flags).
  • Judgment layer over any OpenAI-compatible endpoint (TypeSafe by default), configured in .env, behind one JudgeBackend interface. An unverified backend is best-effort and can never BLOCK or raise a Critical posture on judgment alone.
  • audit - excessive-agency and taint-path exploitability checks, each grounded in a verified static fact.
  • review + posture score - one composed, risk-ranked report (a number and a band over detected findings).
  • Multi-agent detection - the agent graph in map (agents, tools, capabilities, handoffs) across OpenAI Agents SDK, LangGraph, and CrewAI, plus the deterministic PI-AGENT-HANDOFF finding in scan (untrusted input -> agent run -> handoff -> a dangerous-capability agent). Measured on labelled fixtures and gated in CI.
  • Red-team synthesis and gated execution - a Map-driven adversarial attack suite (advisory, offline) that can be fired at a user-provided target with --approve, scored by the judgment backend.
  • Judgment calibration harness (scripts/calibrate.py + a labelled corpus). Preliminary: measured on a 10-case seed corpus (n=4 to 6 per signal), not a benchmark result; the judged layer stays advisory (see corpus/judgment/RESULTS.md). gated is marked known-weak on that seed (the model over-predicts gating) - reported, not trusted to downgrade a finding.
  • SARIF output + GitHub code-scanning Action - scan --sarif emits SARIF 2.1.0 (severity mapped, line-shift-resilient fingerprints); a five-line workflow uploads findings to the Security tab, dogfooded on our own src.

Upcoming (no dates):

  • Guardrail generation - installable guardrail middleware plus a regression test per confirmed finding.
  • Grow the judgment corpus to real repos and calibrate gated; the seed gate is fixture-scale. The deterministic detections (taint rules, PI-AGENT-HANDOFF) already carry a measured, CI-gated precision.
  • Multi-agent recall: crew.kickoff / compiled-LangGraph .invoke entry mapping, conditional edges, and cross-module agent wiring.

The ordering thesis

One rule drives the sequence: at each stage, relieve the single binding constraint preventing the tool from getting or keeping users. For a free, open-source security tool, three facts fix the order:

  1. Trust is scarce and asymmetric. A noisy or crashing scanner gets uninstalled and badmouthed once, forever. → Quality must be measurable before reach.
  2. Adoption is the improvement engine. Real repos are the test corpus; real FP reports are the precision tuning. → Distribute early - but only after you won’t embarrass yourself.
  3. Coverage gates per-user value; expansion dilutes focus. Cover the beachhead deeply before widening it.

Through-line: Measure → Distribute → Cover → Scale → Certify → Expand → Remediate. Trust before reach before depth.

Where we are (v0.5.1)

Phase 0 is complete. The engine, six rules (five taint rules plus PI-AGENT-HANDOFF), both frontends, library mode, the baseline/CI flow and a template-based fix are shipped and pinned by the test suite. 0.5.0, with the judgment tier (audit, review, redteam --execute), is on PyPI; 0.5.1 is the next release. Quality is now measured rather than asserted, against a pinned benchmark corpus of 26 third-party repos (40,466 files Palisade actually scans, 50 repos):

MetricValue
Precision1.000 (tp=2, fp=0)
Recall (train)0.235 (tp=4, fn=13; 17 documented paths)
Recall (held-out)0.000 (tp=0, fn=98; 0 of 45 independent observations)
F10.333

Zero false positives across 40,466 files of real third-party code. Three repos (84 files) contain no untrusted input for taint to start from; they are reported but excluded from the precision claim.

Recall is measured against 10 real paths, each hand-verified at the pinned commit: the Vanna CVE in two releases (found), PandasAI’s CVE-2024-12366 (missed: dynamically dispatched pipeline steps), and 7 paths found by a 2026-09-22 audit of the clean repos (all missed). Of the 8 misses, 4 run in a sandbox by default (autogen, dspy) and 3 reach raw SQL or a shell directly (crewai-tools, griptape, the Anthropic SDK’s bash tool). Every miss stays labelled, so recall stays honest and the gaps stay visible. They point at three engine capabilities, now the top of Phase 2: tool-call arguments as model output, more LLM call shapes, and method calls on objects the engine cannot resolve. Full detail in proof-scans.md.

PhaseThemeStatus
0MeasureDone. P/R published and gated in CI; corpus pinned with recorded SHAs; FP regression harness live; inline suppressions shipped; self-security enforced over an adversarial corpus
1DistributeIn progress - SARIF output + a code-scanning GitHub Action shipped (dogfooded on our own src); pre-commit and a Marketplace action remain. The public-launch gate lives here
2CoverPartial - notebooks, framework breadth, rule-test framework for community PRs
3ScaleOpen - incremental scanning, caching, perf gates
4CertifyStarted - SECURITY.md and release discipline shipped; signing, SBOM, provenance remain
5Expandv1 pulled forward (JS/TS frontend shipped as the architecture proof); the JS benchmark corpus + precision gate remain
6Remediatev1 pulled forward (deterministic template fix); the LLM-assisted, eval-gated diff engine remains

Two items originally sequenced late were deliberately pulled forward in v0.3 with reduced scope, noted in their phases below.

The phases

Phase 0 - MEASURE

Constraint: “We can’t tell if the tool is good, and can’t change it without silently breaking it.”

  • Benchmark corpus: CVE repos pinned at vulnerable and patched commits (PandasAI, Vanna, Langflow, LangChain PAL/LLMMath) + ~25 clean popular Python AI repos (the honesty half - guards against overfitting to the CVE set). Ground-truth labels per repo: file:line → rule → should-flag / should-be-silent.
  • Precision/Recall/F1 harness in CI that fails the build if precision drops below threshold (~90% to start). The published number is a byproduct; the regression gate is the point. (Done: every PR is gated on the fast fixture corpus, and the pinned 26-repo third-party corpus is scored weekly and on demand. The harness now also fails when it measures nothing, after an unlabelled run reported precision=1.000 on tp=0 fp=0 fn=0.)
  • FP regression harness: every reported false positive becomes a permanent must-stay-silent fixture. (Already practiced informally - the test suite grew exactly this way - needs formalizing against the corpus.)
  • Inline suppressions: # palisade: ignore[PI-EXEC] - reviewed, sandboxed, tracked and surfaced in reports. (Done: # and // forms, on the sink line or the one above; suppressed findings are counted, carry their reason into --json, and a comment that stops matching anything is reported as stale. A silent suppression is how a vulnerability quietly returns.)
  • Self-safety guarantees: never-execute and never-crash assertions. (Shipped: tests/fixtures/hostile/ adversarial corpus driven by a sys.addaudithook tripwire - no exec/import of target code, no subprocess, no sockets; plus resource caps and skip-with-warning guarantees. See HARDENING-AUDIT.md.)

Done when: published P/R on N repos; precision gate live in CI; suppressions shipped. Trap: overfitting to the four CVE repos.

Phase 1 - DISTRIBUTE

Constraint: “Even people who like it can’t get it into their workflow.” Directly serves the North Star metric: repos running Palisade in CI.

  • SARIF output (shipped) - the single highest-leverage feature; the standard interface every AppSec pipeline speaks. scan --sarif emits SARIF 2.1.0.
  • GitHub Code Scanning integration (shipped) - a five-line workflow uploads findings to the Security tab, dogfooded on our own src.
  • Published GitHub Action (pinned) on the Marketplace; pre-commit hook; CI recipes for GitLab/CircleCI/Jenkins/Azure.
  • Versioned --json schema (shipped: schema_version: 1, documented); baseline UX polish (--update-baseline, stale reporting).

Done when: a 5-line workflow puts findings in the GitHub Security tab. Trap: unpinned deps in a security tool’s own action.

▶ PUBLIC LAUNCH GATE (fires once, here). Only now do you have a precision number and copy-paste CI integration. Spend the one-time Show HN / Reddit attention spike here - not before. During Phases 0–1, run a private ~5-repo design-partner beta to harvest real FP data.

Phase 2 - COVER

Constraint: “It doesn’t understand my framework / file type / Python version.” Post-launch churn comes from coverage gaps; this phase also opens the community-rule flywheel - the moat.

First, the measured recall gaps (train recall 0.235 over 17 documented paths). These gaps are specified from the train half on purpose: its misses are already explained in docs/proof-scans.md, so building against them costs nothing that was not already spent. The held-out half - 98 paths, 45 independent observations, recall 0.000 - is what measures whether closing them generalizes, so nothing below may be specified from a held-out miss. See corpus/RECALL-PROTOCOL.md.

Read the next measurement on the right axis

Held-out recall is 0.000 over 45 independent observations, and that number says something narrower than “recall is low”. It says the engine models one shape

  • untrusted text reaching a project wrapper that reaches exec - and the field has moved to another: an agent framework whose own code hands model-chosen arguments to a tool it ships.

So the fork has to be named before the fix, not after:

After both gap closures, held-out goes…What it meansWhat follows
0/45 -> a large fractionthe documented gaps were the blockers, and they generalizethe engine works; calibrate and ship
0/45 -> 1-3/45the gaps were real misses and closing them was correct, but they are not what held-out is made ofthe agent-framework shape needs its own rule family, which is a larger piece of work than two gap closures
0/45 -> 0/45same conclusion, stated more sharplyas above

The middle row is the outcome to expect, and it is not a failed fix. Two train gaps closing and held-out barely moving is evidence about what held-out contains, which is exactly what a held-out set is for. Reading it as pass/fail on the gap work would be reading it on the wrong axis - and would create pressure to go looking at held-out misses for the next spec, which is the one thing that would destroy the measurement.

Measured 2026-09-30: the bottom row, not the middle one

Both gaps closed. Train recall rose 0.133 -> 0.235. Held-out moved 0/45 to 0/45 - not a single path, under any of the three denominators.

So the answer is the sharpest available version of the fork: the shapes the documented train gaps describe are not what held-out is made of. The next piece of work is a rule family for the agent-framework shape - a framework’s own code handing model-chosen arguments to a tool it ships, across an object boundary - and it is a larger piece of work than any number of gap closures. Two pieces of evidence point the same way:

  • griptape sql_driver.py:37 is still missed after gap 1, because the tool and the sink are separated by self.sql_driver.execute_query(query). That is the third documented gap - method calls on objects - and it is the shape, not an edge case.
  • dspy’s two misses need the same thing: self.code_generate(...) is only an LLM call because code_generate was assigned dspy.ChainOfThought(...) in __init__. No call signature can see that.

Attribute-type resolution across object boundaries is therefore the prerequisite for the whole class, and it should be specified and built as such rather than as a third gap in a list.

What the pass cost, recorded because it is the part that generalizes. Gap 1 also produced the project’s first false positive in the corpus’s history - precision 1.000 -> 0.750, gate failed - on crewai singlestore_search_tool.py. Triage showed the engine was right and the corpus was incomplete: an unlabelled sibling of an existing label, same shape, same directory. Two labels were added after reading the family by hand, one of which the engine finds and one of which it does not, so the correction cost a miss as well as gaining a hit. Precision is back to 1.000 over 40,466 files.

Measured 2026-09-30: the rule family moved nothing

One run, as pre-committed. No re-run.

beforeafter
precision1.0001.000 (tp=4 fp=0, 40,466 files)
train recall0.2350.235
held-out recall0/450/45

The capability is real - nine tests, and reduced versions of every train shape resolve, including two object boundaries through a base-class template method. It changed zero outcomes across 50 real repositories. That is a result, and a more useful one than a small move would have been: object-boundary resolution is necessary and is not the binding constraint.

What the constraint actually is. Every real instance terminates in something local resolution cannot reach:

  1. Dependency-injected abstractions. griptape’s SqlLoader.sql_driver is declared BaseSqlDriver - abstract, many implementations, concrete driver chosen at a wiring site in another file. The engine refuses to guess among many providers, and that refusal is the precision guarantee. Resolving it needs the value, not the type.
  2. Value flow through constructed containers. dspy builds input_kwargs in a loop and calls self.code_generate(**input_kwargs); taint through dict construction and ** expansion is a separate mechanism from typing.

Both are whole-program problems. Neither is a rule, a signature, or a marker.

The decision taken: ship the narrow tool, make this the roadmap

Three passes, three zeros. That is a scope finding, not an engine to-do, and it settles the launch rather than the next sprint.

The engine is not touched before launch. Precision 1.000 is the entire trust claim of what ships, and the abstract-dispatch shortcut buys recall on a shape we have now measured we do not catch by spending exactly that. An experimental path whose precision is load-bearing for a live product is not an experiment.

The next capability, named honestly: whole-program object tracking. Not the shortcut. The traced constraints are dependency-injected abstractions (the declared type is abstract, the concrete object is chosen at a wiring site in another file) and value flow through constructed containers (**kwargs built in a loop). Both need to follow values, not types. Specified from the four train misses already written down above, no dates.

When it runs, it runs as post-launch research: an experimental branch, measured against the held-out half with precision as a gate, never as a live-product guarantee. And under the same discipline as every pass so far - one specification from train misses, one implementation, one measurement. A search over relaxations, checking held-out between each, is fitting to the test set on a longer feedback loop, and it would be invisible in the final number.

The abstract-dispatch relaxation is recorded as rejected-for-now rather than deleted: it is cheap and it would raise recall, and the reason not to do it is that it trades the one property the shipped product rests on.

The gaps, each specified from a train miss

Written before the code, by coordinate and verbatim sink, so the derivation is auditable and provably contains no held-out path.

Gap 1 - tool-call arguments are not modeled as model output. Diagnosed in docs/proof-scans.md for crewai-tools and autogen. Train misses it targets:

train pathverbatim sink
crewai lib/crewai-tools/…/snowflake_search_tool.py:216cursor.execute(query, timeout=timeout)
autogen …/autogen-ext/…/code_execution/_code_execution.py:83result = await self._executor.execute_code_blocks(
griptape griptape/drivers/sql/sql_driver.py:37results = con.execute(sqlalchemy.text(query))

Fix derived from those three alone: a function registered as an agent tool has model-written parameters, so its parameters are an LLM-output source, not a plain untrusted source. Registration is recognised from the decorators and base classes those three repos use.

Gap 2 - LLM call shapes that are not recognised as LLM calls. Diagnosed for autogen (model_client.create), dspy (module calls) and griptape (prompt_driver.run). Train misses it targets:

train pathverbatim sink
dspy dspy/predict/program_of_thought.py:188result = interpreter.execute(code)
dspy dspy/predict/rlm.py:707return repl.execute(code, variables=dict(input_args))
autogen …/autogen-agentchat/…/agents/_code_executor_agent.py:718result = await self._code_executor.execute_code_blocks(code_blocks

Fix derived from those three alone: add the wrapper signatures those repos call to the recognised LLM-call set, so the source -> LLM -> sink chain becomes visible when the SDK call is not syntactically present.

Measurement discipline for this pass: close both gaps, then measure once. Measuring after each fix and stopping when the number moves is a soft form of fitting, and it is the version of fitting that feels like diligence.

Next: a rule family for the agent-framework shape

Not a gap closure. The measurement on 2026-09-30 showed the two documented gaps cannot move held-out at all, because they do not describe its shape. This does.

The capability being built: resolve the type of an attribute or a local through what was assigned to it, so that a call on that object can be recognised as an LLM call or as a dangerous sink even though neither is visible at the call site. Both halves of the agent-framework shape need it.

Specified from train misses only, by coordinate and verbatim sink

Written before any code, as with the last pass. These are the train paths that need attribute-type resolution and nothing else was consulted.

train pathverbatim sinkwhat the resolution has to see
griptape griptape/drivers/sql/sql_driver.py:37results = con.execute(sqlalchemy.text(query))the tool calls self.sql_driver.execute_query(query); the sink is in the driver, one object boundary away
dspy dspy/predict/program_of_thought.py:188result = interpreter.execute(code)code came from self.code_generate(...), an attribute assigned dspy.ChainOfThought(...) in __init__ - the LLM call is invisible without resolving the attribute
dspy dspy/predict/rlm.py:707return repl.execute(code, variables=dict(input_args))same: self.generate_action / self.extract are dspy.Predict(...) instances
anthropic-sdk src/anthropic/lib/tools/agent_toolset.py:500await stdin.send(wrapped.encode())the sink is a write to a held subprocess’s stdin, reached through a registry-dispatched tool

Those four are the specification. If building it requires looking at a held-out miss to know what to do, stop - that is the signal the train misses are not a sufficient spec, and continuing means fitting to the test set. The held-out set stays closed until the single post-build measurement.

This wall matters more on this pass than the last one. The held-out half is full of this shape and I know it, so the pull toward “just check what they have in common” is far stronger than it was for two narrow gaps. Knowing the shape is common is not a licence to read the instances.

Pre-committed: one measurement, whatever it says

The two gaps could not move held-out. This family describes held-out’s shape, so if it works the number may jump from 0/45 to something substantial in a single run. That is the good outcome and it is also the one most likely to invite a re-run “to check”, then a tweak, then another run.

One measurement. Reported with all three denominators and its upper bound. Then stop. A large favourable move measured once is a result; the same move arrived at by tuning against held-out is fitting with extra steps, and it would be indistinguishable from the real thing in the final number.

Recorded in advance so the commitment predates the temptation.

  • Method calls on objects. Resolve obj.method(x) and self.attr.method(x) through constructor and attribute types (abstract executors, SQL drivers), and model code-execution sinks reached that way (execute_code_blocks, interpreter.execute, writes to a shell’s stdin).

  • Python syntax matrix (3.11–3.13+: match, walrus, type-params) in CI.

  • Jupyter notebook support (.ipynb cells → IR) - a large share of AI code lives in notebooks.

  • Framework-rule breadth, pure data: Django ORM/raw, SQLAlchemy text(), Starlette; LangGraph, CrewAI, AutoGen, LlamaIndex, Haystack, DSPy; template sinks; Bedrock/Vertex/Groq providers.

  • Robustness fuzz against the top ~1000 PyPI AI repos.

  • Rule-test framework (pulled forward from “ecosystem”): community rule PRs cannot land without must-flag + must-be-silent fixtures and a green precision gate. This is what makes community contribution safe.

Done when: matrix green, notebooks scanned, first external rule PRs merged. Trap: coverage sprawl buying recall with false positives.

Phase 3 - SCALE

Constraint: “Too slow or noisy for a large repo / busy CI.” Only bites once Phases 1–2 produce adopters with big repos - optimizing earlier is premature (current baseline: 1,576 files of Langflow in ~9 s on an Apple M3 Pro).

  • Incremental / diff-aware scanning (changed files + their taint neighborhood) - correctness tested against full scans.
  • Content-hash caching; parallel processing with deterministic output.
  • Per-file timeouts + global budget; honest truncation reporting (partially shipped: truncation notes exist).
  • Published perf benchmark on a 500k+ LOC repo with a regression gate.
  • Rules-ecosystem maturity: versioned rule packs, rule-severity SemVer, deprecation policy.

Phase 4 - CERTIFY

Constraint: “Serious security orgs won’t run an unsigned, opaque tool.”

  • Sigstore-signed releases, SLSA provenance, CycloneDX SBOM, reproducible builds, minimal-dependency audit.
  • SECURITY.md + coordinated disclosure (you will receive vuln reports). (Shipped.)
  • Telemetry: off by default or not at all - source never leaves the machine. Default-on telemetry is self-sabotage for a security tool.
  • Release discipline: SemVer, changelog (shipped), deprecation/LTS policy.
  • Optional compliance mapping: CWE + OWASP LLM Top-10 tags per rule.

Phase 5 - EXPAND (v1 pulled forward - deliberately, with scope control)

Original sequencing put JS/TS after the Python wedge was won, and that logic stands for investment. What shipped early in v0.3 is the architecture proof: a tree-sitter JS/TS frontend emitting the same IR with zero engine changes, Express/Node sources and sinks, behind an optional [js] extra so the core stays lean. Remaining Phase 5 work before JS findings deserve equal trust:

  • JS benchmark corpus + its own precision gate (mirror Phase 0).
  • JS/TS rule-pack depth: Next.js request surfaces, innerHTML / dangerouslySetInnerHTML, Vercel AI SDK, LangChain.js.
  • npx distribution wrapper for first-class JS DX.

Reorder trigger honored: the frontend exists; deeper JS investment waits for JS demand or JS-side CVEs.

Phase 6 - REMEDIATE (v1 pulled forward - the safe subset only)

The original vision - and correctly sequenced last for its risky half. What shipped early is the deterministic, offline subset: palisade-sec fix emits templated guardrails (allowlist / argv / SELECT-validator / SSRF-guard) each paired with a regression pytest, and never touches code. The deferred half remains deferred until the scanner’s trust base (Phases 0–1) is formal:

  • LLM-assisted diffs on the user’s own model key, provider-agnostic.
  • Eval-gated patches: generate the guardrail diff and the pytest, run in a sandbox, propose the diff only if the test passes.
  • Human-review-first: PR-style diffs, never auto-apply.
  • Success metric: accept/merge rate on real findings.

Dependency graph

Phase 0 (Measure) ──▶ Phase 1 (Distribute) ──▶ [PUBLIC LAUNCH]
│ ├──▶ Phase 2 (Cover) ┐ interleave, driven by
│ ├──▶ Phase 3 (Scale) ┘ what real users hit
│ └──▶ Phase 4 (Certify) - pull forward if an
│ enterprise partner appears
└── correctness guarantees live HERE, not in Phase 4
Phase 5 (Expand) / Phase 6 (Remediate): v1s shipped as proofs; their
deep halves stay gated on Python adoption being real (metrics, not calendar).

Reorder triggers: enterprise partner → pull Phase 4 signing forward; JS-side CVE wave → deepen Phase 5; any in-the-wild crasher or FP class → jumps the queue into Phase 0’s harness immediately.

Sequencing traps (explicit anti-goals)

  • Building the LLM fix engine before a formally measured scanner.
  • Public launch before SARIF + a precision number - wastes the one-time attention spike.
  • Buying recall with false positives - always net-negative for a security tool.
  • Default-on telemetry - instantly off-brand.
  • Premature performance work before big-repo users exist.

Metrics ladder

PhasePrimary metric
0 MeasurePublished precision/recall; regression gate live
1 Distribute# repos running Palisade in CI (North Star)
2 Cover# frameworks covered; # external rule PRs merged
3 ScaleScan time on 500k LOC; % PRs on incremental mode
4 CertifySigned-release adoption; enterprise evals reaching pilot
5 ExpandJS precision gate green; # JS/TS repos scanned
6 Remediatefix accept/merge rate on real findings