Roadmap
How Palisade goes from a working v0.5.x to applied agentic-safety infrastructure a team puts in front of every PR. This sequences the work and argues why this order - then reports current status honestly against it.
The growth path: one loop, four verbs
We are building the AI safety engineer for a codebase. The complete role is the destination, and exactly one verb of it ships today: SEE, the static detector for prompt-injection-to-execution paths. JUDGE, PROBE, GUARD and WATCH are upcoming, with no dates, and the table below says which is which rather than implying the whole loop exists.
That ordering is not modesty. Three engine passes each moved held-out recall by zero paths, so a claim to cover the agent-framework shape would be a claim a first scan could embarrass. Each capability below maps onto a pillar of the long-horizon agenda, and carries its real status:
| Verb | Status | Palisade capability | Agenda pillar it instantiates |
|---|---|---|---|
| SEE | shipped (v1) | scan / map - the action-boundary surface (input → model → exec/shell/SQL/payments/secrets) | Agentic safety - the model→high-impact-action interface, which is the loss-of-control surface as autonomy scales |
| JUDGE | upcoming - advisory today, calibration preliminary | audit / review + posture score, grounded in verified static facts | Safety cases - a structured, evidence-backed argument about a system’s safety posture |
| PROBE | upcoming - synthesis and gated execution exist, the calibrated measurement does not | redteam synthesis + gated execution + scoring, on a pinned corpus | Evals - a grounded harness for a verifiable failure class |
| GUARD | upcoming - deterministic fix ships; the guardrail and safety-case generators are not built | fix, with a regression test per finding | Oversight - a human at the high-impact boundary, machine doing the labor |
| WATCH | upcoming - nothing built | opt-in self-hosted runtime SDK: monitoring, circuit-breaking, incident capture | Governance - makes safety practice enforceable as an org requirement |
| (cross-cutting) | shipped | advisory + approval gates; SARIF, CI gates, CWE/OWASP-LLM mapping, disclosure/provenance | Oversight + governance at the pipeline boundary |
No dates on any upcoming row, deliberately. The next capability that unblocks SEE on the agent-framework shape is whole-program object tracking - following values rather than types - specified from documented train misses only. The cheaper abstract-dispatch shortcut is recorded as rejected for now: it would raise recall by spending the precision the shipped product rests on.
The scope is deliberate and honest: this is engineering infrastructure at the deployment layer, not frontier alignment research. Its catastrophic-risk relevance is anticipatory - the failure it hardens today is the same shape that scales as agents gain capability and autonomy. It is layered on the deterministic core, never replacing it. It is all MIT and free; the split is keyless-and-offline versus bring-your-own-endpoint.
Shipped:
map- offline inventory of the AI surface (LLM calls, prompts, tools, agents, retrieval, dangerous flags).- Judgment layer over any OpenAI-compatible endpoint (TypeSafe by default),
configured in
.env, behind oneJudgeBackendinterface. An unverified backend is best-effort and can never BLOCK or raise a Critical posture on judgment alone. audit- excessive-agency and taint-path exploitability checks, each grounded in a verified static fact.review+ posture score - one composed, risk-ranked report (a number and a band over detected findings).- Multi-agent detection - the agent graph in
map(agents, tools, capabilities, handoffs) across OpenAI Agents SDK, LangGraph, and CrewAI, plus the deterministicPI-AGENT-HANDOFFfinding inscan(untrusted input -> agent run -> handoff -> a dangerous-capability agent). Measured on labelled fixtures and gated in CI. - Red-team synthesis and gated execution - a Map-driven adversarial attack
suite (advisory, offline) that can be fired at a user-provided target with
--approve, scored by the judgment backend. - Judgment calibration harness (
scripts/calibrate.py+ a labelled corpus). Preliminary: measured on a 10-case seed corpus (n=4 to 6 per signal), not a benchmark result; the judged layer stays advisory (seecorpus/judgment/RESULTS.md).gatedis marked known-weak on that seed (the model over-predicts gating) - reported, not trusted to downgrade a finding. - SARIF output + GitHub code-scanning Action -
scan --sarifemits SARIF 2.1.0 (severity mapped, line-shift-resilient fingerprints); a five-line workflow uploads findings to the Security tab, dogfooded on our ownsrc.
Upcoming (no dates):
- Guardrail generation - installable guardrail middleware plus a regression test per confirmed finding.
- Grow the judgment corpus to real repos and calibrate
gated; the seed gate is fixture-scale. The deterministic detections (taint rules,PI-AGENT-HANDOFF) already carry a measured, CI-gated precision. - Multi-agent recall:
crew.kickoff/ compiled-LangGraph.invokeentry mapping, conditional edges, and cross-module agent wiring.
The ordering thesis
One rule drives the sequence: at each stage, relieve the single binding constraint preventing the tool from getting or keeping users. For a free, open-source security tool, three facts fix the order:
- Trust is scarce and asymmetric. A noisy or crashing scanner gets uninstalled and badmouthed once, forever. → Quality must be measurable before reach.
- Adoption is the improvement engine. Real repos are the test corpus; real FP reports are the precision tuning. → Distribute early - but only after you won’t embarrass yourself.
- Coverage gates per-user value; expansion dilutes focus. Cover the beachhead deeply before widening it.
Through-line: Measure → Distribute → Cover → Scale → Certify → Expand → Remediate. Trust before reach before depth.
Where we are (v0.5.1)
Phase 0 is complete. The engine, six rules (five taint rules plus
PI-AGENT-HANDOFF), both frontends, library mode, the baseline/CI flow and a
template-based fix are shipped and pinned by the test suite. 0.5.0, with the
judgment tier (audit, review, redteam --execute), is on PyPI; 0.5.1 is
the next release. Quality is now measured rather than
asserted, against a pinned benchmark corpus of 26 third-party repos
(40,466 files Palisade actually scans, 50 repos):
| Metric | Value |
|---|---|
| Precision | 1.000 (tp=2, fp=0) |
| Recall (train) | 0.235 (tp=4, fn=13; 17 documented paths) |
| Recall (held-out) | 0.000 (tp=0, fn=98; 0 of 45 independent observations) |
| F1 | 0.333 |
Zero false positives across 40,466 files of real third-party code. Three repos (84 files) contain no untrusted input for taint to start from; they are reported but excluded from the precision claim.
Recall is measured against 10 real paths, each hand-verified at the pinned commit: the Vanna CVE in two releases (found), PandasAI’s CVE-2024-12366 (missed: dynamically dispatched pipeline steps), and 7 paths found by a 2026-09-22 audit of the clean repos (all missed). Of the 8 misses, 4 run in a sandbox by default (autogen, dspy) and 3 reach raw SQL or a shell directly (crewai-tools, griptape, the Anthropic SDK’s bash tool). Every miss stays labelled, so recall stays honest and the gaps stay visible. They point at three engine capabilities, now the top of Phase 2: tool-call arguments as model output, more LLM call shapes, and method calls on objects the engine cannot resolve. Full detail in proof-scans.md.
| Phase | Theme | Status |
|---|---|---|
| 0 | Measure | Done. P/R published and gated in CI; corpus pinned with recorded SHAs; FP regression harness live; inline suppressions shipped; self-security enforced over an adversarial corpus |
| 1 | Distribute | In progress - SARIF output + a code-scanning GitHub Action shipped (dogfooded on our own src); pre-commit and a Marketplace action remain. The public-launch gate lives here |
| 2 | Cover | Partial - notebooks, framework breadth, rule-test framework for community PRs |
| 3 | Scale | Open - incremental scanning, caching, perf gates |
| 4 | Certify | Started - SECURITY.md and release discipline shipped; signing, SBOM, provenance remain |
| 5 | Expand | v1 pulled forward (JS/TS frontend shipped as the architecture proof); the JS benchmark corpus + precision gate remain |
| 6 | Remediate | v1 pulled forward (deterministic template fix); the LLM-assisted, eval-gated diff engine remains |
Two items originally sequenced late were deliberately pulled forward in v0.3 with reduced scope, noted in their phases below.
The phases
Phase 0 - MEASURE
Constraint: “We can’t tell if the tool is good, and can’t change it without silently breaking it.”
- Benchmark corpus: CVE repos pinned at vulnerable and patched commits
(PandasAI, Vanna, Langflow, LangChain PAL/LLMMath) + ~25 clean popular
Python AI repos (the honesty half - guards against overfitting to the CVE
set). Ground-truth labels per repo:
file:line → rule → should-flag / should-be-silent. - Precision/Recall/F1 harness in CI that fails the build if precision drops below threshold (~90% to start). The published number is a byproduct; the regression gate is the point. (Done: every PR is gated on the fast fixture corpus, and the pinned 26-repo third-party corpus is scored weekly and on demand. The harness now also fails when it measures nothing, after an unlabelled run reported precision=1.000 on tp=0 fp=0 fn=0.)
- FP regression harness: every reported false positive becomes a permanent must-stay-silent fixture. (Already practiced informally - the test suite grew exactly this way - needs formalizing against the corpus.)
- Inline suppressions:
# palisade: ignore[PI-EXEC] - reviewed, sandboxed, tracked and surfaced in reports. (Done:#and//forms, on the sink line or the one above; suppressed findings are counted, carry their reason into--json, and a comment that stops matching anything is reported as stale. A silent suppression is how a vulnerability quietly returns.) - Self-safety guarantees: never-execute and never-crash assertions.
(Shipped:
tests/fixtures/hostile/adversarial corpus driven by asys.addaudithooktripwire - no exec/import of target code, no subprocess, no sockets; plus resource caps and skip-with-warning guarantees. SeeHARDENING-AUDIT.md.)
Done when: published P/R on N repos; precision gate live in CI; suppressions shipped. Trap: overfitting to the four CVE repos.
Phase 1 - DISTRIBUTE
Constraint: “Even people who like it can’t get it into their workflow.” Directly serves the North Star metric: repos running Palisade in CI.
- SARIF output (shipped) - the single highest-leverage feature; the
standard interface every AppSec pipeline speaks.
scan --sarifemits SARIF 2.1.0. - GitHub Code Scanning integration (shipped) - a five-line workflow uploads
findings to the Security tab, dogfooded on our own
src. - Published GitHub Action (pinned) on the Marketplace; pre-commit hook; CI recipes for GitLab/CircleCI/Jenkins/Azure.
- Versioned
--jsonschema (shipped:schema_version: 1, documented); baseline UX polish (--update-baseline, stale reporting).
Done when: a 5-line workflow puts findings in the GitHub Security tab. Trap: unpinned deps in a security tool’s own action.
▶ PUBLIC LAUNCH GATE (fires once, here). Only now do you have a precision number and copy-paste CI integration. Spend the one-time Show HN / Reddit attention spike here - not before. During Phases 0–1, run a private ~5-repo design-partner beta to harvest real FP data.
Phase 2 - COVER
Constraint: “It doesn’t understand my framework / file type / Python version.” Post-launch churn comes from coverage gaps; this phase also opens the community-rule flywheel - the moat.
First, the measured recall gaps (train recall 0.235 over 17 documented
paths). These gaps are specified from the train half on purpose: its misses
are already explained in docs/proof-scans.md, so building against them costs
nothing that was not already spent. The held-out half - 98 paths, 45
independent observations, recall 0.000 - is what measures whether closing
them generalizes, so nothing below may be specified from a held-out miss. See
corpus/RECALL-PROTOCOL.md.
Read the next measurement on the right axis
Held-out recall is 0.000 over 45 independent observations, and that number says something narrower than “recall is low”. It says the engine models one shape
- untrusted text reaching a project wrapper that reaches
exec- and the field has moved to another: an agent framework whose own code hands model-chosen arguments to a tool it ships.
So the fork has to be named before the fix, not after:
| After both gap closures, held-out goes… | What it means | What follows |
|---|---|---|
| 0/45 -> a large fraction | the documented gaps were the blockers, and they generalize | the engine works; calibrate and ship |
| 0/45 -> 1-3/45 | the gaps were real misses and closing them was correct, but they are not what held-out is made of | the agent-framework shape needs its own rule family, which is a larger piece of work than two gap closures |
| 0/45 -> 0/45 | same conclusion, stated more sharply | as above |
The middle row is the outcome to expect, and it is not a failed fix. Two train gaps closing and held-out barely moving is evidence about what held-out contains, which is exactly what a held-out set is for. Reading it as pass/fail on the gap work would be reading it on the wrong axis - and would create pressure to go looking at held-out misses for the next spec, which is the one thing that would destroy the measurement.
Measured 2026-09-30: the bottom row, not the middle one
Both gaps closed. Train recall rose 0.133 -> 0.235. Held-out moved 0/45 to 0/45 - not a single path, under any of the three denominators.
So the answer is the sharpest available version of the fork: the shapes the documented train gaps describe are not what held-out is made of. The next piece of work is a rule family for the agent-framework shape - a framework’s own code handing model-chosen arguments to a tool it ships, across an object boundary - and it is a larger piece of work than any number of gap closures. Two pieces of evidence point the same way:
griptapesql_driver.py:37 is still missed after gap 1, because the tool and the sink are separated byself.sql_driver.execute_query(query). That is the third documented gap - method calls on objects - and it is the shape, not an edge case.- dspy’s two misses need the same thing:
self.code_generate(...)is only an LLM call becausecode_generatewas assigneddspy.ChainOfThought(...)in__init__. No call signature can see that.
Attribute-type resolution across object boundaries is therefore the prerequisite for the whole class, and it should be specified and built as such rather than as a third gap in a list.
What the pass cost, recorded because it is the part that generalizes. Gap 1
also produced the project’s first false positive in the corpus’s history -
precision 1.000 -> 0.750, gate failed - on crewai singlestore_search_tool.py.
Triage showed the engine was right and the corpus was incomplete: an unlabelled
sibling of an existing label, same shape, same directory. Two labels were added
after reading the family by hand, one of which the engine finds and one of which
it does not, so the correction cost a miss as well as gaining a hit. Precision
is back to 1.000 over 40,466 files.
Measured 2026-09-30: the rule family moved nothing
One run, as pre-committed. No re-run.
| before | after | |
|---|---|---|
| precision | 1.000 | 1.000 (tp=4 fp=0, 40,466 files) |
| train recall | 0.235 | 0.235 |
| held-out recall | 0/45 | 0/45 |
The capability is real - nine tests, and reduced versions of every train shape resolve, including two object boundaries through a base-class template method. It changed zero outcomes across 50 real repositories. That is a result, and a more useful one than a small move would have been: object-boundary resolution is necessary and is not the binding constraint.
What the constraint actually is. Every real instance terminates in something local resolution cannot reach:
- Dependency-injected abstractions. griptape’s
SqlLoader.sql_driveris declaredBaseSqlDriver- abstract, many implementations, concrete driver chosen at a wiring site in another file. The engine refuses to guess among many providers, and that refusal is the precision guarantee. Resolving it needs the value, not the type. - Value flow through constructed containers. dspy builds
input_kwargsin a loop and callsself.code_generate(**input_kwargs); taint through dict construction and**expansion is a separate mechanism from typing.
Both are whole-program problems. Neither is a rule, a signature, or a marker.
The decision taken: ship the narrow tool, make this the roadmap
Three passes, three zeros. That is a scope finding, not an engine to-do, and it settles the launch rather than the next sprint.
The engine is not touched before launch. Precision 1.000 is the entire trust claim of what ships, and the abstract-dispatch shortcut buys recall on a shape we have now measured we do not catch by spending exactly that. An experimental path whose precision is load-bearing for a live product is not an experiment.
The next capability, named honestly: whole-program object tracking. Not the
shortcut. The traced constraints are dependency-injected abstractions (the
declared type is abstract, the concrete object is chosen at a wiring site in
another file) and value flow through constructed containers (**kwargs built in
a loop). Both need to follow values, not types. Specified from the four train
misses already written down above, no dates.
When it runs, it runs as post-launch research: an experimental branch, measured against the held-out half with precision as a gate, never as a live-product guarantee. And under the same discipline as every pass so far - one specification from train misses, one implementation, one measurement. A search over relaxations, checking held-out between each, is fitting to the test set on a longer feedback loop, and it would be invisible in the final number.
The abstract-dispatch relaxation is recorded as rejected-for-now rather than deleted: it is cheap and it would raise recall, and the reason not to do it is that it trades the one property the shipped product rests on.
The gaps, each specified from a train miss
Written before the code, by coordinate and verbatim sink, so the derivation is auditable and provably contains no held-out path.
Gap 1 - tool-call arguments are not modeled as model output. Diagnosed in
docs/proof-scans.md for crewai-tools and autogen. Train misses it targets:
| train path | verbatim sink |
|---|---|
crewai lib/crewai-tools/…/snowflake_search_tool.py:216 | cursor.execute(query, timeout=timeout) |
autogen …/autogen-ext/…/code_execution/_code_execution.py:83 | result = await self._executor.execute_code_blocks( |
griptape griptape/drivers/sql/sql_driver.py:37 | results = con.execute(sqlalchemy.text(query)) |
Fix derived from those three alone: a function registered as an agent tool has model-written parameters, so its parameters are an LLM-output source, not a plain untrusted source. Registration is recognised from the decorators and base classes those three repos use.
Gap 2 - LLM call shapes that are not recognised as LLM calls. Diagnosed for
autogen (model_client.create), dspy (module calls) and griptape
(prompt_driver.run). Train misses it targets:
| train path | verbatim sink |
|---|---|
dspy dspy/predict/program_of_thought.py:188 | result = interpreter.execute(code) |
dspy dspy/predict/rlm.py:707 | return repl.execute(code, variables=dict(input_args)) |
autogen …/autogen-agentchat/…/agents/_code_executor_agent.py:718 | result = await self._code_executor.execute_code_blocks(code_blocks |
Fix derived from those three alone: add the wrapper signatures those repos call to the recognised LLM-call set, so the source -> LLM -> sink chain becomes visible when the SDK call is not syntactically present.
Measurement discipline for this pass: close both gaps, then measure once. Measuring after each fix and stopping when the number moves is a soft form of fitting, and it is the version of fitting that feels like diligence.
Next: a rule family for the agent-framework shape
Not a gap closure. The measurement on 2026-09-30 showed the two documented gaps cannot move held-out at all, because they do not describe its shape. This does.
The capability being built: resolve the type of an attribute or a local through what was assigned to it, so that a call on that object can be recognised as an LLM call or as a dangerous sink even though neither is visible at the call site. Both halves of the agent-framework shape need it.
Specified from train misses only, by coordinate and verbatim sink
Written before any code, as with the last pass. These are the train paths that need attribute-type resolution and nothing else was consulted.
| train path | verbatim sink | what the resolution has to see |
|---|---|---|
griptape griptape/drivers/sql/sql_driver.py:37 | results = con.execute(sqlalchemy.text(query)) | the tool calls self.sql_driver.execute_query(query); the sink is in the driver, one object boundary away |
dspy dspy/predict/program_of_thought.py:188 | result = interpreter.execute(code) | code came from self.code_generate(...), an attribute assigned dspy.ChainOfThought(...) in __init__ - the LLM call is invisible without resolving the attribute |
dspy dspy/predict/rlm.py:707 | return repl.execute(code, variables=dict(input_args)) | same: self.generate_action / self.extract are dspy.Predict(...) instances |
anthropic-sdk src/anthropic/lib/tools/agent_toolset.py:500 | await stdin.send(wrapped.encode()) | the sink is a write to a held subprocess’s stdin, reached through a registry-dispatched tool |
Those four are the specification. If building it requires looking at a held-out miss to know what to do, stop - that is the signal the train misses are not a sufficient spec, and continuing means fitting to the test set. The held-out set stays closed until the single post-build measurement.
This wall matters more on this pass than the last one. The held-out half is full of this shape and I know it, so the pull toward “just check what they have in common” is far stronger than it was for two narrow gaps. Knowing the shape is common is not a licence to read the instances.
Pre-committed: one measurement, whatever it says
The two gaps could not move held-out. This family describes held-out’s shape, so if it works the number may jump from 0/45 to something substantial in a single run. That is the good outcome and it is also the one most likely to invite a re-run “to check”, then a tweak, then another run.
One measurement. Reported with all three denominators and its upper bound. Then stop. A large favourable move measured once is a result; the same move arrived at by tuning against held-out is fitting with extra steps, and it would be indistinguishable from the real thing in the final number.
Recorded in advance so the commitment predates the temptation.
-
Method calls on objects. Resolve
obj.method(x)andself.attr.method(x)through constructor and attribute types (abstract executors, SQL drivers), and model code-execution sinks reached that way (execute_code_blocks,interpreter.execute, writes to a shell’s stdin). -
Python syntax matrix (3.11–3.13+:
match, walrus, type-params) in CI. -
Jupyter notebook support (
.ipynbcells → IR) - a large share of AI code lives in notebooks. -
Framework-rule breadth, pure data: Django ORM/raw, SQLAlchemy
text(), Starlette; LangGraph, CrewAI, AutoGen, LlamaIndex, Haystack, DSPy; template sinks; Bedrock/Vertex/Groq providers. -
Robustness fuzz against the top ~1000 PyPI AI repos.
-
Rule-test framework (pulled forward from “ecosystem”): community rule PRs cannot land without must-flag + must-be-silent fixtures and a green precision gate. This is what makes community contribution safe.
Done when: matrix green, notebooks scanned, first external rule PRs merged. Trap: coverage sprawl buying recall with false positives.
Phase 3 - SCALE
Constraint: “Too slow or noisy for a large repo / busy CI.” Only bites once Phases 1–2 produce adopters with big repos - optimizing earlier is premature (current baseline: 1,576 files of Langflow in ~9 s on an Apple M3 Pro).
- Incremental / diff-aware scanning (changed files + their taint neighborhood) - correctness tested against full scans.
- Content-hash caching; parallel processing with deterministic output.
- Per-file timeouts + global budget; honest truncation reporting (partially shipped: truncation notes exist).
- Published perf benchmark on a 500k+ LOC repo with a regression gate.
- Rules-ecosystem maturity: versioned rule packs, rule-severity SemVer, deprecation policy.
Phase 4 - CERTIFY
Constraint: “Serious security orgs won’t run an unsigned, opaque tool.”
- Sigstore-signed releases, SLSA provenance, CycloneDX SBOM, reproducible builds, minimal-dependency audit.
SECURITY.md+ coordinated disclosure (you will receive vuln reports). (Shipped.)- Telemetry: off by default or not at all - source never leaves the machine. Default-on telemetry is self-sabotage for a security tool.
- Release discipline: SemVer, changelog (shipped), deprecation/LTS policy.
- Optional compliance mapping: CWE + OWASP LLM Top-10 tags per rule.
Phase 5 - EXPAND (v1 pulled forward - deliberately, with scope control)
Original sequencing put JS/TS after the Python wedge was won, and that logic
stands for investment. What shipped early in v0.3 is the architecture
proof: a tree-sitter JS/TS frontend emitting the same IR with zero engine
changes, Express/Node sources and sinks, behind an optional [js] extra so
the core stays lean. Remaining Phase 5 work before JS findings deserve equal
trust:
- JS benchmark corpus + its own precision gate (mirror Phase 0).
- JS/TS rule-pack depth: Next.js request surfaces,
innerHTML/dangerouslySetInnerHTML, Vercel AI SDK, LangChain.js. npxdistribution wrapper for first-class JS DX.
Reorder trigger honored: the frontend exists; deeper JS investment waits for JS demand or JS-side CVEs.
Phase 6 - REMEDIATE (v1 pulled forward - the safe subset only)
The original vision - and correctly sequenced last for its risky half.
What shipped early is the deterministic, offline subset:
palisade-sec fix emits templated guardrails (allowlist / argv /
SELECT-validator / SSRF-guard) each paired with a regression pytest, and
never touches code. The deferred half remains deferred until the scanner’s
trust base (Phases 0–1) is formal:
- LLM-assisted diffs on the user’s own model key, provider-agnostic.
- Eval-gated patches: generate the guardrail diff and the pytest, run in a sandbox, propose the diff only if the test passes.
- Human-review-first: PR-style diffs, never auto-apply.
- Success metric: accept/merge rate on real findings.
Dependency graph
Phase 0 (Measure) ──▶ Phase 1 (Distribute) ──▶ [PUBLIC LAUNCH] │ ├──▶ Phase 2 (Cover) ┐ interleave, driven by │ ├──▶ Phase 3 (Scale) ┘ what real users hit │ └──▶ Phase 4 (Certify) - pull forward if an │ enterprise partner appears └── correctness guarantees live HERE, not in Phase 4
Phase 5 (Expand) / Phase 6 (Remediate): v1s shipped as proofs; theirdeep halves stay gated on Python adoption being real (metrics, not calendar).Reorder triggers: enterprise partner → pull Phase 4 signing forward; JS-side CVE wave → deepen Phase 5; any in-the-wild crasher or FP class → jumps the queue into Phase 0’s harness immediately.
Sequencing traps (explicit anti-goals)
- Building the LLM
fixengine before a formally measured scanner. - Public launch before SARIF + a precision number - wastes the one-time attention spike.
- Buying recall with false positives - always net-negative for a security tool.
- Default-on telemetry - instantly off-brand.
- Premature performance work before big-repo users exist.
Metrics ladder
| Phase | Primary metric |
|---|---|
| 0 Measure | Published precision/recall; regression gate live |
| 1 Distribute | # repos running Palisade in CI (North Star) |
| 2 Cover | # frameworks covered; # external rule PRs merged |
| 3 Scale | Scan time on 500k LOC; % PRs on incremental mode |
| 4 Certify | Signed-release adoption; enterprise evals reaching pilot |
| 5 Expand | JS precision gate green; # JS/TS repos scanned |
| 6 Remediate | fix accept/merge rate on real findings |