Status: Active planning document. Open work only.
Why the project is where it is — dated decisions, shipped-work narrative, and the measured reasons behind them — is Roadmap History. When an item here ships, strike it with a one-line dated note and move the narrative there.
KnowCode already has the core operational foundations: a canonical MCP
contract, freshness reporting, supported-language checks, knowcode doctor,
local telemetry, and an MCP handshake check. The next release should turn
those foundations into a calibrated, verifiable, and easy-to-adopt product.
The most important open evidence is retrieval routing quality. The committed 60-record corpus is now explicitly calibration-only, with atomic facts, prohibited claims, and AST-resolved source citations. It is not a locked holdout and cannot enable local answering. Runtime routing therefore defaults to an empty task allowlist until independent machine adjudication and the blocking Python external gates pass. See Testing & Evaluation for the evidence contract.
Defects that no workstream owns are recorded in the engineering backlog. Put the next finding there.
- Correctness is the release gate. Token savings, consumer installers, and lifecycle convenience must not make stale or uncalibrated answers appear trustworthy.
- A green
knowcode doctor --mcpis necessary but not sufficient. It proves local artifacts and the MCP transport are healthy; retrieval evaluation and freshness/coverage tests prove that the answer path is trustworthy. - The canonical MCP contract remains the single policy source. Runtime code, agent rules, setup documentation, and tests must agree on it.
- Preserve compatibility deliberately. Changes to MCP tool names or response shapes require a documented migration path and regression coverage.
- Keep telemetry local by default, make its retention and privacy tradeoffs explicit, and use measured data before changing thresholds or budgets.
These hold for every release. The evidence that established them is in Roadmap History.
- MCP operating contract. Contract tests must exercise production-like minimal responses.
- Freshness and coverage safety. Modify/create/delete/rename tests and
doctorfreshness checks remain required. - Local readiness verification.
doctormust stay fast, deterministic, and actionable. - Local telemetry. Event-schema compatibility and failure isolation tests remain required.
- Derived vector plane. A published generation must contain nothing named
vectors.*, and a rebuilt plane must return identical results to a persisted one.
Goal: turn the existing evaluation harness into a release-quality source of truth for local-answer routing.
Why first: the current evaluation data says a score of 0.8 is
over-confident. No payload, onboarding, or automation improvement can make an
uncalibrated local-answer gate safe.
Status: deliberately unblessed. The harness lives in the independent
knowcode-evals repository; KnowCode holds only a checksum-pinned policy
consumer and runtime enforcement tests. No qualifying locked holdout plus full
external baseline has yet passed.
Work:
- Run the real
Agent.smart_answerescalation path over source-cited records. Require two independently configured judge providers to return strict claim verdicts with valid citations, three times each; disagreement, partial support, malformed output, missing evidence, or provider failure is a fail. - Validate the judges with mechanically generated true/false source mutations.
Require macro F1 of at least
0.95and zero false acceptance of critical negative canaries. - Select a global threshold from
0.50through1.00. Enable a task type only when a locked holdout has at least 29 routed cases, zero critical failures, and a one-sided exact-binomial 95% correctness lower bound of at least0.90. - Make Python RepoBench-R archive/v0 and RepoQA blocking. Compare KnowCode with
BM25 over the identical candidate corpus and token budget; every primary
metric's seeded paired-bootstrap lower bound must be at least
-0.02. - Track fixed CrossCodeEval-100 and SWE-bench-Lite-50 A/B suites as non-blocking downstream evidence until three consecutive runs support an explicit promotion decision.
- Publish a versioned machine-verification artifact with the selected policy, floors, source hashes, dataset revisions, provider/model identities, prompt hashes, and canary results. Never describe it as human-reviewed.
Before the next run. The 60 existing records are void: they describe a bundle the product no longer builds. Re-collect rather than reusing them (why).
Threshold input that changed. The routing_quality_floor of 0.90 was
selected against a formula that could only return 1.00 or block. With a fixed
denominator, complete bundles land at 0.95–0.96 and source-less ones at
0.45–0.86. Re-select over that range rather than carrying 0.90 across; the sweep
in work item 3 is where that happens.
Exit criteria: the routing policy is source-verified, independently machine-adjudicated, externally benchmarked, versioned, and enforced by CI. Missing credentials, source or dataset drift, judge instability, BM25 inferiority, and blessed-baseline regression all fail closed.
Goal: prove the canonical MCP policy is what production code actually does.
Work: all three original items shipped 2026-08-11 (agent metadata requests,
end-to-end escalation coverage, and a doctor/release-checklist conformance
audit). See Roadmap History.
Audit re-run 2026-09-09. The 2026-08-11 audit had validated
retrieve_context_for_query, a surface P3's consolidation then replaced, and
the contract's appendix went on describing the old one. The documents were
corrected and all four
release checklist conformance items re-verified
against the running surface: the default tool list is knowcode_retrieve,
knowcode_lifecycle, knowcode_inspect with the flat five reachable only
through tool_definitions(include_legacy=True); the published schema defaults
are max_tokens=1500, limit_entities=1, verbosity="minimal"; the minimal
projection is an allowlist, now pinned to an exact field set by
tests/unit/retrieval/test_orchestrator_profiles.py; and doctor --mcp lists
three tools and calls knowcode_retrieve successfully.
The checklist's boxes are ticked per release rather than once, so the audit is re-run each time — but P2's own exit criteria are met and the workstream is ready to close.
Not evidence for this: the handshake ran against a store carrying BL-32 builder drift. The contract surface, schema defaults, and projection are all code-side and were verified in-process against the running code, so the drift does not weaken them; it does mean the retrieved content described a graph this code did not build.
Exit criteria: agent, MCP, CLI, docs, and tests use one contract with no implicit reliance on fields hidden by minimal mode.
Goal: reduce recurring MCP schema and context costs without weakening the calibrated correctness floor.
Dependencies: P1 and P2. Evaluate every change against the blessed golden baseline and routing-quality gate.
Work:
- End the legacy five-tool surface. Consolidation shipped as three
concern-split tools rather than the single
knowcodetool this item first proposed, because client permissions are per-tool. The five flat tools remain behindmcp-server --legacy-tools, off by default, "for one release" — a release this roadmap has never named. Name it, publish the migration guidance the exit criteria require, then remove the surface. Summary-first response profiles.Shipped 2026-09-09. The rule now lives in the MCP contract.Byte/token caps per action, and payload-size distributions in telemetry.Shipped 2026-09-09.
Exit criteria: default tool/result payloads are measurably smaller, golden retrieval and routing quality do not regress, and migration guidance is published before the legacy tool surface changes.
Goal: make the canonical MCP policy and connection setup installable rather than a collection of hand-maintained per-agent files.
Dependencies: P2; P3 if the consolidated tool becomes the default surface.
Work:
- Add a product-owned
.knowcode/agent-rules.mdthat references the canonical contract and distinguishes semantic queries from direct file/grep work. - Add an idempotent
knowcode install-agent <consumer>flow for verified MCP consumers, beginning with the clients actively supported by the project. - Add
knowcode doctor --agent <consumer>checks for generated configuration, rule inclusion, and an end-to-end tool invocation where the client supports programmatic verification. - Treat each client configuration as an explicit compatibility target. Do not claim support for a consumer until its current configuration syntax and runtime behavior are verified.
Exit criteria: a supported consumer can be configured predictably from one command, its rules point to the canonical policy, and doctor can identify a broken setup with an actionable fix.
Goal: reduce the chance that ordinary repository changes leave KnowCode artifacts stale while preserving explicit user control.
Dependencies: P2. The existing freshness safety remains the fallback if automation is unavailable or fails.
Work:
- Add an optional repo-local freshness manifest that records the source state used to build the store and index, including Git state where available.
- Offer opt-in post-commit and background refresh helpers. They must never block a commit, silently modify unrelated configuration, or hide a refresh failure.
- Document optional always-on watch-service templates for supported local environments, with a normal foreground fallback.
Exit criteria: users can opt into low-friction refresh automation, while
every failure path remains visible through freshness metadata and doctor.
Goal: turn existing JSONL telemetry into a local decision tool.
Dependencies: P1 and P2, so summary metrics use calibrated routing terms.
Work:
- Add
knowcode stats --usage [--since <duration>]backed by the existing telemetry summary support. - Report calls per day, routing rate, mean sufficiency, stale-response count, payload-size distribution, and available per-consumer attribution.
- Add clear retention, redaction, and deletion guidance for local query logs.
Exit criteria: a developer can understand whether KnowCode is used, trusted, fresh, and cost-effective without manually parsing JSONL.
Goal: make an index proportional to the source it describes, and make every document in that source retrievable. The plan of record is Storage Footprint & Optimization Plan; its §17 is the running execution log and carries the measured ledger.
Dependencies: none for the storage phases. The embedding-selection phase depends on P1, because it narrows the semantic candidate set and may only ship behind a measured recall gate.
Governing decision: the plan's DR-4 — no phase may cost retrieval quality, whatever the byte saving.
Phase F and the int8 ANN cache are rejected, not deferred. Both are closed in the backlog with measured reasons. Work item 5's premise — that an unembedded chunk stays reachable through the exact, path, and FTS planes — holds only because F is rejected and the exact plane stays.
Where this leaves the stream. Everything DR-4 permits inside a generation has shipped, and one generation now measures 50.78 MB. The larger remaining lever is retention: two generations hold mostly identical bytes, which is Phase G, and it is a different axis from making a generation smaller.
Work, in order:
Document identity and chunking correctness.Shipped 2026-08-29, Phase B. Prose coverage 41.6% to 95.7%; the generation fell 2.18 MB while indexing 2.3 times as much prose.Shipped 2026-08-29, Phase A2, 3.79 MB.VACUUMbefore the manifest is digested.Lossless encoding.Shipped 2026-08-29, Phase C, 10.0 MB across both artifacts.Stop persisting the rest of the derived data.Shipped 2026-08-30, Phase D: D2 the contentless FTS5 term index, 4.46 MB; D3entities.source_coderesolved from disk, 3.67 MB.- Embedding-selection policy, gated on the retrieval evaluation harness. Chunks below a content-size threshold stay stored and stay reachable through the exact, path, and FTS planes; they leave only the semantic candidate set. Under DR-4 it ships on a measured no-loss result, never on a loss judged small enough to accept.
The lossless remainder.Shipped 2026-09-08, 2.80 MB measured.- Decide Phase G, or retire the plan without it. Content addressed across retained generations, so retention costs the delta rather than the whole. Lossless by construction and worth more than everything above combined, but larger than the rest of the plan put together. The correctness debt that was meant to close before the stream does either way is now paid: a chunk is content-addressed by SHA-256 (BL-11) and a staged rewrite witnesses its own losslessness (BL-8).
Exit criteria: every tracked document in a repository is retrievable, one generation is a small multiple of the source it describes rather than an order of magnitude, and no phase past the first changes what an existing query returns without a measured recall number to justify it. DR-4 tightened the last clause: a measured recall number is no longer a licence to ship a loss, only evidence that there is none.
P1 and the implementation work in P2 can proceed in parallel, but the
release-candidate evaluation must run against the P2 production response
shape. P3 follows once the correctness baseline and contract are stable.
P4, P5, and P6 can proceed independently after their listed dependencies
are met. P7's remaining storage phases are independent of P1 through P6, and
its embedding-selection phase is gated on P1's evaluation harness.
The next trust release ships only when all of the following are true:
- P1 and P2 exit criteria are met.
- The full test suite, including the server-extra API contract suite, passes.
knowcode doctor --mcppasses for the supported local configuration.- Freshness and language-coverage checks report no unresolved correctness warnings for the target repository.
- No
Criticalitem is open in the engineering backlog. Read the backlog rather than trusting any roster written here: a previous version of this line namedBL-1long after Phase B fixed it, and claimed nothing was open right up until an audit found more. - Preflight on this repository's own graph keeps the unresolved-reference
resolution rate at 0.790 or better — the BL-34
baseline, ratcheted. Of the 4,077 holes at that baseline, 2,867 are
receiver-qualified calls whose file states no type (unannotated
parameters, cross-function dataflow — a type checker's residue, not a
linking defect), 1,210 are ambiguous or unknown bare names, and 37 are
typed-but-unbound holes the evidence reports as
unresolved_typed_receivers; a change to the rate in either direction updates the baseline here in the same commit that moved it.
P3 through P7 improve efficiency, adoption, and footprint, but they are not permitted to weaken these release gates.
- A shared or hosted multi-tenant knowledge service.
- Making the HTTP gateway the primary path for this roadmap.
- Automatically rewriting agent configuration without an explicit user command.
- Declaring compatibility for an agent consumer before its current setup and invocation behavior are verified.
- Cost optimization that bypasses the retrieval and routing-quality gates.