Last Updated: 2026-09-22
This is the canonical operating policy for agents that use the KnowCode MCP server. Keep agent rules, setup guides, and prompts pointed here instead of redefining thresholds or token budgets in multiple places.
Agents should minimize expensive context generation by asking KnowCode for the smallest useful repository context first, then escalating only when the reduced context is not enough to answer safely.
Three consolidated tools, each selecting a capability with an action:
| Tool | Actions | Nature |
|---|---|---|
knowcode_retrieve |
query, search, context, trace, semantic_search |
Read-only. The hot path. |
knowcode_lifecycle |
build, index, export |
Writes artifacts. Confirmed per call. |
knowcode_inspect |
job_status, doctor, freshness, quality, stats, preflight, history, telemetry |
Read-only. |
Tool schemas are injected into every LLM request, so schema size is a
recurring per-turn cost. Measured at ~4 chars/token: the previous five flat
tools cost ~650 tokens for 5 capabilities; this surface costs ~1,110 for 14;
one tool per capability would cost ~2,000. The ceiling is enforced by
tests/unit/mcp/test_consolidated_surface.py.
The split is by concern because client permissions are per-tool
(mcp__knowcode__<tool>): knowcode_retrieve can be allowlisted so ordinary
questions never prompt, while knowcode_lifecycle still asks.
Deliberately not exposed: telemetry deletion (irreversible), server /
mcp-server / install (host-level process control), and ask (pays a
second LLM — an agent should consume context via query instead).
The five original flat tools (search_codebase, get_entity_context,
trace_calls, retrieve_context_for_query, assess_codebase_quality) remain
available for one release behind --legacy-tools, off by default.
The server starts whether or not artifacts exist. A repository KnowCode has never seen is bootstrapped through the surface itself:
{"tool": "knowcode_lifecycle", "action": "build"}Indexing embeds pending chunks in cross-file batches with bounded concurrency
and retry/backoff (src/knowcode/indexing/embedding_batch.py); large first
builds still take a while, which is why lifecycle actions return a job_id
immediately and are polled via knowcode_inspect action=job_status. Poll
until terminal:
{"tool": "knowcode_inspect", "action": "job_status", "job_id": "j-…"}Never report a build as successful before state is succeeded and
result.published is true. A build that fails after parsing leaves the
previously published generation current, so a non-zero entity count does not
mean retrieval improved.
Two lifecycle guardrails: knowcode_lifecycle action="export" requires an
output argument or validation fails (mcp/lifecycle.py:148-149), and only
one lifecycle job may run at a time — a concurrent submit returns
code="job_already_running" (mcp/jobs.py:57-66).
From a terminal, the equivalent remains:
uv run knowcode build .
uv run knowcode doctor --store . --mcpTwo store layouts are valid, and both count as ready: a published generation
carrying knowledge.db, or the legacy flat knowcode_knowledge.json. A
current build produces the former and writes the latter only with
export_json.
Use knowcode_retrieve with action="query" whenever the current
conversation does not already contain enough repository context.
Default MCP arguments:
{
"action": "query",
"query": "<user question>",
"task_type": "auto",
"max_tokens": 1500,
"limit_entities": 1,
"expand_deps": false,
"verbosity": "minimal"
}Use larger starting budgets only when the question clearly needs more breadth:
| Query type | max_tokens |
limit_entities |
expand_deps |
|---|---|---|---|
| Locate or explain one symbol | 1500 | 1 | false |
| Debug a concrete failure | 2000 | 2 | true |
| Review or extend a feature area | 3000 | 2-3 | true |
| Trace callers, callees, or impact | 2000 | 2 | true |
Retrieval never builds. If artifacts are missing it returns
code="missing_knowledge_store" with a hint naming the lifecycle call —
it will not silently spend minutes and embedding quota on your behalf.
verbosity="minimal" is the default for IDE agents. In minimal mode,
KnowCode summarizes context and omits raw source/evidence metadata where it can.
Escalate only when the returned context_text is not enough:
- Keep
verbosity="minimal"and raisemax_tokensorlimit_entitiesif the answer needs more breadth. - Use
verbosity="standard"if implementation detail or raw source is missing. - Use
verbosity="verbose"if ranking evidence or retrieved chunk provenance is needed. - Use
verbosity="diagnostic"only for tests and debugging the retrieval system, not as an agent default.
The local-answer threshold is configured in aimodels.yaml:
config:
sufficiency_threshold: 0.8Agents should use that configured value. The recommended starting value is
0.8; tune it later from eval or telemetry data, not by hard-coding competing
thresholds in agent prompts.
If sufficiency_score >= sufficiency_threshold and context_text is non-empty,
the agent may answer from the retrieved context without sending repository
source to an external LLM.
If the score is below threshold, the agent should first use the verbosity ladder when the missing information is likely available locally. Only fall back to a larger external LLM prompt after the local context has clearly failed or the user explicitly asks for a broader synthesis.
Prefer action="query" for natural-language questions. Use the others only
for focused follow-up:
search: find entities by known name or pattern.context: fetch context for a specific entity after its ID is known.trace: inspect callers or callees. Accepts a bare name and resolves it; an unresolvable name returnscode="entity_not_found"rather than an empty list that would read as "nothing calls this".semantic_search: raw ranked chunks, capped per chunk. Bypasses the sufficiency projection, so it is a follow-up, not a first choice.knowcode_inspect action="quality": the persisted pre-flight report when you need to judge how much to trust local context.knowcode_inspect action="freshness": whether artifacts lag the working tree. Check this before trusting context on an actively edited repository.
Typed failures surface as a code plus a hint naming the next action.
Beyond the two cases noted above, the catalog:
| Code | Raised by | Meaning |
|---|---|---|
missing_knowledge_store |
query |
no knowledge store; run the lifecycle build the hint names |
entity_not_found |
trace |
bare name resolved to no entity |
missing_semantic_index |
semantic_search |
no usable vector plane in the index generation |
path_outside_root |
root resolution | requested path escapes the MCP server root |
unknown_job / no_jobs |
inspect action=job_status |
job id not found / no jobs recorded yet |
missing_preflight_report |
inspect action=quality |
generation has no preflight_report.json |
job_already_running |
knowcode_lifecycle |
one lifecycle job at a time; submit again after it finishes |
Use this compact rule in agent-specific config files:
When repository context is needed, follow docs/mcp-contract.md.
Start with knowcode_retrieve action=query, verbosity=minimal, and the smallest
budget that fits the task. Escalate to standard or verbose only when the minimal
context is insufficient. Use the configured sufficiency_threshold from
aimodels.yaml to decide whether to answer from local context.
If artifacts are missing or stale, run knowcode_lifecycle action=build and poll
knowcode_inspect action=job_status until it succeeds before retrying.This document outlines the sources of token overhead when using the KnowCode MCP server and provides five concrete strategies to reduce this overhead by approximately 90%.
The MCP approach can be token-expensive due to the following overhead sources:
The estimates below were measured against the superseded pre-consolidation five-tool surface, before Strategies 1–5 shipped. They record the baseline the strategies were justified against, not today's per-call cost.
| Overhead Source | Est. Tokens/Call | Notes |
|---|---|---|
| 5 tool schemas injected into every LLM prompt | ~650 | IDE injects ALL tool definitions into the system prompt on every turn |
context_text (up to 4000 tokens of source code) |
~2000-4000 | Full source code for 3 entities with callers/callees |
evidence array (up to 15 entries × 5 fields each) |
~300-500 | rank, chunk_id, entity_id, score, source per chunk |
selected_entities metadata (per-entity duplication) |
~150-300 | Re-echoes entity_id, task_type, total_tokens, truncated, sufficiency_score |
Echoed fields (query, max_tokens, retrieval_mode, etc.) |
~100-200 | The query text itself is echoed back |
json.dumps(indent=2) whitespace |
~200-400 | Pretty-printing doubles the character count |
| Total per call | ~4,500–8,000 |
Implementation Status: all five strategies have shipped (see Status section below).
The 5 original flat tools (search_codebase, get_entity_context, trace_calls,
retrieve_context_for_query, assess_codebase_quality) were each injected into
every LLM request.
Change (shipped): the default surface is three concern-split tools with
action enums — knowcode_retrieve, knowcode_lifecycle, knowcode_inspect.
The split is by concern rather than one knowcode tool because client
permissions are per-tool, so retrieval can be allowlisted while builds stay
confirmed. The five flat tools remain behind --legacy-tools, off by default.
Consolidation shipped together with a capability expansion, so the recurring schema cost went from ~650 tokens for 5 capabilities to ~1,110 for 14 rather than down to ~200. Cost per capability improved ~2.3x; absolute per-turn cost did not fall until Strategy 5 shipped. A ceiling test guards the schema at ~1,140 tokens under a 1,200 limit.
The response from the query path returned 12 fields. The agent realistically only needs 2–3 of these fields to proceed:
context_text(the actual content)sufficiency_score(the decision metric)total_tokens(for budget awareness)
Recommendation: Omit all other fields (query echo, task_confidence, retrieval_mode, max_tokens, truncated, evidence[], selected_entities[], errors[]) by default, or gate them behind a verbose=true flag.
The previous defaults in the MCP server were max_tokens=6000 and limit_entities=3. max_tokens defaults: 1500 for knowcode_retrieve action=query (mcp/server.py:380); 4000 for direct RetrievalOrchestrator use, the legacy flat tool, and the REST query endpoint. For most day-to-day queries, this is excessive.
Recommendation: Update your agent rules (.agent/rules/context.md) to use tiered budgets:
max_tokens=1500, limit_entities=1is sufficient for "locate" and "explain" queries.max_tokens=2000, limit_entities=2for "debug" queries.- Only use
max_tokens=3000+for broad "extend" or "review" queries.
Tool results in src/knowcode/mcp/server.py were originally serialized with
json.dumps(result, indent=2), roughly doubling the character count of every
response.
Change (shipped): serialization now uses a condensed format, instantly removing hundreds of unnecessary whitespace tokens:
return json.dumps(result, separators=(',', ':'))Full source_code used to be dumped into context_text on every call.
Change (shipped): response profiles are one definition,
retrieval/response_profiles.py, that every source-bearing retrieval consumer
answers to. Summary-first is the default on query and context; escalation
rides the existing verbosity ladder rather than a new schema enum, because the
schema is paid on every turn. semantic_search is the explicit source request
and stays raw. Measured on this repository, default context fell from 4,313 to
674 bytes on the probe entity.
debug and review include raw source even at minimal verbosity. This is
a floor, not a default: minimal is the default, so a floor an explicit
default could cancel is no floor at all.
This set is exactly {debug, review}. It is the canonical statement of the
rule — retrieval/response_profiles.py implements it and
tests/unit/retrieval/test_response_profiles.py pins the membership against
this section, so widening the set is a deliberate contract change rather than a
quiet edit. Adding a task type here takes source away from every default
response of that type.
If all strategies are implemented, the token savings would be dramatic. The
first row's projection assumed the superseded five-tool surface; the shipped
consolidated surface is pinned under 1200 tokens by
tests/unit/mcp/test_consolidated_surface.py:79:
| Strategy | Est. Token Savings |
|---|---|
| Tool consolidation (3 tools vs the old 5 schemas) | schema pinned under 1200 tokens per turn |
| Stripped response metadata | ~800 tokens saved per call |
| Lower default token limits | ~3000 tokens saved per call |
| Compact JSON formatting | ~300 tokens saved per call |
| Summaries vs full source code | ~2000 tokens saved per call |
| Overall Reduction | ~6,500 tokens → ~800 tokens (≈88% reduction) |
The following optimizations have been fully implemented in the KnowCode codebase:
-
Stripped Response Metadata (Strategy 2 — Implemented): The default
minimalmode returnscontext_text,sufficiency_score, andtotal_tokens, plus: thefreshnessblock the service appends to every query response;reduction_summarywhen the response was summarized, orsource_included: truewhen raw source was included (task typesdebug/review, or verbosity ≥standard); anderrorswhen present (src/knowcode/retrieval/orchestrator.py:342-364,src/knowcode/service.py:643). Strip-and-compact narratives must treat these fields as part of the minimal payload. All non-essential fields (such as query echo, task confidence, evidence lists, etc.) are excluded, saving ~800 tokens per call. -
Lowered default token limits (Strategy 3 — Implemented):
max_tokensdefaults: 1500 forknowcode_retrieve action=query(mcp/server.py:380); 4000 for directRetrievalOrchestratoruse, the legacy flat tool, and the REST query endpoint. -
Compact JSON Formatting (Strategy 4 — Implemented): Responses in
server.pyare serialized usingjson.dumps(result, separators=(',', ':')), eliminating unnecessary whitespace and saving ~300 tokens per call. -
Tool consolidation (Strategy 1 — Implemented): the default surface is
knowcode_retrieve,knowcode_lifecycle, andknowcode_inspect; the five flat tools are available behind--legacy-toolsfor one release, off by default. A ceiling test guards the recurring schema cost. -
Summary-first responses (Strategy 5 — Implemented, 2026-09-09): default retrieval payloads are summaries unless the agent escalates
verbosityor the task type is one of the source-hungry task types. Payload sizes are capped by regression test and recorded per action in local telemetry.