Skip to content

Commit 4b79cff

Browse files
authored
feat: incremental planning and verified completion (#253)
* feat: require tool evidence for declared plan checks * docs: explain acceptance checks and runtime evaluations * feat: revise plans incrementally without dropping acceptance checks * fix: resolve release CodeQL findings in planning and evals
1 parent 1a66bd4 commit 4b79cff

22 files changed

Lines changed: 2653 additions & 108 deletions

‎Makefile‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -42,6 +42,10 @@ test-internal: ## Run only internal package tests (excludes cmd/odek)
4242
test-cmd: ## Run cmd/odek unit tests (env-gated E2E/sandbox suites skipped)
4343
$(GO) test -short -count=1 -timeout 600s ./cmd/odek -skip 'TestE2E_|TestMCPE2E|TestSandbox'
4444

45+
.PHONY: eval
46+
eval: ## Run deterministic local runtime evaluations (no credentials or external network)
47+
$(GO) run ./cmd/odek-eval
48+
4549
.PHONY: test-cli-serve
4650
test-cli-serve: ## Serve/WS/REST headless-run surface only (~1 min)
4751
$(GO) test -short -count=1 -timeout 300s -run 'TestServe|TestWS|TestRestRun|TestPrompt|TestHandlePrompt' ./cmd/odek

‎README.md‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -201,7 +201,8 @@ odek run "@README.md what does this project do?"
201201
| [Extensions](docs/EXTENSIONS.md) | `odek-extension/v1` contract: MCP limits, artifact refs, event stream, external refs, budgets |
202202
| [Maintenance](docs/MAINTENANCE.md) | Storage janitor: retention, log rotation, `odek cleanup` |
203203
| [Extended Memory](docs/EXTENDED_MEMORY.md) | Atomic long-term memory layer (opt-in) |
204-
| [Planning](docs/PLANNING.md) | Plan tool, protected plan message, security model |
204+
| [Planning](docs/PLANNING.md) | Incremental plan revisions, acceptance checks, completion evidence |
205+
| [Runtime Evals](docs/EVALS.md) | Deterministic task scenarios, independent checks, JSON reports |
205206
| [Tool Selection](docs/TOOL_SELECTION.md) | Tool whitelist/blacklist guide and names reference |
206207
| [Daily Worker](docs/DAILY-WORKER.md) | Headless scheduled-worker patterns |
207208
| [Providers](docs/PROVIDERS.md) | go-llm-sdk registry, `--provider`, v2 knobs |
@@ -238,6 +239,7 @@ The full `Config` struct supports: `Provider`, `Providers`, `BaseURL` (selected-
238239
go test ./... # full suite, no setup required
239240
go test -race ./... # also clean under the race detector
240241
go test -cover ./... # per-package coverage report
242+
make eval # scripted runtime evaluations, no model credentials
241243
ODEK_E2E=1 go test ./cmd/odek/ # opt-in Docker / subprocess E2E suite
242244
```
243245

‎cmd/odek-eval/main.go‎

Lines changed: 24 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,24 @@
1+
// odek-eval runs deterministic, localhost-only runtime evaluations.
2+
package main
3+
4+
import (
5+
"context"
6+
"encoding/json"
7+
"fmt"
8+
"os"
9+
10+
"github.com/BackendStack21/odek/internal/eval"
11+
)
12+
13+
func main() {
14+
report := eval.Run(context.Background(), eval.Scenarios())
15+
enc := json.NewEncoder(os.Stdout)
16+
enc.SetIndent("", " ")
17+
if err := enc.Encode(report); err != nil {
18+
fmt.Fprintln(os.Stderr, err)
19+
os.Exit(2)
20+
}
21+
if report.Failed != 0 {
22+
os.Exit(1)
23+
}
24+
}

‎docs/CONFIG.md‎

Lines changed: 5 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -408,7 +408,11 @@ Gives the agent a protected plan tool and a plan message that survives context t
408408
|-------|---------|---------|----------|-------------|
409409
| `planning.enabled` | `true` | `ODEK_PLANNING` | `--planning` / `--no-planning` | Enable the plan tool and protected plan message |
410410
| `planning.max_steps` | `12` | — | — | Plan steps allowed (clamped 1–50) |
411-
| `planning.max_render_chars` | `2000` | — | — | Cap on the rendered plan shown in the UI (clamped 200–8000) |
411+
| `planning.max_render_chars` | `2000` | — | — | Cap on the protected plan render (clamped 200–8000); checked or revised plans must fit in full, including reserved evidence space |
412+
413+
Acceptance checks need no additional flag: declare them through `plan create`
414+
while planning is enabled. Large declarations may need a higher operator-set
415+
`max_render_chars`; project config cannot raise this cap.
412416

413417
Feature behavior, verbs, and the security model are documented in [PLANNING.md](PLANNING.md).
414418

‎docs/DEVELOPMENT.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -172,6 +172,15 @@ go test -v -count=1 ./cmd/odek/ -run "TestSubagent|TestDelegateTasks"
172172

173173
Zero external test dependencies — tests use `httptest`, `testing`, and the standard library only.
174174

175+
### Runtime evaluations
176+
177+
Run `make eval` (or `go run ./cmd/odek-eval`) from the repository root. The
178+
eleven scripted scenarios use the production loop with localhost fixture tools
179+
and independent outcome checks. The command writes a JSON report and exits
180+
nonzero if a scenario assertion fails; expected task failures can still pass
181+
the scenario. No model credentials or external provider calls are needed.
182+
See [EVALS.md](EVALS.md) for report fields, limitations, and adding scenarios.
183+
175184
### Test layers
176185

177186
| Layer | Runner | What's tested |

‎docs/EVALS.md‎

Lines changed: 56 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,56 @@
1+
# Runtime evaluation harness
2+
3+
`make eval` runs `cmd/odek-eval`, a deterministic harness around the
4+
production `internal/loop.Engine`. It starts an OpenAI-compatible provider on
5+
localhost, feeds scripted assistant replies, and exposes small stateful
6+
fixture tools. No credentials, external network, or live model is used.
7+
8+
Each scenario keeps its model messages separate from its oracle. The oracle
9+
checks fixture state and observed tool calls, so an assistant saying “done”
10+
cannot by itself make a task pass. The JSON report contains per-case scenario
11+
status, independently determined `task_success`, tool calls, synthetic
12+
scripted-provider token counts, and elapsed milliseconds. `false_completion_rate` is the fraction of
13+
cases whose oracle explicitly marked a success claim while required fixture
14+
state was absent; it is a regression signal, not a model-quality score.
15+
`cost_known` is always false for the shipped harness; it does not configure
16+
prices or estimate cost.
17+
18+
The initial suite covers verified artifact work, a failed read followed by a
19+
false success claim, unrelated reads after a write, transient failure and
20+
recovery, cancellation, and plan acceptance checks for success, failed
21+
evidence, missing evidence, and incremental revision behavior. Negative cases can still be scenario passes
22+
when the oracle correctly records that the task did not succeed. The plan
23+
cases use the production `plan` tool and `PlanStore`, including the runtime
24+
incomplete marker for failed or missing evidence.
25+
26+
The reusable `internal/eval.RunWithOptions` API accepts a per-case client
27+
factory. An application may use that adapter to compare a separately
28+
authorized real provider, but it must provide its own credentials, network
29+
policy, and model-message adapter. The shipped CLI intentionally does not
30+
evaluate live-model intelligence. Cost reporting stays unknown because the
31+
harness does not configure token prices.
32+
33+
## Running and extending the suite
34+
35+
```bash
36+
make eval
37+
# Save only the JSON report (without make's command echo):
38+
go run ./cmd/odek-eval > eval-report.json
39+
```
40+
41+
Exit status is 0 when every scenario passes, 1 for scenario failures, and 2
42+
if the report cannot be encoded. The eleven-case baseline includes a deliberate
43+
unguarded false-success control: all scenarios pass while
44+
`false_completion_rate` is 1/11 (about 0.091). That expected control is not a failure of
45+
the checked-plan guard or a live-model benchmark.
46+
47+
Add a case to `internal/eval.Scenarios` with fresh fixture state, scripted
48+
responses, tools, and an independent oracle. Assert the required state and
49+
observed tool outcomes, including expected failures; do not accept a success
50+
claim as proof. `task_success` answers whether the fixture task was completed;
51+
`scenario_passed` answers whether the runtime behaved as the test expected.
52+
Add regression assertions in `internal/eval/eval_test.go`, then run:
53+
54+
```bash
55+
go test -count=1 -timeout=120s ./internal/eval ./internal/loop
56+
```

0 commit comments

Comments
 (0)