|
| 1 | +# Runtime evaluation harness |
| 2 | + |
| 3 | +`make eval` runs `cmd/odek-eval`, a deterministic harness around the |
| 4 | +production `internal/loop.Engine`. It starts an OpenAI-compatible provider on |
| 5 | +localhost, feeds scripted assistant replies, and exposes small stateful |
| 6 | +fixture tools. No credentials, external network, or live model is used. |
| 7 | + |
| 8 | +Each scenario keeps its model messages separate from its oracle. The oracle |
| 9 | +checks fixture state and observed tool calls, so an assistant saying “done” |
| 10 | +cannot by itself make a task pass. The JSON report contains per-case scenario |
| 11 | +status, independently determined `task_success`, tool calls, synthetic |
| 12 | +scripted-provider token counts, and elapsed milliseconds. `false_completion_rate` is the fraction of |
| 13 | +cases whose oracle explicitly marked a success claim while required fixture |
| 14 | +state was absent; it is a regression signal, not a model-quality score. |
| 15 | +`cost_known` is always false for the shipped harness; it does not configure |
| 16 | +prices or estimate cost. |
| 17 | + |
| 18 | +The initial suite covers verified artifact work, a failed read followed by a |
| 19 | +false success claim, unrelated reads after a write, transient failure and |
| 20 | +recovery, cancellation, and plan acceptance checks for success, failed |
| 21 | +evidence, missing evidence, and incremental revision behavior. Negative cases can still be scenario passes |
| 22 | +when the oracle correctly records that the task did not succeed. The plan |
| 23 | +cases use the production `plan` tool and `PlanStore`, including the runtime |
| 24 | +incomplete marker for failed or missing evidence. |
| 25 | + |
| 26 | +The reusable `internal/eval.RunWithOptions` API accepts a per-case client |
| 27 | +factory. An application may use that adapter to compare a separately |
| 28 | +authorized real provider, but it must provide its own credentials, network |
| 29 | +policy, and model-message adapter. The shipped CLI intentionally does not |
| 30 | +evaluate live-model intelligence. Cost reporting stays unknown because the |
| 31 | +harness does not configure token prices. |
| 32 | + |
| 33 | +## Running and extending the suite |
| 34 | + |
| 35 | +```bash |
| 36 | +make eval |
| 37 | +# Save only the JSON report (without make's command echo): |
| 38 | +go run ./cmd/odek-eval > eval-report.json |
| 39 | +``` |
| 40 | + |
| 41 | +Exit status is 0 when every scenario passes, 1 for scenario failures, and 2 |
| 42 | +if the report cannot be encoded. The eleven-case baseline includes a deliberate |
| 43 | +unguarded false-success control: all scenarios pass while |
| 44 | +`false_completion_rate` is 1/11 (about 0.091). That expected control is not a failure of |
| 45 | +the checked-plan guard or a live-model benchmark. |
| 46 | + |
| 47 | +Add a case to `internal/eval.Scenarios` with fresh fixture state, scripted |
| 48 | +responses, tools, and an independent oracle. Assert the required state and |
| 49 | +observed tool outcomes, including expected failures; do not accept a success |
| 50 | +claim as proof. `task_success` answers whether the fixture task was completed; |
| 51 | +`scenario_passed` answers whether the runtime behaved as the test expected. |
| 52 | +Add regression assertions in `internal/eval/eval_test.go`, then run: |
| 53 | + |
| 54 | +```bash |
| 55 | +go test -count=1 -timeout=120s ./internal/eval ./internal/loop |
| 56 | +``` |
0 commit comments