Skip to content

Add the arbiter family (hiteshluke/arbiter-4b) - #49

Merged
cobanov merged 19 commits into
ollaya-dev:mainfrom
codemanhitesh:add-arbiter-family
Oct 6, 2026
Merged

cobanov merged 19 commits into
ollaya-dev:mainfrom
codemanhitesh:add-arbiter-family

Conversation

@codemanhitesh

@codemanhitesh codemanhitesh commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Add the arbiter family (hiteshluke/arbiter-4b)

Arbiter is Codekins Pvt Ltd / Zyot Lab's 4B decision model: Gemma 3 4B IT + LoRA + a fixed 24-slot pointer head
(arbiter-fixed-v1). One forward pass per question; noul reads slots 0..1, choice slots 2..17 (up to 16
options), score slots 18..23 (exactly 6 levels). Full write-up: docs/families/arbiter.md.

This update addresses the review point by point.

Review items

  1. Rust runtime. Layout in crates/ollaya-decision/src/arbiter.rs (port of layout.py, with unit tests);
    engine in crates/ollaya-runner/src/arbiter.rs, registered in engine.rs. A choice with more than 16 options
    and a score with other than 6 levels are rejected with a 422, never truncated.
  2. Parity.
    • check.py: the Rust-layout port against the reference prompt, id for id, on the shared request set plus 6
      extra Arbiter cases, with two tokenizer library versions: 420 rows identical, 0 mismatches.
    • goldens.py: the reference model in fp32 (token ids, last position, all 24 slot scores, option logits,
      probabilities).
    • crates/ollaya-runner/examples/parity_arbiter.rs: identical rows and rejections, every decision the same,
      slot scores within 1e-3.
  3. Registry. registry/v2/library/arbiter/manifests/* removed.
  4. Licensing. The catalog license field now reads "Apache-2.0 (LoRA adapter and head) and the Gemma Terms of
    Use (Gemma 3 base model)"; docs match. unsloth/gemma-3-4b-it@bf46152 is byte-identical to Google's
    google/gemma-3-4b-it for both safetensors shards and tokenizer.json (same sha256 and size per the Hub API);
    only metadata files differ. The original is gated (manual approval) and Ollaya's client has no Hugging Face
    token support, which is why the manifest points at the copy. Happy to repoint if you prefer.
  5. Conventions. docs/pr/arbiter.md folded into docs/families/arbiter.md; scripts/arbiter_hash_hf.py
    removed.

Found while porting

  • Training adds <bos>; the earlier layout did not. Fixed.
  • The head has no bias; the earlier reference loader expected one. Fixed.
  • The model repo's tokenizer.json truncates at 255 tokens, so the runtime takes the tokenizer from the base repo.

Accuracy

Measured on a T4 with the training prompt, one row at a time. Model-card numbers use the 4-bit base the adapter was
trained on; the same adapter and head were also run on the same rows with the unquantized base (fp32 compute, as
Ollaya runs it):

Benchmark n 4-bit base Unquantized base
BoolQ (validation) 1,000 0.849 0.853
ARC-Challenge (test) 500 0.738 0.762
CommonsenseQA (validation) 500 0.706 0.720
OpenBookQA (test) 500 0.722 0.748

Running the full BF16 base does not cost accuracy relative to the model card.

Status

  • Not run yet: the export, goldens and both parity runs on the real weights (they need a GPU), so there are no
    parity numbers yet; and typed-decisions accuracy on the 400 test states. Both are runnable with the commands
    in the docs. The pipeline (export, goldens, ONNX parity) was exercised end to end on a small random stand-in
    model: ONNX within about 1e-5 of the reference, all decisions agree.
  • Rust: the new Rust code has not been compiled or linted on our side (no Rust toolchain on our machine), so
    please treat CI as the first check; we will fix anything it reports.
  • Coverage: every score question in the shared set's 40 typed-decisions rows has 4 or 5 levels, so the fixed
    6-level head cannot answer them and they are rejected with a 422.

worldkingk777 and others added 12 commits October 1, 2026 23:54
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…t head)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
… slot scores)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…erance 1e-4)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…ensing)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…ding export)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Pins the upstream Hugging Face shas in ref.MODELS, catalog.py and the
registry/v2 arbiter manifests:
  base   unsloth/gemma-3-4b-it     bf46152c47f5dd20b896357cb51abc4c03b8ee8c
  model  hiteshluke/arbiter-4b     0c44271c59f89758e3cae17b032e98a9140093e9

ONNX export was not run in this environment (torch.export on the full
Gemma 3 4B causal LM is multi-hour and memory-heavy, and the pinned
convert/ deps still have to resolve from source), so the
sha256:ARBITER_*_PLACEHOLDER digests in the two manifests are left as
placeholders for a maintainer to fill after running
`python -m ollaya_convert.families.arbiter.export arbiter-4b
--out out/arbiter-4b` against the pinned revisions. The parity gate
(python -m ollaya_convert.families.arbiter.parity) is deferred to the
same step for the same reason.

Also scrubs incidental cross-family references in the arbiter sources
so the new family stands on its own.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Hash both base shards from unsloth/gemma-3-4b-it @ bf46152c, and the
adapter, head.pt and tokenizer from hiteshluke/arbiter-4b @ 0c44271c,
replacing the ARBITER_WEIGHTS_PLACEHOLDER_{1..4} and
ARBITER_TOKENIZER_PLACEHOLDER entries in registry/v2/library/arbiter/
manifests/{4b,latest}. The ONNX graph digest and the ollaya-hosted
config/decision/calibration/license blobs are produced by the export +
packaging step on GPU and stay as placeholders until then.

scripts/arbiter_hash_hf.py downloads each file at its pinned revision
with huggingface_hub, streams it through hashlib.sha256, prints the
JSON report and deletes the local copy; rerun it to verify.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Rewrite the "What maintainers need to finish" section into a split
"Done on the fork" / "Left for a maintainer with sufficient GPU"
and point at the explicit uv command plus the parity gate. Fill in
the pinned base and adapter SHAs in the family doc's status row and
rephrase the parity note so it says what still needs to run.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@codemanhitesh
codemanhitesh marked this pull request as ready for review October 1, 2026 20:11

@cobanov cobanov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@codemanhitesh, thanks for submitting Arbiter, and sorry for the wait. The Python side is a good start. Every family in Ollaya goes through the same path before it ships, and this PR is not there yet. Here is what it needs, in order.

1. The Rust runtime. Nothing runs this model today. arbiter-fixed-v1 needs:

  • a layout in crates/ollaya-decision that builds the prompt rows from a request (port layout.py);
  • an engine in crates/ollaya-runner that feeds the graph and reads the 24 slots per question type, registered in engine.rs.

kev.rs and jeeves.rs in both crates are close templates (decoder rows, last-position readout). Requests the fixed head cannot answer must be rejected with a clear 422, not truncated: a choice with more than 16 options, a score with other than 6 levels.

2. Parity, which is the gate.

  • Prompt check. A check.py that compares the layout port with your reference prompt, id for id, on our shared request set: the edge cases and 40 typed-decisions rows, as a JSONL here.
  • Goldens. goldens.py, written from your reference model in fp32 on those requests: token ids, slot scores and probabilities.
  • Runtime parity. A crates/ollaya-runner/examples/parity_arbiter.rs that runs the Rust runtime against those goldens: identical rows and rejections, every decision the same, scores within 1e-3. parity_jeeves.rs is a template.

We run the export, the goldens and the parity again on our own GPU before merging, so they only need to be runnable.

3. Leave the registry to us. Please remove registry/v2/library/arbiter/manifests/*. We write manifests with package.py when a tag ships, from the export we checked, so a manifest never carries placeholder digests.

4. Licensing.

  • License field. The base is Gemma 3, under the Gemma Terms of Use, so the catalog's "license" cannot say Apache-2.0 alone. Your license_text already names both; the field and the docs should match it.
  • Base repository. Ollaya pulls weights from the original author's repository. Please say whether unsloth/gemma-3-4b-it@bf46152 is byte-identical to Google's google/gemma-3-4b-it, and why the manifest should point at the copy.

5. Repository conventions.

  • Docs. Docs for a family live in docs/families/arbiter.md: sequence, graph, differences from upstream, parity, quality. Please fold docs/pr/arbiter.md into it or into this PR's description.
  • Hashing script. One-off tools like scripts/arbiter_hash_hf.py stay out of the repository. package.py already verifies every file against the Hub's sha256.
  • Accuracy. Please add typed-decisions accuracy (all 400 test states, argmax against the majority label), the number every model in the library is compared on. BoolQ and ARC can stay next to it.

Once 1 and 2 are in, we will do the GPU checks and help with the rest. Thanks again for the contribution.

worldkingk777 and others added 5 commits October 5, 2026 10:52
Manifests are written by package.py when a tag ships, from the checked export, and
package.py already verifies every file against the Hub's sha256.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… and parity on the shared set

- ref.py: the training script's three templates verbatim, tokenized by transformers as in
  training (with <bos>); Gemma3ForConditionalGeneration in fp32 + peft adapter + the
  bias-free 24-slot head, one unpadded row at a time.
- layout.py: the port the Rust runtime follows ([bos] + tokenizers, slots from
  decision.json, noul read as [false, true]); rejects a choice over 16 options, a score
  with other than 6 levels and a row over 8,192 tokens instead of truncating.
- check.py: port vs reference prompt id for id on the shared request set (JSONL) plus
  arbiter edge cases: 127 requests, 420 identical rows, 0 mismatches.
- goldens.py: token ids, all 24 slot scores, option logits and probabilities per question.
- export.py: Gemma 3 trunk (llm_common/gemma3.py, sliding-window layers) with the LoRA
  unmerged; tokenizer.json from the base repository.
- parity.py: ONNX Runtime vs the fp32 reference on the same requests.
- eval_refs.py: typed-decisions quality for arbiter, with answered/total coverage.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…iter

- ollaya-decision/arbiter.rs: the training prompt after <bos>, one causal row per
  question, slots from decision.json. A choice over 16 options is TOO_MANY_OPTIONS, a
  score with other than 6 levels or a row over max_row_tokens is INVALID_REQUEST (422),
  never truncated.
- ollaya-runner/arbiter.rs: feeds input_ids + last_pos, reads the 24 slot scores and
  returns each question's option logits (noul as [false, true]); registered in engine.rs.
- examples/parity_arbiter.rs: identical rows and rejections against the goldens, all 24
  slot scores within 1e-3, the same decision on every question.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ilies/arbiter.md

- catalog: license is "Apache-2.0 (LoRA adapter and head) and the Gemma Terms of Use
  (Gemma 3 base model)"; the license text adds the Gemma notice and policy links; the
  tokenizer comes from the base repository (the model repository's tokenizer.json has
  truncation to 255 tokens on).
- docs/families/arbiter.md: sequence, graph, differences from upstream, parity, quality,
  limits; the base repository's files compared with google/gemma-3-4b-it by Hub sha256.
  Typed-decisions accuracy is marked TODO until it is measured.
- docs/pr/arbiter.md removed.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…it ones

Same adapter, head and rows on both bases: equal or better on BoolQ, ARC-Challenge, CommonsenseQA and OpenBookQA.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@codemanhitesh
codemanhitesh marked this pull request as draft October 5, 2026 14:47
…y page

The GPU steps of the parity procedure, run by the maintainers on the RTX 4090 machine (transformers
4.57.6 and peft 0.19.1 for the reference), on the shared set: 127 requests, 6 accepted whole, 121
rejected by both, 420 rows.

- check.py: 420 rows identical, 0 mismatches, as reported in the PR.
- Export against the reference (parity.py): every decision the same, slot scores within 5.6e-5,
  probabilities within 9.3e-6.
- Runtime (parity_arbiter): identical rows and rejections, 420/420 decisions; slot scores within
  6.0e-5 on the x86-64 CPU and 8.5e-5 on CUDA (Microsoft ONNX Runtime 1.28.2 from the CUDA 13 pack),
  probabilities within 1.0e-5. A request of three or more questions takes 154 ms at the median in the
  runner on the RTX 4090.
- Typed-decisions from the fp32 reference (eval_refs.py): the fixed head answers 1,200 of 2,000
  questions, since none of the 800 score questions has 6 levels; on those, 0.620 (choice 0.563, noul
  0.677), ECE 0.149 with no temperature. The docs say it is not comparable with full scores.

Every built-in preset has a score question with 3 or 4 levels, so Arbiter rejects them all. The
library page says so and its usage example passes its own questions; the site's generated CLI
example does the same for a model whose overlay sets noPresets. The catalog's parity text replaces
PARITY-PENDING; the README row and license note (the base is under the Gemma Terms of Use), the
family count (19) on the home page and FAQ, and the skill's model table include arbiter.
@cobanov
cobanov marked this pull request as ready for review October 6, 2026 12:08
@cobanov

cobanov commented Oct 6, 2026

Copy link
Copy Markdown
Member

@codemanhitesh, thanks for working through the review point by point. The remaining steps needed weights and a GPU, so we ran them on our side and pushed the results to this branch:

  • Merged main (0.11.0). The only conflicts were engine.rs and lib.rs, where clef and decima also register; fmt, clippy and the 251 workspace tests pass.
  • Parity, on the shared set (127 requests, 6 accepted whole, 121 rejected by both, 420 rows), with transformers 4.57.6 and peft 0.19.1 for the reference:
    • check.py: 420 rows identical, 0 mismatches, as you reported.
    • Export (parity.py): every decision the same, slot scores within 5.6e-5, probabilities within 9.3e-6.
    • Runtime (parity_arbiter): identical rows and rejections, 420/420 decisions, slot scores within 6.0e-5 on the CPU and 8.5e-5 on CUDA (RTX 4090), probabilities within 1.0e-5. A request of three or more questions takes 154 ms in the runner.
  • Typed-decisions (eval_refs.py): the head answers 1,200 of the 2,000 questions (none of the 800 score questions has 6 levels); on those, 0.620 (choice 0.563, noul 0.677).
  • Presets. Every built-in preset has a 3- or 4-level score question, so Arbiter rejects them all. The new library page says so, and its usage example passes its own questions.
  • Docs and site: the numbers in docs/families/arbiter.md, the catalog's parity text, the library page, the README row and the Gemma license note.

The unsloth/gemma-3-4b-it copy is fine with us, since you checked it is byte-identical to Google's. Marking this ready; it merges once CI is green and ships in the next release.

@cobanov cobanov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The review points are addressed and the parity gate passes on the CPU and on CUDA (see the comment above). Thanks.

@cobanov
cobanov merged commit 798a9b5 into ollaya-dev:main Oct 6, 2026
7 checks passed
@cobanov

cobanov commented Oct 6, 2026

Copy link
Copy Markdown
Member

Merged, thanks @codemanhitesh. arbiter:4b goes into the registry with the next release, together with its library page; the release notes will credit you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants