Repository navigation
Add the arbiter family (hiteshluke/arbiter-4b) - #49
Conversation
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…t head) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
… slot scores) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…erance 1e-4) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…ensing) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…ding export) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Pins the upstream Hugging Face shas in ref.MODELS, catalog.py and the registry/v2 arbiter manifests: base unsloth/gemma-3-4b-it bf46152c47f5dd20b896357cb51abc4c03b8ee8c model hiteshluke/arbiter-4b 0c44271c59f89758e3cae17b032e98a9140093e9 ONNX export was not run in this environment (torch.export on the full Gemma 3 4B causal LM is multi-hour and memory-heavy, and the pinned convert/ deps still have to resolve from source), so the sha256:ARBITER_*_PLACEHOLDER digests in the two manifests are left as placeholders for a maintainer to fill after running `python -m ollaya_convert.families.arbiter.export arbiter-4b --out out/arbiter-4b` against the pinned revisions. The parity gate (python -m ollaya_convert.families.arbiter.parity) is deferred to the same step for the same reason. Also scrubs incidental cross-family references in the arbiter sources so the new family stands on its own. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Hash both base shards from unsloth/gemma-3-4b-it @ bf46152c, and the
adapter, head.pt and tokenizer from hiteshluke/arbiter-4b @ 0c44271c,
replacing the ARBITER_WEIGHTS_PLACEHOLDER_{1..4} and
ARBITER_TOKENIZER_PLACEHOLDER entries in registry/v2/library/arbiter/
manifests/{4b,latest}. The ONNX graph digest and the ollaya-hosted
config/decision/calibration/license blobs are produced by the export +
packaging step on GPU and stay as placeholders until then.
scripts/arbiter_hash_hf.py downloads each file at its pinned revision
with huggingface_hub, streams it through hashlib.sha256, prints the
JSON report and deletes the local copy; rerun it to verify.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Rewrite the "What maintainers need to finish" section into a split "Done on the fork" / "Left for a maintainer with sufficient GPU" and point at the explicit uv command plus the parity gate. Fill in the pinned base and adapter SHAs in the family doc's status row and rephrase the parity note so it says what still needs to run. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
cobanov
left a comment
There was a problem hiding this comment.
@codemanhitesh, thanks for submitting Arbiter, and sorry for the wait. The Python side is a good start. Every family in Ollaya goes through the same path before it ships, and this PR is not there yet. Here is what it needs, in order.
1. The Rust runtime. Nothing runs this model today. arbiter-fixed-v1 needs:
- a layout in
crates/ollaya-decisionthat builds the prompt rows from a request (portlayout.py); - an engine in
crates/ollaya-runnerthat feeds the graph and reads the 24 slots per question type, registered inengine.rs.
kev.rs and jeeves.rs in both crates are close templates (decoder rows, last-position readout). Requests the fixed head cannot answer must be rejected with a clear 422, not truncated: a choice with more than 16 options, a score with other than 6 levels.
2. Parity, which is the gate.
- Prompt check. A
check.pythat compares the layout port with your reference prompt, id for id, on our shared request set: the edge cases and 40 typed-decisions rows, as a JSONL here. - Goldens.
goldens.py, written from your reference model in fp32 on those requests: token ids, slot scores and probabilities. - Runtime parity. A
crates/ollaya-runner/examples/parity_arbiter.rsthat runs the Rust runtime against those goldens: identical rows and rejections, every decision the same, scores within 1e-3.parity_jeeves.rsis a template.
We run the export, the goldens and the parity again on our own GPU before merging, so they only need to be runnable.
3. Leave the registry to us. Please remove registry/v2/library/arbiter/manifests/*. We write manifests with package.py when a tag ships, from the export we checked, so a manifest never carries placeholder digests.
4. Licensing.
- License field. The base is Gemma 3, under the Gemma Terms of Use, so the catalog's
"license"cannot sayApache-2.0alone. Yourlicense_textalready names both; the field and the docs should match it. - Base repository. Ollaya pulls weights from the original author's repository. Please say whether
unsloth/gemma-3-4b-it@bf46152is byte-identical to Google'sgoogle/gemma-3-4b-it, and why the manifest should point at the copy.
5. Repository conventions.
- Docs. Docs for a family live in
docs/families/arbiter.md: sequence, graph, differences from upstream, parity, quality. Please folddocs/pr/arbiter.mdinto it or into this PR's description. - Hashing script. One-off tools like
scripts/arbiter_hash_hf.pystay out of the repository.package.pyalready verifies every file against the Hub's sha256. - Accuracy. Please add typed-decisions accuracy (all 400 test states, argmax against the majority label), the number every model in the library is compared on. BoolQ and ARC can stay next to it.
Once 1 and 2 are in, we will do the GPU checks and help with the rest. Thanks again for the contribution.
Manifests are written by package.py when a tag ships, from the checked export, and package.py already verifies every file against the Hub's sha256. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… and parity on the shared set - ref.py: the training script's three templates verbatim, tokenized by transformers as in training (with <bos>); Gemma3ForConditionalGeneration in fp32 + peft adapter + the bias-free 24-slot head, one unpadded row at a time. - layout.py: the port the Rust runtime follows ([bos] + tokenizers, slots from decision.json, noul read as [false, true]); rejects a choice over 16 options, a score with other than 6 levels and a row over 8,192 tokens instead of truncating. - check.py: port vs reference prompt id for id on the shared request set (JSONL) plus arbiter edge cases: 127 requests, 420 identical rows, 0 mismatches. - goldens.py: token ids, all 24 slot scores, option logits and probabilities per question. - export.py: Gemma 3 trunk (llm_common/gemma3.py, sliding-window layers) with the LoRA unmerged; tokenizer.json from the base repository. - parity.py: ONNX Runtime vs the fp32 reference on the same requests. - eval_refs.py: typed-decisions quality for arbiter, with answered/total coverage. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…iter - ollaya-decision/arbiter.rs: the training prompt after <bos>, one causal row per question, slots from decision.json. A choice over 16 options is TOO_MANY_OPTIONS, a score with other than 6 levels or a row over max_row_tokens is INVALID_REQUEST (422), never truncated. - ollaya-runner/arbiter.rs: feeds input_ids + last_pos, reads the 24 slot scores and returns each question's option logits (noul as [false, true]); registered in engine.rs. - examples/parity_arbiter.rs: identical rows and rejections against the goldens, all 24 slot scores within 1e-3, the same decision on every question. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ilies/arbiter.md - catalog: license is "Apache-2.0 (LoRA adapter and head) and the Gemma Terms of Use (Gemma 3 base model)"; the license text adds the Gemma notice and policy links; the tokenizer comes from the base repository (the model repository's tokenizer.json has truncation to 255 tokens on). - docs/families/arbiter.md: sequence, graph, differences from upstream, parity, quality, limits; the base repository's files compared with google/gemma-3-4b-it by Hub sha256. Typed-decisions accuracy is marked TODO until it is measured. - docs/pr/arbiter.md removed. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…it ones Same adapter, head and rows on both bases: equal or better on BoolQ, ARC-Challenge, CommonsenseQA and OpenBookQA. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…y page The GPU steps of the parity procedure, run by the maintainers on the RTX 4090 machine (transformers 4.57.6 and peft 0.19.1 for the reference), on the shared set: 127 requests, 6 accepted whole, 121 rejected by both, 420 rows. - check.py: 420 rows identical, 0 mismatches, as reported in the PR. - Export against the reference (parity.py): every decision the same, slot scores within 5.6e-5, probabilities within 9.3e-6. - Runtime (parity_arbiter): identical rows and rejections, 420/420 decisions; slot scores within 6.0e-5 on the x86-64 CPU and 8.5e-5 on CUDA (Microsoft ONNX Runtime 1.28.2 from the CUDA 13 pack), probabilities within 1.0e-5. A request of three or more questions takes 154 ms at the median in the runner on the RTX 4090. - Typed-decisions from the fp32 reference (eval_refs.py): the fixed head answers 1,200 of 2,000 questions, since none of the 800 score questions has 6 levels; on those, 0.620 (choice 0.563, noul 0.677), ECE 0.149 with no temperature. The docs say it is not comparable with full scores. Every built-in preset has a score question with 3 or 4 levels, so Arbiter rejects them all. The library page says so and its usage example passes its own questions; the site's generated CLI example does the same for a model whose overlay sets noPresets. The catalog's parity text replaces PARITY-PENDING; the README row and license note (the base is under the Gemma Terms of Use), the family count (19) on the home page and FAQ, and the skill's model table include arbiter.
|
@codemanhitesh, thanks for working through the review point by point. The remaining steps needed weights and a GPU, so we ran them on our side and pushed the results to this branch:
The |
cobanov
left a comment
There was a problem hiding this comment.
The review points are addressed and the parity gate passes on the CPU and on CUDA (see the comment above). Thanks.
|
Merged, thanks @codemanhitesh. |
Add the
arbiterfamily (hiteshluke/arbiter-4b)Arbiter is Codekins Pvt Ltd / Zyot Lab's 4B decision model: Gemma 3 4B IT + LoRA + a fixed 24-slot pointer head
(
arbiter-fixed-v1). One forward pass per question;noulreads slots 0..1,choiceslots 2..17 (up to 16options),
scoreslots 18..23 (exactly 6 levels). Full write-up:docs/families/arbiter.md.This update addresses the review point by point.
Review items
crates/ollaya-decision/src/arbiter.rs(port oflayout.py, with unit tests);engine in
crates/ollaya-runner/src/arbiter.rs, registered inengine.rs. A choice with more than 16 optionsand a score with other than 6 levels are rejected with a 422, never truncated.
check.py: the Rust-layout port against the reference prompt, id for id, on the shared request set plus 6extra Arbiter cases, with two tokenizer library versions: 420 rows identical, 0 mismatches.
goldens.py: the reference model in fp32 (token ids, last position, all 24 slot scores, option logits,probabilities).
crates/ollaya-runner/examples/parity_arbiter.rs: identical rows and rejections, every decision the same,slot scores within 1e-3.
registry/v2/library/arbiter/manifests/*removed.licensefield now reads "Apache-2.0 (LoRA adapter and head) and the Gemma Terms ofUse (Gemma 3 base model)"; docs match.
unsloth/gemma-3-4b-it@bf46152is byte-identical to Google'sgoogle/gemma-3-4b-itfor both safetensors shards andtokenizer.json(same sha256 and size per the Hub API);only metadata files differ. The original is gated (manual approval) and Ollaya's client has no Hugging Face
token support, which is why the manifest points at the copy. Happy to repoint if you prefer.
docs/pr/arbiter.mdfolded intodocs/families/arbiter.md;scripts/arbiter_hash_hf.pyremoved.
Found while porting
<bos>; the earlier layout did not. Fixed.tokenizer.jsontruncates at 255 tokens, so the runtime takes the tokenizer from the base repo.Accuracy
Measured on a T4 with the training prompt, one row at a time. Model-card numbers use the 4-bit base the adapter was
trained on; the same adapter and head were also run on the same rows with the unquantized base (fp32 compute, as
Ollaya runs it):
Running the full BF16 base does not cost accuracy relative to the model card.
Status
parity numbers yet; and typed-decisions accuracy on the 400 test states. Both are runnable with the commands
in the docs. The pipeline (export, goldens, ONNX parity) was exercised end to end on a small random stand-in
model: ONNX within about 1e-5 of the reference, all decisions agree.
please treat CI as the first check; we will fix anything it reports.
6-level head cannot answer them and they are rejected with a 422.