Show and Tell: bit-jev — structured decisions (option scoring, no answer decoding) on a BitNet backbone, exported as I2_S GGUF with a native CPU runner #634
Zeaulo (Zeaulo)
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi all,
I'm sharing bit-jev, a small project that uses a BitNet b1.58 backbone for structured decisions rather than text generation. The caller declares questions and candidate options. A pointer head reads the hidden states and scores the options directly, so no answer tokens are generated. It supports three question types:
choice,noul(yes/no), andscore(ordinal).This is an independent project. It is not affiliated with or endorsed by Microsoft. The backbone and the native inference base come from microsoft/BitNet.
How it relates to bitnet.cpp
head.f32).test/bootstrap_bitnet.pyfetches pinned BitNet / llama.cpp commits and applies a ReLU² compatibility patch. A small native runner (core/native/main.cpp) then loads the GGUF backbone and scores the options with the f32 head.Numbers (single-question case, with caveats)
There is one fixed dev question: 703 input tokens, 77 candidate options, native I2_S CPU path. Native latency excludes model loading.
An EPYC 16-thread check on the same host averaged about 3,210.12 ms. The EPYC build used g++ 11.4.0, CMake 3.31.10, Release,
GGML_NATIVE=ON, andGGML_OPENMP=ON. The per-run timings, source commit, and model SHA-256 are in docs/benchmark-data. This is a single small dev case. It does not represent general latency or accuracy, and the input text is not public.As a separate check, I also measured the public microsoft/bitnet-b1.58-2B-4T-gguf I2_S file at a pinned older revision. I chose that revision because of the reports in #608. With
llama-bench -p 128 -n 32 -r 5 -t 8 -b 128 -ub 128on a Ryzen 7 4800H (8 threads, CPU-only), the median prefill was 56.23 tok/s, the median decode was 4.57 tok/s, and peak RSS was 1,230.19 MiB. The commands, the compatibility patch, and the per-run samples are in docs/BENCHMARKS.md.Try it
pip install bit-jev, thenpython -m bit_jev.demo(local Gradio page). The wheel ships a prebuilt CPU runner.Limits
Where help would be useful
--batch, build flags, andlatency_ms.llama-benchmeasurement on other CPUs, so the I2_S numbers have more reference points.run_inference.py.Thanks for BitNet. Feedback and reproduction reports are welcome.
All reactions