M.01Open-weight · MIT
GLM-5.2 · colibri int4-g64Grouped int4 · int8 MTP head
Grouped int4 experts (one scale per 64 weights) with an int8 multi-token-prediction head for speculative decoding. Validated token-exact against the transformers reference, and the engine's reference grouped-int4 container.
87.0%HellaSwag acc_norm, against 83.5% for per-row int4 (n=200)
- Base
- GLM-5.2 · 744B MoE
- Size
- 429 GB
- Downloads
- 23,000+
- Licence
- MIT
- Built with
- Claude Code
M.02Open-weight · MIT · experimental
GLM-5.2 · colibri E8/IQ33.06 bits per weight · int8 MTP head
The same weights, converted from the FP8 parent: a third smaller, and 22–33% faster on hosts that stream experts from disk, with no measurable quality loss across HellaSwag, ARC-Challenge and MMLU. Slower when every expert already fits in memory — the model card says so.
1.57tokens/s decode on a single 16 GB consumer GPU, against 1.18 for the int4-g64 build
- Base
- GLM-5.2 · 744B MoE
- Size
- 289 GB
- Downloads
- 1,400+
- Licence
- MIT
- Built with
- Claude Code
M.03Open-weight · Apache-2.0 · Arabic
ALLaM-7B · Arabic SFTQLoRA fine-tune · bf16 + GGUF builds
An Arabic fine-tune of ALLaM-7B-Instruct, trained on 39,576 Arabic conversations on a single consumer GPU in 10.2 hours. Modest gains on Arabic knowledge benchmarks, with GGUF builds for llama.cpp, LM Studio and Ollama.
49.2%ArabicMMLU (0-shot), against 47.8% for the ALLaM-7B base. Arab-culture score (ACVA) 76.5 vs 77.6, within noise. All models run in 4-bit, so absolute scores sit below leaderboard figures.
The first version gained knowledge but lost 6 points on Arab-culture questions, because the translated training data crowded out Arab-specific knowledge. Applying the learned changes at half strength kept most of the gain and recovered the culture score, with no retraining.
- Base
- ALLaM-7B-Instruct-preview
- Recommended
- Q5_K_M · 5.0 GB · 149 tok/s
- Fastest
- Q4_K_M · 4.3 GB · 168 tok/s
- Licence
- Apache-2.0
- Built with
- Claude Code