Request: Quantized versions (MLX / GGUF) for Apple Silicon – happy to benchmark

#1
by Blackbeard82 - opened

Hi, and thanks for releasing Kolibri-1 under Apache 2.0 — a German-focused MoE with only 3.46B active parameters is exactly what many of us have been waiting for.

Would it be possible to publish (or point us to) lower-bit quantized versions for consumer hardware? We run local inference on a 64 GB Apple Silicon Mac (M2 Max). The FP8 weights (78 GB) don't fit, but a 2.5–4-bit quant (26–39 GB) would.

What would help most:

MLX quantizations (3-bit and 4-bit, ideally keeping router and embeddings at higher precision) and the mlx-lm version that supports the architecture.
GGUF quantizations (Q3_K_M / IQ3 / Q4_K_M, a Q2 variant if quality allows) and the llama.cpp version or PR that supports Kolibri.
If available, quality numbers for the quantized versions versus FP8.

Once working builds exist, we'd be glad to run a short benchmark (speed, memory use, tool calling, German test set) and share the results here.

Aleph Alpha org

Thanks for for the amazing interest in our model! Whilst we don't have official quants yet, the amazing open source community has already created a few. You can find them here: https://huggingface.co/models?other=base_model:quantized:Aleph-Alpha/Kolibri-1

Follow-up: I measured multi-session throughput (4 sessions ≈ 84 tok/s total on an M2 Max), memory use and a KV-cache pitfall in mlx-lm on the MLX 2-bit build here: https://huggingface.co/velaia/Kolibri-1-MLX-2bit/discussions/1

Getting mlx-serve (zig kernel) working over here: https://github.com/ddalcu/mlx-serve/pull/738
no real speed gains yet, using https://huggingface.co/here-be-dragons-ai/Kolibri-1-MLX-3bit on M5 Pro 64GB.

I’ve published a new low-bit GGUF series for Kolibri-1:

Sakura-MicroQuality Kolibri-1

  • IQ2_XS: 20.94 GiB / 2.30 bpw
  • ~59% lower KL divergence than my standard IQ2_XS baseline
  • 90.5% top-token agreement against the near-lossless Q8_0 reference

I also released a separate 365E expert-pruned variant at 19.99 GiB / 74.4B parameters.

Q3 and Q4 variants are currently uploading and should follow within the next few hours.
Repo:
https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-GGUF
365E repo:
https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-365E-GGUF

Happy to share more measurements or compare results on different hardware.

Sign up or log in to comment