GELab-Zero-4B-preview-Sico-Evolution-MLX-mxfp8

An MLX MXFP8 quantization (microscaling FP8, group_size = 32) of microsoft/GELab-Zero-4B-preview-Sico-Evolution, a Qwen3-VL vision-language model, converted from the original bf16 weights for Apple Silicon.

  • Base model: microsoft/GELab-Zero-4B-preview-Sico-Evolution
  • Architecture: Qwen3VLForConditionalGeneration (qwen3_vl) — vision-language
  • Quantization: MLX MXFP8 — mode=mxfp8, 8-bit, group size 32 (~8.98 bits/weight incl. block scales)
  • Size on disk: ~4.7 GB

MXFP8 keeps a per-32-element E8M0 microscale, which better preserves dynamic range than affine int8 at a similar footprint.

Use with mlx-vlm

pip install mlx-vlm
python -m mlx_vlm generate \
  --model unigilby/GELab-Zero-4B-preview-Sico-Evolution-MLX-mxfp8 \
  --prompt "Describe this image." \
  --image path/to/image.png

Text-only prompting works as well.

License

Inherits the base model's Apache-2.0 license.

Downloads last month
14
Safetensors
Model size
2B params
Tensor type
U8
·
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for unigilby/GELab-Zero-4B-preview-Sico-Evolution-MLX-mxfp8

Quantized
(5)
this model