Model Card for Model ID

This is a mixed BF16-INT8 AWQ layer quantization, with working MTP (speculative decoding) via llmcompressor.

Model Details

📖 Model Description:

The "CK" in the name refers to "cyankiwi/calibration" dataset used for this quant.

Fixed chat_template with "froggeric/Qwen-Fixed-Chat-Templates"

Working MTP with VLLM flag:

--speculative-config '{"method":"mtp","num_speculative_tokens":2}'

Tested with VLLM 0.19.1 and transformers 5.6.2

Recommended flags:

--enable-auto-tool-choice
--reasoning-parser qwen3
--tool-call-parser qwen3_xml

🏆 Ranking: Qwopus3.5-9B-v3.5 Series (just my benchmarks):

Rank Model / Dataset HumanEval (Code) ↑ Winogrande (Logic) ↑ HellaSwag (Context) ↑ WikiText (PPL) ↓ Verdict
1st AWQ-CK (CyanKiwi) 0.6829 0.7395 0.7807 9.6310 v3.5 Savior (Best for Code)
2nd AWQ-NM (NeuralMagic) 0.6707 0.7466 0.7811 9.6322 Best for Reasoning
3rd AWQ-UC (Ultrachat) 0.6646 0.7427 0.7814 9.6325 Best for Context/Chat
4th Base (BF16) 0.6463 0.7443 0.7805 9.6306 Reference (Slow)

🔍 Key Takeaways:

  • CyanKiwi (CK): A game-changer for the v3.5 series. It successfully recovered the coding performance, jumping from 0.64 (Base) to 0.68, making it the superior version for technical tasks.
  • NeuralMagic (NM): Demonstrated the best logical consistency in Winogrande, surpassing the base model.
  • Ultrachat200k (UC): While showing slightly higher perplexity, it led in contextual understanding (HellaSwag).
  • Efficiency: Quantized versions are not only faster but actually outperform the Base BF16 model in almost every technical metric.

🙏 Acknowledgements:

Downloads last month
57
Safetensors
Model size
10B params
Tensor type
I64
·
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Arien0/Qwopus3.5-9B-v3.5-AWQ-BF16-INT8-CK-MTP

Finetuned
Qwen/Qwen3.5-9B
Quantized
(11)
this model

Dataset used to train Arien0/Qwopus3.5-9B-v3.5-AWQ-BF16-INT8-CK-MTP