Qwen3.5-35B-A3B AWQ quant planned?

#1
by amidwestnoob - opened

Hey β€” I'm running your Qwen3-30B-A3B-Instruct-2507-AWQ on dual 3090s
with vLLM and it's been rock solid. Thank you for the clean quant.

Are you planning an AWQ quant of Qwen3.5-35B-A3B? If you are, I'll
wait for yours rather than using another quantizer. If not, no worries
β€” just want to plan my upgrade path.

Thanks!
β€” @amidwestnoob

Hi!

Not yet. The last quants I made, I made because no one else had made them at that time. ;-) Like the nemotron.

I think the cyanwiki team did already a AWQ quant. Those are pretty solid. And the are normally faster than me.

I only tested Qwen3.5-35B-A3B yet on a Spark/GB10 as FP8 and "pure" as BF16 on a dual L40 config. The last vLLM versions made trouble with AWQ under Blackwell. There a still a lot of open issues.

I could try a llm-compressor run as soon as I have free GPU resources.

Kind regards, cos

Thanks cos β€” appreciate the quick reply and the context on the Blackwell/AWQ issues. Your nemotron quant is actually the checker in a dual-brain setup on my rig β€” Qwen3.5 GGUF as the primary builder on one 3090, your nemotron AWQ as an independent fact-checker on the other. Different training lineage catches things the builder misses. Working great. I'll try cyankiwi for the 3.5 AWQ β€” currently on GGUF via llama.cpp since the AWQ had a VL config loading issue on vLLM 0.17.0. If you do end up running one later I'd be happy to test it on dual 3090s. Cheers.

Sign up or log in to comment