Flow-Drafter-Qwen3.5-27B-v2

Expanded-data revision of Flow-Drafter-Qwen3.5-27B β€” a speculative-decoding draft model for Qwen/Qwen3.5-27B (Chained-Flow, joint-VAE drafter).

What changed in v2: trained on a broader, diversity-weighted mix (~30k rows, 10 sources) adding multi-turn chat (UltraChat), diverse instructions (No-Robots), creative prose (WritingPrompts), summarization (CNN/DailyMail) and translation (OPUS) on top of the original math/code/STEM core, to lift acceptance on the low-predictability domains.

Lossless by construction: a proposed token is emitted only if the target itself would have emitted it. Outputs can still differ from unspeculated decode, because verifying K+1 positions at once changes the target's own fp16 reduction order β€” the observed divergence is comparable to the run-to-run difference between two unspeculated runs.

Usage

pip install chained-flow

CF_DRAFTER_DIR=selimaktas/Flow-Drafter-Qwen3.5-27B-v2 \
vllm serve Qwen/Qwen3.5-27B --async-scheduling --max-num-seqs 64 \
  --speculative-config '{"method":"custom_class","model":"chained_flow.vllm_plugin.flow_proposer.FlowDrafterProposer","num_speculative_tokens":5}'

Measured under vLLM β€” acceptance and speedup, per domain

Condition: batch 1 (concurrency 1), chain proposer at K=5, greedy, fixed 256 output tokens per request (ignore_eos) so both arms do identical work, async scheduling on for every arm including the baseline, 3 repeats (pooled spread ≀0.1%). Benchmark: RedHatAI/speculator_benchmarks, 8 distinct domains Γ— 25 prompts, one run per domain, never pooled into a single split. Harness: guidellm 0.6.0 driving vllm serve (vLLM 0.25.1). acceptance = 1 + num_accepted_tokens / num_drafts from vLLM's own Prometheus counters (bonus token included, so a non-speculative baseline is 1.00). speedup = chain tok/s Γ· base tok/s, prefill included in the denominator.

domain acceptance (tokens/step) speedup vs base
HumanEval (code) 2.42 2.00x
math_reasoning 3.05 2.51x
qa (short free-form) 2.00 1.67x
question (MT-bench) 2.09 1.74x
rag 2.13 1.74x
summarization 2.07 1.70x
tool_call 2.13 1.75x
translation (de->en) 1.57 1.31x
POOLED (all 8 domains) 2.12 1.75x

Quote the pooled figure, not the best domain. The spread runs from 2.51x on math_reasoning down to 1.31x on translation. Unlike the 4B and 9B siblings, no domain is a regression here β€” the weakest still returns 1.31x. The pooled 1.75x is the number to plan with.

At 27B the base decode step is expensive enough that even the weakest domain pays for its draft pass. This is the size where speculative decoding is easiest to justify.

CF_SPEC_MAX_BATCH ships on by default and resolves to 16 here, disengaging speculation above that decode batch and holding a loaded server near parity rather than below it.

Files

  • model.safetensors β€” the drafter (flow experts, Markov head, path head, jointly-trained VAE)
  • chained_flow_tree_config.json β€” architecture + loss config
  • vae/ β€” the base VAE checkpoint

Drafts all K future hidden states in one flow pass and the base model verifies them. Trained 2-GPU DDP.

Companion models: Flow-Drafter-4B-v2, Flow-Drafter-9B-v2.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
2B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for selimaktas/Flow-Drafter-Qwen3.5-27B-v2

Base model

Qwen/Qwen3.5-27B
Finetuned
(299)
this model