ARCHITECTURE Signal 503
DFlash 2 yields 16 to 25% more tokens per verification pass at ~1% latency cost
DFlash 2 improves parallel speculative decoding by better selecting from existing candidate lists, delivering 16 to 25% more output per verification pass with roughly 1% added cycle latency and provably unchanged output.
Inference cost dominates the economics of agent workloads, which consume tokens for hours. A throughput gain this large at negligible latency cost directly reduces serving cost for production LLM endpoints. The approach is already integrated into four mainstream inference engines, so adoption is a config change, not a rewrite.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
DFlash 2 produces 16 to 25% more output per verification pass for around 1% added cycle latency, with output provably unchanged.
With the released Qwen3.8-27B drafter, SGLang serves at 2.7 to 3.4× the throughput of autoregressive decoding at batch size 1.
DFlash 2 is available now in SGLang, vLLM, llama.cpp, and oMLX with the incoai/Qwen3.8-27B-DFlash2 draft model.
THE CLUSTER
↗