← All series
ESSAY SERIES
Qwen Hybrid Architecture
3 essays. Start with the first piece or jump to the question you came for.
03essays
- 01
Attention Should Be Allowed to Say No
Qwen's Gated Attention separates where an attention head reads from whether its output should influence the model.
7 min read ↗ - 02
Why Qwen Rotates Only a Quarter of Each Attention Head
Partial RoPE gives Qwen a compact relative-position channel alongside a larger RoPE-free similarity subspace.
6 min read ↗ - 03
The Tiny Draft Model Hidden Inside Qwen
Qwen ships a one-layer multi-token predictor that vLLM and SGLang can use for native speculative decoding.
6 min read ↗