← All series

ESSAY SERIES

Qwen Hybrid Architecture

3 essays. Start with the first piece or jump to the question you came for.

03essays
  1. 01

    Attention Should Be Allowed to Say No

    Qwen's Gated Attention separates where an attention head reads from whether its output should influence the model.

    7 min read
  2. 02

    Why Qwen Rotates Only a Quarter of Each Attention Head

    Partial RoPE gives Qwen a compact relative-position channel alongside a larger RoPE-free similarity subspace.

    6 min read
  3. 03

    The Tiny Draft Model Hidden Inside Qwen

    Qwen ships a one-layer multi-token predictor that vLLM and SGLang can use for native speculative decoding.

    6 min read