The Tiny Draft Model Hidden Inside Qwen
Qwen ships a one-layer multi-token predictor that vLLM and SGLang can use for native speculative decoding.
A home for technical writing, work in progress, and the occasional useful rabbit hole.
Divyansh
Agrawal
Research and engineering notes on LLM inference, attention mechanisms, agent systems, memory, evaluation, and efficient model serving.
SELECTED BLOG
ALL POSTS ↗Qwen ships a one-layer multi-token predictor that vLLM and SGLang can use for native speculative decoding.
What an evolutionary multi-agent simulation reveals about cultural persistence, selection, migration, and the danger of values that never bend.
Why language models need to distinguish a manufactured echo from genuinely independent confirmation.
Partial RoPE gives Qwen a compact relative-position channel alongside a larger RoPE-free similarity subspace.
Qwen's Gated Attention separates where an attention head reads from whether its output should influence the model.
Why serious coding agents should separate judgment from execution—and spend frontier-model capability where it matters most.
CURRENT & PAST WORK
Small research systems with an opinion about the future.
A synthetic social web for testing whether AI agents can distinguish correlated repetition from independent evidence.
An ultra-low-latency LLM gateway for routing, caching, budgets, analytics, and forecasting.
Plug-and-play persistent memory and recall for LLM applications.
A LITTLE CONTEXT
I'm Divyansh Agrawal, an AI/ML researcher and builder. I care about how learning systems reason, retrieve, evaluate themselves, and earn trust in the hands of real people.
This is a home for work in progress: ideas that are not yet papers, systems that are not yet products, and the questions that sit between them.
Say hello ↗