Articles on models,
methods, and
making things work.

A home for technical writing, work in progress, and the occasional useful rabbit hole.

Divyansh
Agrawal

Research and engineering notes on LLM inference, attention mechanisms, agent systems, memory, evaluation, and efficient model serving.

SELECTED BLOG

ALL POSTS
06

The Tiny Draft Model Hidden Inside Qwen

Qwen ships a one-layer multi-token predictor that vLLM and SGLang can use for native speculative decoding.

04

Do LLM Agent Societies Adapt Their Values, or Eventually Die by Them?

What an evolutionary multi-agent simulation reveals about cultural persistence, selection, migration, and the danger of values that never bend.

03

Consensus Is Not Corroboration

Why language models need to distinguish a manufactured echo from genuinely independent confirmation.

05

Why Qwen Rotates Only a Quarter of Each Attention Head

Partial RoPE gives Qwen a compact relative-position channel alongside a larger RoPE-free similarity subspace.

02

Attention Should Be Allowed to Say No

Qwen's Gated Attention separates where an attention head reads from whether its output should influence the model.

01

The Case for Planner–Executor Architectures in Agentic Coding

Why serious coding agents should separate judgment from execution—and spend frontier-model capability where it matters most.

A LITTLE CONTEXT

I'm interested in the distance between a capable model and a useful system.

I'm Divyansh Agrawal, an AI/ML researcher and builder. I care about how learning systems reason, retrieve, evaluate themselves, and earn trust in the hands of real people.

This is a home for work in progress: ideas that are not yet papers, systems that are not yet products, and the questions that sit between them.

Say hello