123b1f2479
Per-token cost of the batched head forward keeps falling until the flush is large enough to reach the fastest kernels: NAX matmul tiles for dense heads, and the segmented gather path for MoE heads, which needs tokens*topK/experts >= 4. Measured across the qwen3.6 heads, 256 is the smallest cap past every threshold and within a few percent of each head's per-token floor. The cost is bounded: up to 2.5 MiB of pinned hiddens per request and a flush stall under one decode step.