"How routing compute per-token rather than maintaining monolithic layer execution unlocks a 3x throughput efficiency in modern transformer architectures."
Introduction
For the past five years, the scaling paradigm of transformer models has relied on static compute budgets: every single token passes through every single layer regardless of its information entropy. Mixture-of-Depths (MoD) flips this premise by dynamically routing FLOPs based on token predictability.
The Cost of Homogeneous Layer Processing
In standard autoregressive generation, simple grammatical punctuation marks consume the same matrix multiplication budgets as complex mathematical reasoning steps. By implementing a learned router at each residual block, MoD caps the active token budget to a static threshold, allowing unselected tokens to bypass dense feed-forward computations via identity residual shortcuts.
def route_tokens(hidden_states, router_weights, capacity_limit):
scores = torch.sigmoid(router_weights(hidden_states))
top_k_indices = torch.topk(scores, k=capacity_limit, dim=1).indices
routed_tokens = torch.gather(hidden_states, 1, top_k_indices)
return routed_tokens, top_k_indices
Benchmarking Throughput and Memory Latency
In distributed cluster evaluations, MoD models achieved parity with dense baselines on MMLU while demonstrating a 42% reduction in memory bandwidth pressure during high-concurrency inference streams. This makes MoD especially potent for real-time agentic workflows.
Key Takeaways
• MoD allocates dynamic compute budgets based on token entropy instead of static homogeneous processing.
• Simple tokens bypass heavy feed-forward blocks via residual identity routing, freeing memory bandwidth.
• Achieves up to 3x throughput improvements on inference clusters without degrading downstream reasoning benchmarks.


