Pit NeitemeierAlessio Serra
Designing Kolibri: Architecture Trade-Offs from First Principles
TL;DR
This blogpost investigates the impact of the main architectural choices in autoregressive language models on training and deployment cost, focusing on the allocation of parameters and FLOPs, how FLOPs scale with context length and sequence-mixer1 state2 size. We derive these quantities and analyse how they vary across recent open-weight architectures. It also includes an interactive tool to configure your own model and compare it with recent open-weight releases.
Introduction
A language model’s architecture decides how its capacity and computation are distributed. Many of these choices affect quality, training cost and serving cost at the same time. The effect of a choice on quality has to be measured empirically, but the costs follow directly from the configuration:
- Total parameters set the weight-memory footprint, and with it the minimum hardware needed to deploy the model.
- FLOPs per token, multiplied by the training-token horizon,3 set training compute. They are also a useful proxy for prefill cost, which is typically compute-bound.
- Sequence-mixer state2, e.g. the key-value (KV) cache for an attention-based sequence mixer, sets how many sequences, and how much context, fit in a decode batch, and how many bytes each decode step has to read.
Which of these becomes the serving bottleneck depends on the workload. Small-batch, short-context decoding is limited by reading the weights. Prefill and large-batch decoding of short queries are limited by compute. Long-context decoding is limited by reading the sequence-mixer state.
In this blogpost we derive each quantity for MoE decoders and analyse recent open-weight models:
- Parameter allocation: how active parameters are allocated between the FFN (a Mixture of Experts (MoE) layer in all models we analyse here), the sequence-mixer (e.g. attention) and the Language Modelling (LM) head.
- FLOP allocation: how FLOPs per token are allocated between components and scale with sequence length.
- Sequence-mixer state: how much state each sequence holds and each decode step reads, depending on sequence length and sequence-mixer type.
We visualise these quantities in an interactive explorer below that lets you change the configuration and see the effect on all three axes.
Architecture Configurations
The explorer below is a simplified subset of an internal tool we use when designing model architectures. It helps us quantify the trade-offs between different architectural choices before committing to a configuration. Inspired by Sebastian Raschka’s LLM Architecture Gallery, we compare recent open-weight architectures. We initialise the explorer with two Aleph Alpha models, Kolibri Origin and Kolibri, the latter of which was recently released under the Apache 2.0 licence.4 These provide two concrete reference points for the design choices discussed throughout the post. You can modify either configuration using the sliders or add your own with +.
- Total
- 78.1B
- Active / token
- 3.5B
- FLOPs / token · 16K
- 28.3G
- State · 16K
- 189M
Parameter Allocation
In MoE architectures, total and active parameters can differ by an order of magnitude. Total parameters mainly affect the memory footprint, hence the minimum deployment requirements of a model. Active parameters typically account for most of the FLOPs during training and prefill. At small batch sizes, they also largely determine how many weight bytes each decode step reads from memory.5
If you play with the sliders, you’ll notice that every slider except the full-attention/SWA split and the window size affects total or active parameters. We now derive each component’s contribution to the two totals.
Total and Active Parameters per Model
| 78.1B | 3.5B | 4.4% | |
| 30.6B | 3.3B | 10.7% | |
| 31.6B | 3.2B | 10.2% | |
| 120.7B | 12.2B | 10.1% | |
| 549.3B | 55.0B | 10.0% | |
| 125.1B | 6.0B | 4.8% | |
| 1.57T | 48.8B | 3.1% | |
| 551.5B | 15.8B | 2.9% | |
| 2.78T | 104.2B | 3.7% | |
| 308.8B | 14.8B | 4.8% | |
| 116.8B | 5.1B | 4.4% | |
| 20.9B | 3.6B | 17.3% | |
| 743.4B | 40.3B | 5.4% | |
| 426.2B | 24.7B | 5.8% |
Total counts every stored weight of the text decoder. Active / token counts the weights each token passes through: sequence mixer, the top-k routed and shared experts, dense FFN layers, router, norms and LM head; the input embedding is a lookup and is left out.
FFN: Dense Layers, Routed and Shared Experts
Routed expert count and expert width determine the capacity of the FFN. Top- and shared experts determine how much of that capacity is active for one token.
A SwiGLU expert uses three linear projections: gate, up and down. The gate and up projections map from the hidden dimension to the intermediate dimension , while down projection maps back from to . This gives parameters per expert. Multiplying this number by the number of experts and the number of MoE layers gives the routed total parameters. For the active routed parameters we substitute with , since only the top- experts will be used.
In the same way we can derive the shared experts parameters. Note that multiple shared experts are equivalent to a single shared expert of larger width , since all shared experts are always active and their outputs are combined. So we only consider a single shared expert here.
Many MoE decoders keep their first layers dense6. They can have their own width . In the sliders above is tied to the active expert width, so , which is a common choice among the models analysed here (e.g. MiMo-V2.5, MiniMax M3, Kolibri Origin).
The router maps the full residual stream to logits, so its size grows with the expert count and hidden size. However, it is a very small part of total parameters, and we count it with the FFN.
Sequence Mixer: Attention Projections
The Query, Key, Value and Output projections of one layer depend on the hidden width, the query width and the KV width. Their total then scales with the number of attention layers.
In a GQA layer we project the hidden state of width into a query of width and two KV tensors of width , then project the attention result from back to .
With the group size , MHA () and MQA (, so ) are the two edge cases of GQA.
Norms
Each decoder layer commonly carries two (only pre-norm) to four (sandwich norm) RMSNorms, with scale vectors of width , plus one final norm before the LM head. The factor four below assumes sandwich norms.
Embedding and LM Head
The product of vocabulary size and hidden width sets both the LM head and the embedding. The latter only affects the total parameters, since it determines only a light lookup operation instead of a matrix multiplication.
An untied input embedding and LM head each contain weights. Tying them stores one shared matrix instead of two, but does not remove the dense vocabulary projection from training compute. You can verify how only the total parameters change using the embedding-tying toggle.
Active Parameter Composition per Token
Parameter Allocation Across Architectures
As Figure 2 shows, the fraction of active parameters allocated to the sequence mixer varies significantly across the selected architectures, from roughly 18% to 49%.
Sorting the models by total parameters, active parameters or the active-to-total ratio reveals no obvious relationship with the sequence-mixer share. The choice of sequence-mixer share comes primarily down to empirically measured quality. Additionally, increasing the sequence-mixer compute while keeping the state size constant increases arithmetic intensity and makes the primarily memory bound kernels more efficient.
The ratio between sequence mixer and FFN parameters is mainly controlled by the relative widths of the attention projections and the intermediate active width of the FFN. Using the parameter counts derived above,
so their relative allocation is
The residual-stream width cancels out. Changing alone therefore leaves this split unchanged, while changing the attention or expert width directly shifts the allocation.
Finally, the LM head’s share shrinks as models grow. Most weight matrices grow quadratically with the width, such as with , while the LM head () grows only linearly, because the vocabulary size usually stays fixed. You can see this clearly with the LM head stacked first and the models sorted by active parameters.
FLOP Allocation
Training compute per token has two main components: parameter FLOPs and attention FLOPs.
Parameter FLOPs come from multiplying activations with weight matrices and are independent of the context length. Attention FLOPs come from the and products, which involve no learned weight matrix but depend on the attention width and on the context length.
The equations in this section reuse the parameter-allocation symbols defined above.
The parameter-FLOPs term depends only on the active weights: the 6 comes from three matrix multiplications of equal size, each costing two FLOPs per weight per token, one multiply and one add. The forward pass () computes . Backpropagation then needs two more products of the same shape. The gradient with respect to the activations () propagates the error signal back into the preceding layer, while the gradient with respect to the weights () measures how the loss depends on each weight and how it should be updated to reduce the loss.
Embedding lookup, norms and activations are not considered by the previous equation since, at this scale, their contribution is negligible.7
Attention FLOPs involve no weight matrices: computes one dot product per token and applies one value vector per token. This results in FLOPs per query-key pair, and the same multiplier accounts for , and . We count the keys a query attends under the causal mask, on average in a full-attention layer. FLOPs grow with the sequence length because the query attends every retained key, which is all previous tokens in a full-attention layer and at most the window in a SWA layer. Note that GQA leaves and the attention FLOPs unchanged: it only narrows the KV width, which reduces the sequence-mixer state size, discussed in the next section.
Beyond Full and Sliding-Window Attention
The closed-form derivations below focus on standard and sliding-window attention, but several architectures in our comparison use other sequence mixers. We only give the intuition of these alternatives here and refer interested readers to the original papers for their detailed formulations.
At a high level, these alternatives follow two main approaches. Sparse attention, such as Native Sparse Attention (NSA), reduces the number of tokens accessed from a longer context. Recurrent or linear sequence mixers, such as Mamba-2 and Gated DeltaNet, instead compress information from the prefix into a fixed-size state. Recent models often combine these mechanisms with attention in hybrid architectures. In our comparison, DeepSeek V4, GLM-5.3, MiniMax M3 and Qwen3.8-Flash-Next use sparse attention, while Qwen3.8-Flash-Next (Gated DeltaNet), Kimi K3 (Kimi Delta Attention) and Nemotron 3 (Mamba-2) combine attention with recurrent layers.
FFN
The FFN transforms each token independently and, as we have seen before, is where most of the capacity is spent. At fixed top-, inactive routed expert matrices add capacity that costs memory footprint but no FLOPs.
Each MoE layer adds one small router matrix on top.
Sequence-Mixer Projections
The sequence mixer models interaction between tokens. Its Query (Q), Key (K), Value (V) and Output (O) linears depend on hidden width, attention width, KV width and depth, but are independent of sequence length per token.
Pure Attention
It includes the attention score computation and the value aggregation , which depend on query width and retained keys. A full-attention layer retains every previous token, so a query attends keys on average.
A sliding-window layer retains at most of them.
LM Head
The output projection scales with hidden width times vocabulary size.
Training FLOP Composition at 16K
FLOP Allocation vs. Context Length
Adjusting the sequence-length slider in Figure 3, you can see how the FLOP composition changes significantly with context length. At 16K context, pure attention accounts for roughly 3–51% of training FLOPs across the selected models. As context approaches 1M tokens, this contribution can become dominant, reaching 98% for .
Parameter-bound FLOPs grow linearly with the number of tokens, while full-attention FLOPs grow quadratically with sequence length. As a result, the FLOP mix shifts increasingly toward pure attention as context grows, especially in architectures with many full-attention layers.
Across models, the total fraction of FLOPs spent in the sequence mixer at 16K sequence length is more constrained, ranging from roughly 33% to 65%. As pure-attention FLOPs are reduced through e.g. sparse attention, projection linears account for a larger fraction of total FLOPs. follows this pattern by reducing the number of full-attention layers and replacing them with sliding window attention, substantially lowering the pure-attention share from 51% to 27% at 16K sequence length compared to .
Training FLOPs Over Sequence Length
FLOP Scaling vs. Context Length
As Figure 4 shows, differences in compute cost become increasingly pronounced as context length grows. has the highest FLOPs across the full range, combining a large active model size with 24 full-attention layers out of 93. sits at the opposite end at long context, with only a small fraction of its sequence mixer scaling quadratically with context length.
and need similar FLOPs at 16K context, but at 1M tokens MiMo-V2.5 needs about 5× more: its full-attention layers attend to every previous token, while DeepSeek V4.1-Flash attends to a fixed number of selected, compressed tokens and scores the full prefix only with a light indexer.
follows the same trend. Compared with , it uses sliding-window attention and roughly 5× fewer full-attention layers, so the compute advantage widens as context length increases.
reaches the lowest FLOPs/token at long contexts in this comparison. Additionally, together with it is the flattest curve over context length, with only 13% and 33% increase respectively from 4K to 1M.
Multiplying FLOPs by the token horizon gives us the total training FLOPs, 8 and dividing by your achieved training FLOP/s gives an estimate of your training duration. Consider that the real training cost is far more complex than just FLOPs. Some of the other major factors, not addressed here, are the overlap of computation and communication, the quantisation of your weights, kernel efficiency and kernel launch overhead, activation recomputation and expert imbalance.
FLOPs are a useful proxy for prefill cost, but they do not directly translate into faster decoding, where memory movement becomes more important. This motivates the next section on sequence-mixer state size.
Sequence-Mixer State Size
At every decode step, the new query must interact with the representations of earlier tokens. Those key and value representations were already computed during the previous steps and never change, so we can cache them instead of re-projecting, trading FLOPs for HBM capacity and more I/O. The attention computation itself is unaffected, since at each step we simply append the K and V vectors of the current token.
What limits decode speed depends on the regime: A decode step multiplies every weight matrix with only one token per sequence, so at short context and small batch it is dominated by streaming the model weights from HBM to the compute units (memory bound); for MoE, the weights read are those of every expert selected by at least one sequence, which approach all experts as the batch grows. At large batch and short context, weight reuse raises arithmetic intensity and compute can become the limit (compute bound). Once contexts grow long, the bottleneck is the I/O cost of reading the state for every generated token, which, unlike the weights, is not shared between sequences (memory bound). These are only general regimes: the limiting factor depends on the arithmetic intensity of the operation and on the hardware specifications, in particular the bandwidth (byte/s) and the throughput (FLOP/s), both of which also depend on the numerical precision. For a deeper explanation of these topics, read our post Inference performance from first principles.
Cached Representation
Standard attention stores separate key and value vectors (). Multi-head latent attention (MLA, introduced in DeepSeek-V2) instead caches a single vector per token (): a low-rank content latent plus the positional component required by its decoupled RoPE path.
Batch
In a naive implementation every sequence of the batch has its own independent state. Prefix caching instead reuses a shared prefix (e.g. the system prompt) across sequences, so one copy of that state serves many requests.
Bytes per Element
Values in the cache are commonly stored in 1 byte (FP8) or 2 bytes (BF16). Methods such as KIVI and KVQuant go down to 2–4 bits, but are not yet common in production serving. Recurrent states are often kept in FP32 instead.
KV Width
The number of KV heads times the head dimension sets . GQA shares one KV head across a group of query heads, while MQA uses a single KV head for all queries.
Attention Layers
By default every attention layer keeps its own state. Hybrids replace most full-attention layers with sliding-window or recurrent layers, and research directions such as cross-layer attention (CLA) let groups of layers share one layer’s K and V. DeepSeek V4.1-Flash, for example, computes its compressed cache in four layers and lets the layers in between reuse it.
Retained Tokens
Full attention retains the entire prefix. A common and robust variation is the sliding window, which keeps only the newest tokens. Selective attention, such as NSA, reads only a sparse subset of the prefix at each step. This cuts the bytes read per step, but not the bytes stored, since any block may be selected later.
Recurrent Layers
Mamba-2 and Gated DeltaNet layers fold the whole prefix into a fixed state per head, elements for GDN and for Mamba-2, plus the last few inputs of a short convolution. A Qwen3.8-Flash-Next GDN layer keeps 48 heads of , 0.79M elements, as many as one of its attention layers caches (K, V and indexer key) for about 680 tokens, but the same at any context length.
Sequence-Mixer State Over Sequence Length
Sequence-Mixer State Across Architectures
Figure 5 shows the total sequence-mixer state and the amount of state read at each decode step, which capture two different constraints.9 Total state size determines how many sequences and how much context can fit in a decode batch, while the state read per token is more directly tied to decode speed in the long-context, memory-bound regime.
This distinction is particularly important for sparse attention. Many sparse mechanisms reduce the number of tokens accessed at each step, lowering memory traffic and FLOPs, without reducing the total state that must be stored. At 1M tokens, holds 72B elements per sequence but reads 11B per decode step. Sliding-window and recurrent layers reduce both, while techniques such as KV compression or cross-layer sharing can also shrink the stored state itself.
This produces a substantially different ranking from FLOPs. , for example, has lower long-context FLOPs than , while the latter achieves a smaller sequence-mixer state through KV compression and reuse. and DeepSeek V4.1-Flash similarly occupy comparable parameter scales but very different points in terms of state size: at 1M tokens, 72B elements per sequence for MiniMax M3 against 1.7B for DeepSeek V4.1-Flash.
also improves substantially over by replacing most full-attention layers with sliding-window attention, reducing the state that grows with context length: at 1M tokens it holds 10.8B elements per sequence, 5× less than Kolibri Origin’s 53.7B.
Conclusion
In this blogpost, we present our Model Explorer, used for designing the Kolibri architecture. We analyse how parameters and FLOPs are allocated across the main model components for popular open models and how the sequence-mixer state grows with sequence length under different architectures. Additionally, we derive formulas for FLOPs and parameters for our Kolibri models. We are excited to share this with the community, so others can explore architecture trade-offs and reason about architecture from first principles.
Acknowledgements
We are grateful to Ahmed Hammam, Fabien Benureau, Jonas Knupp, Simon Thel, Thomas Burns and Yasser Jadidi for their thoughtful review. We thank Alexander Wortmeier and Noé Beckerle Vallejo for their support bringing it to our website and Helena Treeck and Svenja Fahlisch for keeping everything on track.
References
- Ainslie, J. et al. (2023). GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245.
- Aleph Alpha Research. Inference performance from first principles. Aleph Alpha blog.
- Anthony, Q., Biderman, S., Schoelkopf, H. (2023). Transformer Math 101. EleutherAI blog.
- Beltagy, I., Peters, M. E., Cohan, A. (2020). Longformer: The Long-Document Transformer. arXiv:2004.05150.
- Brandon, W. et al. (2024). Reducing Transformer Key-Value Cache Size with Cross-Layer Attention. arXiv:2405.12981.
- Dai, D. et al. (2024). DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models. arXiv:2401.06066.
- Dao, T., Gu, A. (2024). Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. arXiv:2405.21060.
- DeepSeek-AI (2024). DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434.
- Ding, M. et al. (2021). CogView: Mastering Text-to-Image Generation via Transformers. arXiv:2105.13290.
- Gu, A., Dao, T. (2023). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752.
- Hooper, C. et al. (2024). KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization. arXiv:2401.18079.
- Liu, Z. et al. (2024). KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache. arXiv:2402.02750.
- Raschka, S. (2026). LLM Architecture Gallery. sebastianraschka.com.
- Shazeer, N. et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. arXiv:1701.06538.
- Shazeer, N. (2019). Fast Transformer Decoding: One Write-Head is All You Need. arXiv:1911.02150.
- Shazeer, N. (2020). GLU Variants Improve Transformer. arXiv:2002.05202.
- Su, J. et al. (2021). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864.
- Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762.
- Yang, S., Kautz, J., Hatamizadeh, A. (2024). Gated Delta Networks: Improving Mamba2 with Delta Rule. arXiv:2412.06464.
- Yuan, J. et al. (2025). Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. arXiv:2502.11089.
- Zhang, B., Sennrich, R. (2019). Root Mean Square Layer Normalization. arXiv:1910.07467.
- Zheng, L. et al. (2023). SGLang: Efficient Execution of Structured Language Model Programs. arXiv:2312.07104.