Viyan

Viyan AI

Why Mixture-of-Experts Models Overfit on Repeated Data

Mixture-of-Experts models suffer from early router stabilization when trained on repeated data, leading to specialization that prevents robust learning.

Mixture-of-Experts (MoE) models reduce compute costs by activating only a small subset of their total parameters for each token. Recent research reveals a vulnerability in this architecture when training sets are repeated. Models ranging from 80 million to 1 billion active parameters, often within larger structures containing up to 8.5 billion total parameters, overfit more aggressively than dense Transformers under these conditions. The router, which directs tokens to specific experts, stabilizes prematurely. This forces experts to specialize on idiosyncratic patterns in the data rather than learning features that transfer to unseen sequences.

The Mechanism of Expert Lock-in

When a model processes repeated data, the router quickly settles on a mapping that minimizes loss for specific sequences. In a dense model, every parameter is updated regardless of the input, which necessitates the creation of compressed, generalizable representations. In an MoE, the router partitions the input space. Once the router favors a specific expert for a recurring sequence, that expert becomes the primary recipient of the gradient updates for that data. This creates a cycle where the expert becomes increasingly proficient at reconstructing the training sequence, while other experts never encounter that data. The model treats each sequence as an isolated case to be stored in an expert's weights rather than learning general rules from the distribution. The total parameter count exacerbates this. With many experts available, the model finds it easier to allocate a unique expert to a unique segment of the data than to force a single weight set to handle multiple, potentially conflicting inputs.

Regularization as a Forcing Function

To prevent the router from collapsing into a fixed mapping, training must force the model to handle inputs with experts that are not the currently favored ones. Researchers utilize masking-based regularization to achieve this. By randomly dropping specific routing paths or experts during training, you force the input into a sub-optimal expert. This sub-optimal expert must now update its internal representations to accommodate an input it was not specifically assigned to handle. This process prevents any single expert from hoarding a training sequence. By requiring multiple experts to understand the same input distribution, the model is compelled to distribute its knowledge across a wider range of parameters.

Comparing Repetition Susceptibility

Model Type Safe Repetition Limit Susceptibility to Overfitting
Dense (80M) ~8x Low
Sparse (MoE) ~4x High

The Unknowns in Scaling

We know that expert specialization correlates with poor generalization, and that masking forces broader weight updates, but the exact limit of this technique remains unclear. It is not yet determined whether increasing the number of experts continues to provide marginal gains in capacity once a certain density of regularization is applied, or if there is a point where the overhead of forcing generalization cancels out the efficiency gains of the sparse architecture. Furthermore, most empirical data on this failure mode focuses on smaller active parameter counts. Whether these same dynamics dominate at the scale of 100-billion-parameter MoE models is a subject for further investigation. There is currently no definitive evidence that masking-based regularization can perfectly replicate the generalization performance of dense models trained on purely unique datasets.

Sources