Viyan

Viyan AI

Mapping Information Routing in Transformers

Recent research reveals how Qwen, Llama, and Gemma process internal routing signals versus factual content during inference.

Large language models map input text into high-dimensional hidden states where internal processing is not a uniform blend of information. Research into Qwen, Llama, and Gemma identifies that these models follow specific causal trajectories when retrieving knowledge. Models maintain vector directions within the hidden state space that function as steering signals. By projecting hidden states onto these directions, one can measure how a model balances query-specific intent against general factual knowledge.

The Mechanism of Internal Routing

When a model processes a prompt, its weight matrices perform linear transformations on activation vectors. A target direction represents a specific semantic intent, such as the relationship between a country and its continent. This direction is identified by analyzing how internal activations shift when a prompt contains a specific query structure versus a neutral one. The model essentially stores these semantic associations in the weights; when an input triggers an activation pattern that aligns with this learned target vector, it creates a coherent path for information flow.

This alignment acts as a functional filter. In the research, this gain is a post-hoc intervention where the activation vector is adjusted by adding a scaled steering vector to the residual stream. This addition amplifies the signal along the target direction, effectively pushing the state toward the desired subspace. Imagine a hallway with multiple branching paths; the gain factor acts like a guide that makes one specific door much easier to enter while the others become harder to notice. If these vectors are orthogonal, they remain isolated; if they are not, they become entangled, making it difficult for the model to separate the specific fact from the broader context.

Global fitted directions serve as a stable, broad-reaching routing guide, while pair-conditioned directions are specific to individual entity pairs like a country and its continent. In Qwen, the pair-conditioned direction persists late into the inference process, allowing the final layers to maintain access to the intent. Other architectures dissolve this routing information much earlier in the forward pass, causing the signal to be subsumed by general processing.

Structural Variations in Processing

Model Routing Window Duration Handoff Pattern
Qwen Extended Distinct
Gemma Mid-layer Partial overlap
Llama None Saturated

Architecture dictates the longevity of these routing windows. Qwen maintains a sustained routing-effect window where the intent signal remains distinguishable through the final layers. Gemma shows a mid-layer routing window with partial overlap between the routing signal and the content generation. Llama lacks a sustained routing-effect window, meaning the routing signals are not held in isolated trajectories during the forward pass.

Consider an evaluation on a query like "Which continent is Norway in?". In Qwen, the internal routing direction for that specific pair persists, keeping the hidden state aligned with the target vector and ensuring the final token prediction remains on track. In Llama, the lack of a sustained window means the "Norway" and "continent" signals are not held in isolated trajectories. The distinct intent signal is quickly lost, meaning the prediction relies on the model’s aggregate knowledge rather than a guided, prioritized path.

What remains unclear is the exact threshold at which these steering signals shift from being functional pathways to noise. While we can now observe these vectors, we do not yet have a precise method to force a model to prioritize one routing trajectory over another without significant performance degradation in other tasks. Future work will likely focus on whether we can prune the layers that dissolve these signals too early to improve model reliability.

Sources