Viyan

Viyan AI

Shrinking Model Verbosity via LOCUS Subspace Selection

LOCUS reduces model verbosity by optimizing LoRA subspaces for brevity while maintaining performance.

Instruction-tuned models are often pathologically verbose. When models undergo Reinforcement Learning from Human Feedback (RLHF) or Direct Preference Optimization (DPO), they frequently mimic the lengthy, conversational examples common in training datasets. Because inference costs scale linearly with the number of generated tokens, this tendency creates a direct financial and latency burden for production applications. LOCUS addresses this by selecting a task-aware low-rank adaptation subspace that minimizes total output-token cost while maintaining performance on utility objectives.

The Mechanism of Subspace Selection

Traditional Low-Rank Adaptation (LoRA) injects small, trainable matrices into a frozen model backbone. While standard LoRA is used to teach a model a new skill or style, LOCUS treats the configuration of these low-rank updates as an optimization problem focused on brevity. The algorithm constructs multiple candidate low-rank adaptation subspaces by varying the initialization and training objectives. It then evaluates these candidates by measuring the average output-token length against a validation dataset.

This is not a manual search through infinite orientations. Instead, LOCUS operates by selecting the specific subspace that achieves the lowest average token count while keeping the model’s performance on a utility metric, such as the Anthropic HH-RLHF dataset, within a predefined threshold of the original performance. If a subspace reduces token count but causes the model to hallucinate or fail at logical tasks, it is discarded. The method essentially finds a "cost-efficient" orientation within the LoRA parameter space that satisfies both brevity and accuracy.

Influencing Generation Patterns

By selecting a subspace that specifically optimizes for brevity, the model undergoes a shift in how it weights the continuation of a response. When a model generates text, it calculates probabilities for the next token at every step. A standard model may assign high probability to "filler" tokens—conversational pleasantries or repetitive restatements—because it associates these with "helpful" completions. A LOCUS-adjusted model, by contrast, operates within a subspace where the probability mass for tokens that signify completion or satisfy the prompt directly is higher, relative to the verbose filler tokens.

Model Method Parameter Update % Reduction in Length
Pythia-2.8B LOCUS ~0.24% 39.84%
Qwen2.5-3B LOCUS ~0.28% 14.87-17.58%

Practical Implementation for Builders

For engineers deploying models at scale, LOCUS allows you to treat brevity as a tunable variable, decoupled from the base model's core intelligence. Because you keep the backbone frozen, you avoid the risks of catastrophic forgetting that often occur during full-weight fine-tuning. Consider an automated customer support system. A base model might provide a paragraph of empathy before solving the issue. By applying a LOCUS adapter trained to prioritize brevity, you shift the generation bias.

In practice, this means the model reaches the end of the query sooner. A 39% reduction in tokens on a 2.8B model, as shown in the provided performance data, translates directly to a lower invoice from your inference provider. It also reduces time-to-first-token for your users. The mechanism relies on finding that "sweet spot" in the parameter space where the model remains functional but loses the habit of rambling. If you are building for high-volume, structured tasks, this is a way to trim costs without the compute overhead of retraining your entire model.

Sources