Detecting VLM Hallucinations via Two-Token Transitions
Using transition features between consecutive hidden states allows smaller models to detect hallucinations in larger vision-language models without full-model inference.
A 4-billion parameter Vision Language Model acts as a per-token hallucination detector by analyzing transitions between consecutive token hidden states. This architecture replaces the need for costly full-model inference or black-box external metrics by integrating a smaller, specialized classifier directly into the evaluation loop. For builders, this means detecting model errors in real-time without the overhead of re-running massive generative models for every token.
The Mechanism of Two-Token Features
The detector identifies hallucinations by mapping the trajectory of the model's latent state. In transformer architectures, hidden states represent a position in high-dimensional embedding space. When a model generates a token, the transition from the previous state to the current one creates a vector that encodes the model's certainty. During a hallucination, this transition vector often deviates from the manifold of 'factual' transitions observed during training. The detector treats the difference between the hidden state at $t-1$ and $t$ as a feature, capturing the geometric shift in the model's confidence. If the transition vector magnitude or orientation falls outside a learned threshold of validity, the token is flagged as a likely hallucination.
Training on Synthetic Error Modes
To train this 4B classifier, researchers use synthetic data produced by a 400B parameter model. The large model is prompted to generate captions with varied fidelity, intentionally injecting factual inconsistencies based on source images. Crucially, the 4B model is trained on these specific pairings: pairs where the 400B model produces a correct label, and pairs where it generates a known hallucination. The smaller model learns to classify these based on the internal latent transition signals, effectively becoming a specialist in identifying the specific failure signatures that the 400B model produces when it loses track of the image content. The 4B classifier essentially learns to recognize the 'untruthful' internal states that the 400B model manifests even when the larger model fails to catch its own error.
Implementation in Ensembles
At inference time, the 4B classifier and the 400B model operate as a single ensemble. As the 400B model produces a sequence, the 4B classifier scans the transition between each generated token. This system functions as a gate; the 4B classifier provides a binary flag on the validity of the token, which can then be used to prune the generation, flag it for human review, or trigger a re-generation for the specific segment. The compute budget is optimized because the expensive 400B inference is coupled with a lightweight, per-token pass that adds negligible latency.
Operational Trade-offs
For an application, this implies that inference latency will fluctuate depending on the model's output quality. A sequence with high-confidence tokens incurs only the marginal cost of the 4B pass, while flagged tokens suggest a need for further downstream logic. A concrete example of where this breaks is in scenes with high visual density, such as dense document processing. When the model encounters text-heavy images, the OCR-driven features can dominate the signal, potentially masking subtler geometric deviations in the visual embeddings. While this provides a mechanism to monitor for hallucinations, it is currently unknown how this detector performs against adversarial prompts designed to trigger high-confidence, high-probability hallucinations that mimic the geometric footprint of factual text.