Bayesian Backward Reasoning for LLM Consensus
By checking if an LLM's conclusion implies its evidence, you can filter out faulty logic without needing labeled data.
When multiple LLM agents provide conflicting answers, traditional methods like majority voting or LLM-based judging rely on forward reasoning. These techniques map symptoms to conclusions in a single direction. Diverse groups of agents often share the same training-induced biases, which leads to identical correlated errors. The paper 'When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making' proposes a dual-factorized system that incorporates Bayesian backward reasoning. By generating a reverse posterior alongside the standard forward prediction, you measure cross-path consistency to determine which agent is correct without labeled data.
The Mechanism of Backward Verification
The framework treats the reasoning process as a probabilistic flow between evidence and a conclusion. In a forward path, the model takes a set of symptoms and predicts a diagnosis. To compute the reverse posterior, the model performs the inverse operation. It takes the proposed diagnosis as a given fact and evaluates the probability of the original symptoms.
In practice, the model does not run a formal Bayesian inference engine. Instead, it uses the LLM's ability to output probability scores for individual tokens. When the model is tasked with generating the likelihood, it assesses how likely the input symptoms are given a specific label. If the model predicts a diagnosis of 'influenza', it then calculates the likelihood of 'fever' and 'cough' being present. If the forward reasoning led to a diagnosis that makes the evidence seem statistically impossible, the divergence between the forward and backward probabilities spikes, signaling a failure in logic.
Aggregation Strategies
Existing aggregation strategies are vulnerable to consensus-seeking errors where independent agents fail in identical ways. The research highlights three methods for utilizing these dual signals: MinJS, FwdJS, and LogLin.
| Method | Approach | Relies on Labeled Data | Error Mitigation |
|---|---|---|---|
| Majority Voting | Forward Aggregation | No | Low |
| LLM Judge | Forward Aggregation | No | Medium |
| MinJS | Hard selection | No | High |
| FwdJS | Hard selection | No | High |
| LogLin | Log-linear fusion | No | High |
MinJS operates by selecting the agent with the lowest Jensen-Shannon divergence between its forward and backward paths, discarding any agent whose logic does not mirror itself. FwdJS selects the agent based on the highest forward confidence, essentially using the backward path as a secondary filter.
LogLin (log-linear fusion) takes a different approach by mathematically combining the forward and backward probabilities. It calculates the final label by taking the weighted sum of the log-probabilities from both the forward and backward paths. Because the backward path acts as an independent anchor, it prevents the final result from being skewed by a single, overconfident forward prediction. This mathematical integration effectively penalizes agents that produce high-confidence, internally inconsistent chains, resulting in lower error rates on the DDXPlus benchmark compared to standard voting.
Evaluation and Scaling
The results from the DDXPlus benchmark show that reverse posterior signaling consistently outperforms standard voting methods across five different model architectures. The accuracy gains are specifically driven by the independence of the backward chain, which acts as a secondary verification layer that is not susceptible to the same biases as a standard forward prediction.
Despite these gains, the method faces a specific hurdle in implementation. The authors rely on an explicit likelihood to reverse the reasoning chain, but defining this likelihood in a way that remains consistent across different domains remains difficult. If the task is a medical diagnosis, mapping symptoms to a disease is straightforward. If the task is less structured, determining what constitutes a 'valid symptom' for a specific outcome becomes subjective and prone to hallucination. While the framework demonstrates a clear path toward verifying agent outputs, it is not yet known how to automate the generation of this likelihood function for tasks that do not easily map to clear, probabilistic classification categories.