Exact Policy Optimization Outperforms Sampled Learning
Full-Group Policy Optimization (FGPO) replaces stochastic sampling with exhaustive enumeration to stabilize agent training in domains with limited tool sets.
Full-Group Policy Optimization (FGPO) replaces stochastic sampling with exhaustive enumeration when selecting tools for frozen-reasoner agents. Researchers identified a structural mismatch in scientific domains like genomics, where the number of possible tool subsets is small enough to evaluate completely, yet common training methods like Group Relative Policy Optimization (GRPO) continue to sample from that space. This approach solves the reward collapse that occurs as models converge on specific tool subsets during training.
In standard RL training for agent tool use, models rely on sampling rollouts to estimate which tool combinations yield the best results. GRPO calculates the advantage for each rollout by comparing its reward against the mean reward of the group. As the policy begins to prefer specific subsets, the sampling process repeatedly selects the same configurations. When every rollout in a group returns the identical tool subset, the variance of the group's rewards drops to zero. Because the advantage is defined by the difference between the sample reward and the group mean, the signal effectively disappears, leaving the policy update with no gradient to follow. In genomic reasoning tasks, this behavior caused the failure rate, where questions receive no reward signal, to balloon from 0.2% under a uniform policy to 20.8% after training with GRPO.
How the Update Function Shifts
FGPO abandons sampling in favor of calculating the exact expectation across all possible tool subsets. Instead of calculating an advantage based on the variance within a batch of sampled rollouts, the model updates the policy by taking the weighted sum of rewards across the entire power set of tools. Specifically, the objective function uses the probability of each subset, assigned by the current policy, multiplied by its precomputed reward. This makes the update deterministic. Because the model uses the global context of all possible rewards, it avoids the feedback loop where a local sampling bias masks superior tool combinations. This method requires 2.4 times fewer frozen-reasoner evaluations than a standard on-demand GRPO schedule, as all outcomes are precomputed into a table before training begins.
Comparing Training Strategies
| Feature | GRPO (Sampled) | FGPO (Full-Group) |
|---|---|---|
| Action Space | Sampled subset | Exhaustive set |
| Reward Estimation | Statistical | Exact expectation |
| Reasoner Calls | Iterative/Per-step | Precomputed table |
| Reward Signal Decay | High | Zero |
Implementation Trade-offs
To decide if your project warrants FGPO, you must evaluate the size of your tool library. If your agent chooses from N tools, the total number of combinations is 2 to the power of N. For an agent selecting subsets of up to 4 tools from a library of 8, there are 163 possible combinations. If this number is manageable for your compute budget, FGPO provides a stable, deterministic signal that consistently outperforms sampling. Across three genomic benchmarks and five frozen reasoners, FGPO outperformed GRPO by an average of 6.75 points, peaking at 14.20 across all 15 experimental settings. On the GenomeQA benchmark, this change reduced the average number of tools invoked per question from 2.36 to 1.40, yielding a more efficient and accurate reasoning path.
Consider an agent designed to query a database of protein structures using three different analysis tools. With a library of four tools, there are only 15 possible combinations. In a sampled approach, the model might gravitate toward a single tool that provides a safe but mediocre answer, causing the advantage signal to collapse as that tool appears in every sampled group. With FGPO, the model evaluates all 15 combinations. It sees that a different combination of two tools yields a significantly higher reward. Because it evaluates the entire set, it incorporates this superior path into its policy update regardless of its current sampling bias.
Scaling this technique is limited by the exponential growth of the power set. Once the tool library exceeds roughly 10 to 15 options, the combinatorial explosion makes precomputation impractical. At that threshold, the memory requirements for the reward table overwhelm the performance gains, and a return to stochastic sampling becomes necessary. It remains an open question whether hybrid approaches, which perform partial enumeration for common tool clusters and sampling for rarer configurations, can extend these benefits to larger agentic libraries.