Pareto Frontier of LLM Test-time Compute: Adaptive Rollouts and Voting Under Token Budgets

Large language models improve test-time accuracy by generating multiple rollouts and aggregating these answers. As generating additional rollouts consumes more tokens, test-time compute faces a fundamental cost accuracy trade-off, naturally raising the question: given a fixed token budget, how to allocate the budget to maximize accuracy? Existing approaches typically adapt the number of rollouts for … Read more