Large language models improve test-time accuracy by generating multiple rollouts and aggregating these answers. As generating additional rollouts consumes more tokens, test-time compute faces a fundamental cost accuracy trade-off, naturally raising the question: given a fixed token budget, how to allocate the budget to maximize accuracy? Existing approaches typically adapt the number of rollouts for each query while keeping the voting rule fixed, i.e., the rule used to aggregate the generated rollouts into a final answer. As different voting rules can favor different queries and induce different accuracy ceilings, we show that the voting rule itself should also adapt across queries. In particular, commonly used fixed voting rules such as self-consistency and best-of-n are special cases of a unified Boltzmann-weighted vote family, parameterized by a scalar τ ≥ 0. We therefore formulate an adaptive query dependent choice of jointly the number of rollouts and the voting rule under a token-budget constraint. We show that no fixed voting rule is uniformly optimal across budgets, whereas adaptive rollout and voting yields a cost–accuracy Pareto frontier that dominates every fixed voting rule’s frontier. We further characterize the large-budget asymptotics of this Pareto frontier, identifying both its attainable accuracy ceiling with an unlimited budget, and the convergence rate to the ceiling as the token budget grows. Building on these structural results, we propose PRICE (Priced Rollouts and Inference-time voting-rule Choice under a token budgEt), a practical algorithm that implements the adaptive, query-level joint decisions under a token budget. On MATH-500, PRICE consistently outperforms six test-time compute baselines across all evaluated budgets and achieves comparable accuracy using 2.8–3.5× fewer tokens.