The results reveal a trade-off in allocating customer interactions across BLMP’s two stages: more interactions in the first stage generally improve final pricing quality, whereas allocating relatively more interactions to the second stage increases revenue collected during learning. These findings show how exploiting the structure embedded in customer choices can make bundle-price learning more tractable and shed light on the allocation of a finite budget of customer interactions.
Bundle Pricing via Learning the Market from Customer Preferences
\(\)We study the problem of learning revenue-maximizing bundle prices when customer valuations and market composition are unknown and the firm observes only customer choices. This problem is challenging because the number of bundle-price decisions grows exponentially with the number of items offered. We develop BLMP (Bundle Pricing via Learning the Market from Customer Preferences), a learning-and-optimization approach that transforms observed bundle choices into item-level revealed-preference information, sequentially estimates valuation characteristics and segment fractions, and then solves a mixed-integer bundle-pricing model using the estimated parameters. The framework accommodates both deterministic and bounded probabilistic valuations and reduces the dimensionality of the valuation learning problem from $O(2^n)$ to $O(n)$. We benchmark BLMP against CMA-ES and GP-UCB across different problem dimensions and valuation distributions. BLMP outperforms CMA-ES and remains competitive with GP-UCB while avoiding the latter’s computational difficulties as the number of observations increases. BLMP produces higher-quality final pricing solutions and moves closer to the full-information benchmark, whereas CMA-ES shows little improvement.
The results reveal a trade-off in allocating customer interactions across BLMP’s two stages: more interactions in the first stage generally improve final pricing quality, whereas allocating relatively more interactions to the second stage increases revenue collected during learning. These findings show how exploiting the structure embedded in customer choices can make bundle-price learning more tractable and shed light on the allocation of a finite budget of customer interactions.
The results reveal a trade-off in allocating customer interactions across BLMP’s two stages: more interactions in the first stage generally improve final pricing quality, whereas allocating relatively more interactions to the second stage increases revenue collected during learning. These findings show how exploiting the structure embedded in customer choices can make bundle-price learning more tractable and shed light on the allocation of a finite budget of customer interactions.