Beyond Shadow Weights: Quantization-Aware Training as Quantized-Endpoint Descent

Quantization-aware training (QAT) updates a full-precision shadow weight \(\mathbf{x}\) but deploys the quantized endpoint \(Q(\mathbf{x})\). Existing explanations for QAT largely view its success through the lens of shadow weights: QAT can move \(\mathbf{x}\) toward flatter basins, gain robustness from quantization-induced oscillations, or balance the shadow loss \(f(\mathbf{x})\) against the quantization error \(\|\mathbf{x}-Q(\mathbf{x})\|_2\). These perspectives do … Read more

Subgradient Regularization: A Descent-Oriented Subgradient Method for Nonsmooth Optimization

In nonsmooth optimization, a negative subgradient is not necessarily a descent direction, making the design of convergent descent methods based on zeroth-order and first-order information a challenging task. The well-studied bundle methods and gradient sampling algorithms construct descent directions by aggregating subgradients at nearby points in seemingly different ways, and are often complicated or lack … Read more

Variational Theory and Algorithms for a Class of Asymptotically Approachable Nonconvex Problems

We investigate a class of composite nonconvex functions, where the outer function is the sum of univariate extended-real-valued convex functions and the inner function is the limit of difference-of-convex functions. A notable feature of this class is that the inner function may fail to be locally Lipschitz continuous. It covers a range of important yet … Read more

A Decomposition Algorithm for Two-Stage Stochastic Programs with Nonconvex Recourse

In this paper, we have studied a decomposition method for solving a class of nonconvex two-stage stochastic programs, where both the objective and constraints of the second-stage problem are nonlinearly parameterized by the first-stage variable. Due to the failure of the Clarke regularity of the resulting nonconvex recourse function, classical decomposition approaches such as Benders … Read more