- Dissertation
- 10.32657/10356/204489
Generalizable neural solvers for vehicle routing problems
- Jan 01, 2025
- Jianan Zhou
Vehicle routing problems (VRPs) are a fundamental class of combinatorial optimization problems (COPs) in computer science and operations research, with diverse real-world applications in logistics, transportation, and manufacturing. The intrinsic NP-hard nature makes VRPs exponentially expensive to be solved by exact solvers. As an alternative, heuristic solvers deliver suboptimal solutions within reasonable time but require extensive hand-crafted rules and domain expertise tailored to each specific problem. Recently, neural combinatorial optimization (NCO), which leverages machine learning to learn heuristics in a data-driven manner, has gained significant attention. These neural solvers demonstrate strong performance while reducing computational overhead and reliance on domain expertise compared to traditional solvers. However, they face significant generalization challenges. For example, their performance may degrade substantially when applied to instances with different data distributions, problem scales, or constraints than those encountered during training. This generalization issue severely limits their practical applicability. This thesis aims to systematically address these challenges and advance the development of generalizable neural solvers for VRPs. The thesis begins by enhancing the cross-distribution generalization of neural VRP solvers through an adversarial training approach. We introduce an ensemble-based Collaborative Neural Framework (CNF) that adversarially trains multiple models in a collaborative manner, promoting robustness against adversarial attacks while also boosting performance on clean instances. In this framework, an attacker generates hard instances by perturbing the data distribution, challenging the models to adapt and thereby strengthening cross-distribution generalization. Additionally, a neural router is designed to efficiently distribute training instances among the models, thereby improving load balancing and collaborative efficacy. Extensive experiments on two classic routing problems -- the travelling salesman problem (TSP) and the capacitated vehicle routing problem (CVRP) -- demonstrate the effectiveness and versatility of CNF in enhancing the cross-distribution generalization. Next, the thesis further expands neural solvers to simultaneously address both cross-size and cross-distribution generalization challenges. We propose Omni-VRP, a generic meta-learning framework that enables effective training of an initial model with the capacity for rapid adaptation to new tasks during inference, where each task corresponds to a set of instances with the same problem scale and data distribution. Additionally, we introduce a simple yet efficient approximation method to reduce the training overhead associated with second-order derivative calculations. Extensive experiments on both synthetic and benchmark instances of TSP and CVRP validate the effectiveness of our approach. Finally, the thesis pioneers the study of cross-problem generalization, pushing the boundaries of neural VRP solvers in learning generalizable representations across diverse constraints. We propose MVMoE, a unified neural solver capable of solving 16 VRP variants simultaneously in a zero-shot manner through attribute composition. To efficiently enhance model capacity, we employ a mixture-of-experts architecture with a hierarchical gating mechanism, striking a balance between empirical performance and computational complexity. Experimentally, our method significantly promotes zero-shot generalization on 10 unseen VRP variants, and showcases decent results on the few-shot setting and real-world benchmark instances. Overall, this thesis represents a pioneering exploration and substantial advancement in developing generalizable neural solvers for VRPs. It investigates innovative designs across various components, including model architectures, training algorithms, and problem settings, to enhance the generalization of neural solvers across diverse contexts such as varying data distributions, problem scales, and constraints. These contributions encourage neural solvers to learn more robust and generalizable representations, paving the way for the first generation of foundation models capable of addressing a wide range of COPs. Ultimately, the insights gained here not only propel the field of NCO forward but also enrich the broader landscape of learning-based optimization methods.
Read more