TY - GEN
T1 - Probe-then-Commit Multi-Objective Bandits
T2 - 24th International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and WirelessNetworks, WiOpt 2026
AU - Shi, Ming
N1 - Publisher Copyright:
© 2026 IFIP.
PY - 2026
Y1 - 2026
N2 - We study an online resource-selection problem motivated by multi-radio access selection and mobile edge computing offloading. In each round, an agent chooses among K candidate links/servers (arms) whose performance is a stochastic d-dimensional vector (e.g., throughput, latency, energy, reliability). The key interaction is probe-then-commit (PtC): the agent may probe up to q > 1 candidates via control-plane measurements to observe their vector outcomes, but must execute exactly one candidate in the data plane. This limited multi-arm feedback regime strictly interpolates between classical bandits (q = 1) and full-information experts (q = K), yet existing multi-objective learning theory largely focuses on these extremes. We develop PtC-P-UCB, an optimistic probe-then-commit algorithm whose technical core is frontier-aware probing under uncertainty in a Pareto mode, e.g., it selects the q probes by approximately maximizing a hypervolume-inspired frontier-coverage potential and commits by marginal hypervolume gain to directly expand the attained Pareto region. We prove a dominated-hypervolume frontier error of O(KPd/√qT), where KP is the Pareto-frontier size and T is the horizon, and scalarized regret O(Lφd√(K/q)T), where φ is the scalarizer. These quantify a transparent 1/√ q acceleration from limited probing. We further extend to multi-modal probing: each probe returns M modalities (e.g., CSI, queue, compute telemetry), and uncertainty fusion yields variance-adaptive versions of the above bounds via an effective noise scale.
AB - We study an online resource-selection problem motivated by multi-radio access selection and mobile edge computing offloading. In each round, an agent chooses among K candidate links/servers (arms) whose performance is a stochastic d-dimensional vector (e.g., throughput, latency, energy, reliability). The key interaction is probe-then-commit (PtC): the agent may probe up to q > 1 candidates via control-plane measurements to observe their vector outcomes, but must execute exactly one candidate in the data plane. This limited multi-arm feedback regime strictly interpolates between classical bandits (q = 1) and full-information experts (q = K), yet existing multi-objective learning theory largely focuses on these extremes. We develop PtC-P-UCB, an optimistic probe-then-commit algorithm whose technical core is frontier-aware probing under uncertainty in a Pareto mode, e.g., it selects the q probes by approximately maximizing a hypervolume-inspired frontier-coverage potential and commits by marginal hypervolume gain to directly expand the attained Pareto region. We prove a dominated-hypervolume frontier error of O(KPd/√qT), where KP is the Pareto-frontier size and T is the horizon, and scalarized regret O(Lφd√(K/q)T), where φ is the scalarizer. These quantify a transparent 1/√ q acceleration from limited probing. We further extend to multi-modal probing: each probe returns M modalities (e.g., CSI, queue, compute telemetry), and uncertainty fusion yields variance-adaptive versions of the above bounds via an effective noise scale.
KW - hypervolume
KW - limited multi-arm feedback
KW - multi-modal feedback
KW - multi-objective bandits
KW - online resource selection
KW - Pareto frontier
KW - probe-then-commit (PtC)
KW - scalarization
UR - https://www.scopus.com/pages/publications/105043607829
U2 - 10.23919/WiOpt71098.2026.11568166
DO - 10.23919/WiOpt71098.2026.11568166
M3 - Conference contribution
AN - SCOPUS:105043607829
T3 - Proceedings of the International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks, WiOpt
BT - 2026 24th International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and WirelessNetworks, WiOpt 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 3 June 2026 through 6 June 2026
ER -