Skip to main navigation Skip to search Skip to main content

Optimal Control of the Future via Prospective Learning with Control

  • Yuxin Bai
  • , Aranyak Acharyya
  • , Ashwin De Silva
  • , Zeyu Shen
  • , James Hassett
  • , Joshua T. Vogelstein
  • Johns Hopkins University

Research output: Contribution to journalConference articlepeer-review

Abstract

Optimal control of the future is the next frontier for AI. Current approaches to this problem are typically rooted in reinforcement learning (RL). RL is mathematically distinct from supervised learning, which has been the main workhorse for the recent achievements in AI. Moreover, RL typically operates in a stationary environment with episodic resets, limiting its utility. Here, we extend supervised learning to address learning to control in non-stationary, reset-free environments. Using this framework, called "Prospective Learning with Control" (PLuC), we prove that under certain fairly general assumptions, empirical risk minimization (ERM) asymptotically achieves the Bayes optimal policy. We then consider a specific instance of prospective learning with control: foraging, a canonical task relevant to both natural and artificial agents. We illustrate that modern RL algorithms, which assume stationarity, struggle in these non-stationary reset-free environments. Even with time-aware modifications, they converge orders of magnitude slower than our prospective foraging agents on a simple 1-D foraging benchmark. Code is available at: https://github.com/neurodata/procontrol.

Original languageEnglish
JournalProceedings of Machine Learning Research
Volume331
StatePublished - 2026
Event8th Annual Learning for Dynamics and Control Conference, 2026 - Los Angeles, United States
Duration: Jun 17 2026Jun 19 2026

Keywords

  • adaptive control
  • reset-free single-episode reinforcement learning
  • statistical learning for dynamical and control systems

Fingerprint

Dive into the research topics of 'Optimal Control of the Future via Prospective Learning with Control'. Together they form a unique fingerprint.

Cite this