Workshop at the Neural Information Processing Systems 2026 conference (Paris)
Reinforcement learning has achieved major breakthroughs in simulated environments, yet transferring these advances to real experimental systems remains challenging. This workshop explores how RL can enable adaptive experimentation and real-world scientific discovery while addressing the simulation-to-reality gap.
Many scientific experiments can naturally be formulated as sequential decision processes in which each experimental action influences future observations. Reinforcement learning provides a principled framework to optimize such adaptive decision loops.
Differences between simulations and physical systems create major challenges for deploying reinforcement learning in laboratory and field experiments.
Many scientific experiments are costly, slow, or resource-constrained. Reinforcement learning methods can help allocate experimental effort more efficiently, accelerating discovery while reducing material, time, and financial costs.
Adaptive experiment design, autonomous laboratories, and sequential optimization in scientific experiments.
Domain randomization, system identification, hybrid training pipelines, and robust policy transfer.
Sample-efficient reinforcement learning, safe exploration, uncertainty-aware decision making.
Agricultural experimentation, ecological monitoring, and environmental sensing systems.
Human–AI collaboration in experimentation and participatory research frameworks.
Simulation environments, digital twins, and benchmarking platforms for RL in experimental science.
University of Salzburg
Particle accelerators generate vast amounts of historical data from logs, yet learning-based control often still relies on risky online optimisation. To better utilise this data and avoid online exploration, we present an offline reinforcement learning (RL) workflow. First, we use XSuite to generate high-fidelity trajectories for steering tasks across representative scenarios, including optics variations, alignment errors and jitter, yielding a synthetic dataset of expert and non-expert behaviour. Second, we learn an uncertainty-aware, Koopman-stabilised world model from this data, in which nonlinear beam dynamics are lifted into a latent space with approximately linear, spectrally constrained evolution and a regularised residual term. This structure provides numerically stable long-horizon rollouts and estimates of epistemic uncertainty in latent space. The resulting surrogate environment enables model-based offline RL, where policies are optimised entirely on pre-generated data while epistemic uncertainty is used to detect distribution shift and enforce soft safety constraints. We benchmark these offline RL policies against a PPO agent trained directly in simulation. Results show that policies trained purely offline on the Koopman world model can match online PPO performance without requiring any interaction with the real machine. This demonstrates a safe, reproducible pathway for turning historical accelerator data into effective learning-based control policies.
University of Cambridge
In machine learning literature, multi-armed bandits and reinforcement learning are heavily celebrated for optimizing the "learning-earning" trade-off—balancing the exploration of unknown actions to gather data with the exploitation of known actions to maximize immediate rewards. When applying these sequential decision-making frameworks to confirmatory clinical trial design, this trade-off manifests as a profound ethical and statistical tension: maximizing care for current trial participants versus optimizing scientific data gathered to benefit future patients.
However, translating multi-armed bandits and adaptive optimization algorithms into real-world clinical trial architectures introduces significant methodological bottlenecks. Most notably, adaptively collected data destroys the independent and identically distributed (i.i.d.) assumptions of classical statistics, making error control notoriously difficult. While clinical trial efficiency is paramount—especially given small sample size constraints—strict regulatory standardisation demands that Type I error rate inflation must be completely mitigated. To date, the adoption of these methods in confirmatory trials is practically zero, despite gaining some traction in early-phase drug development.
In this talk, I will provide a comprehensive overview of the mathematical, operational, and regulatory challenges that arise when bridging advanced reinforcement learning concepts with clinical trial design. Finally, I will present a suite of specific solutions developed in our group. These approaches successfully close the loop by blending multi-objective adaptive optimization during the design phase with robust, adjusted statistical analysis at the study’s conclusion, ensuring both trial efficiency and regulatory compliance.
Université de Montréal & Mila, Canada
TBD
AstraZeneca, Gothenburg
TBD
09:00 – 09:15
Opening remarks and introduction
09:15 – 10:00
Invited talk: RL for Autonomous Scientific Discovery
10:00 – 10:45
Invited talk: Sim-to-Real RL in Robotics and Experimental Platforms
10:45 – 11:15
Coffee break
11:15 – 12:15
Contributed spotlight talks (6 × 10 min)
12:15 – 12:45
Panel discussion: “Can RL Become a Standard Tool for Experimental Science?”
14:00 – 14:45
Invited talk: Adaptive Experiment Design with RL
14:45 – 15:30
Invited talk: Learning in Physical Laboratory Environments
15:30 – 16:00
Coffee break
16:00 – 16:45
Poster session and discussion to facilitate networking
16:45 – 17:00
Closing discussion: “Open challenges in RL for experimental sciences”
We invite submissions describing reinforcement learning methods applied to real-world experimental systems, including research papers, system descriptions, benchmarks, and position papers. We will put efforts to include diverse participants. We encourage participation from both machine learning researchers and experimental scientists, practitioners, and interdisciplinary scientists.
Formatting Instructions: We solicit workshop papers following the NeurIPS 2026 main conference paper template. The main text of a submitted paper is limited to nine content pages, including all figures and tables. Additional pages containing references, optional technical appendices and mandatory paper checklist do not count as content pages. The maximum size of submissions is 50 MB. While your submission can contain a supplement or appendix, please note that reviewers are not obliged to review supplementary material.
Reviews: The review process will be double-blind. All submissions must be anonymized and the leakage of any identification information is prohibited. All submissions will undergo peer review by the Program Committee. Each submission will receive at least three reviews, followed by a discussion phase among reviewers and organizers.
Presentation: Accepted contributions will be presented as posters, with a subset selected for spotlight presentations.
What can be submitted: Following the convention of most workshops, we welcome submissions of papers already published in journals or presented at non–machine-learning conferences or workshops. If a paper has appeared in a machine-learning venue, we will consider it only if it includes substantial extensions or new results.
Submission deadline: August 29, 2026
Notification: September 29, 2026
Camera-ready: November 29, 2026
Workshop: December 12 or 13, 2026
INRIA & University of Lille, France
He is a senior researcher at Inria Lille and head of the MARLEL group, working on the foundations of reinforcement learning and sequential statistics with a focus on regret, robustness, and scientific experimentation. He actively develops interdisciplinary applications of RL in fields such as agronomy and medicine.
Technical University of Leoben, Austria
He is an Associate Professor at the Technical University of Leoben specializing in theoretical reinforcement learning, particularly regret analysis in Markov decision processes. His work focuses on designing sample-efficient algorithms suited to the stringent data regimes of experimental sciences.
Université Laval, Canada
She is an Associate Professor at Université Laval and Mila research affiliate, working on interactive machine learning, including reinforcement learning, bandits, and active learning. Her research emphasizes learning systems that actively shape their data acquisition, with applications in healthcare and real-world decision systems.
University of Copenhagen, Denmark
He is an Assistant Professor at the University of Copenhagen whose research focuses on uncertainty-aware, risk-sensitive reinforcement learning and decision-making. He develops methods for reliable adaptive experimentation under uncertainty, with applications to safety-critical and model-based scientific systems.
Higher Polytechnic School, Cheikh Anta DIOP University, Senegal
He is an Associate Professor at Cheikh Anta Diop University of Dakar, working on AI applications in environmental, agricultural, and health domains. His work includes reinforcement learning for sustainable agriculture, ecological modeling, and language technologies for local African languages.
BaysicLabs, France
He is a machine learning engineer and co-founder of Baysic Labs, with a PhD in computational neuroscience. He specializes in Bayesian modeling and human-in-the-loop AI, translating theoretical advances into real-world applications in biology, medicine, and industry.








This workshop brings together researchers from reinforcement learning and experimental sciences across multiple disciplines and institutions.