[Robot Learning Team Seminar] Talks by Prof. Matteo Papini (University of Milan) and Dr. Gianmarco Genalti (Politecnico di Milano) on July 23
RIKEN AIP online seminar: two Italian researchers on policy gradient sample efficiency and graph-triggered bandits, July 23.
- When
- Thu, July 23, 2026 · 15:00–16:30 JST
- Where
- Open Space at the RIKEN Nihonbashi Office (on-site for RIKEN members only) · Online
- Organizer
- RIKEN Center for Advanced Intelligence Project
- Language
- EN
- Source
- Doorkeeper
Summary
This is an online seminar hosted by the Robot Learning Team, with an on-site option at RIKEN's Nihonbashi office limited to RIKEN members. Two invited talks cover recent theoretical work in reinforcement learning.
Prof. Matteo Papini (University of Milan) presents "Reusing Data in Policy Gradients to Improve Sample Efficiency," introducing an actor-only policy gradient algorithm based on a multiple importance sampling estimator that reuses trajectory data from the k most recent policies, with sample efficiency guarantees and early results on data reuse in Proximal Policy Optimization.
Dr. Gianmarco Genalti (Politecnico di Milano) presents "Bridging Rested and Restless Bandits with Graph-Triggering," proposing Graph-Triggered Bandits, a framework that generalizes rested and restless bandit settings using a graph over arms to model how pulling one arm triggers changes in another, with algorithms and guarantees for rising and rotting bandit cases.
About the community
This seminar series brings international researchers to present cutting-edge work in reinforcement learning and sequential decision-making to an audience of AI researchers and engineers in Japan. Talks combine formal theoretical results with practical algorithmic guarantees, and sessions are held online via Zoom so remote participants across Japan can join alongside a small in-person audience.
#reinforcement-learning#policy-gradient#multi-armed-bandits#robot-learning#academic-seminar#machine-learning#research-seminar#online-seminar#online-learning