JapanTech

[Robot Learning Team Seminar] Talks by Prof. Matteo Papini (University of Milan) and Dr. Gianmarco Genalti (Politecnico di Milano) on July 23

RIKEN AIP online seminar: two Italian researchers on policy gradient sample efficiency and graph-triggered bandits, July 23.

When
Thu, July 23, 2026 · 15:00–16:30 JST
Where
Open Space at the RIKEN Nihonbashi Office (on-site for RIKEN members only) · Online
Organizer
RIKEN Center for Advanced Intelligence Project
Language
EN
Source
Doorkeeper
Summary
This is an online seminar hosted by the Robot Learning Team, with an on-site option at RIKEN's Nihonbashi office limited to RIKEN members. Two invited talks cover recent theoretical work in reinforcement learning. Prof. Matteo Papini (University of Milan) presents "Reusing Data in Policy Gradients to Improve Sample Efficiency," introducing an actor-only policy gradient algorithm based on a multiple importance sampling estimator that reuses trajectory data from the k most recent policies, with sample efficiency guarantees and early results on data reuse in Proximal Policy Optimization. Dr. Gianmarco Genalti (Politecnico di Milano) presents "Bridging Rested and Restless Bandits with Graph-Triggering," proposing Graph-Triggered Bandits, a framework that generalizes rested and restless bandit settings using a graph over arms to model how pulling one arm triggers changes in another, with algorithms and guarantees for rising and rotting bandit cases.
About the community

This seminar series brings international researchers to present cutting-edge work in reinforcement learning and sequential decision-making to an audience of AI researchers and engineers in Japan. Talks combine formal theoretical results with practical algorithmic guarantees, and sessions are held online via Zoom so remote participants across Japan can join alongside a small in-person audience.

#reinforcement-learning#policy-gradient#multi-armed-bandits#robot-learning#academic-seminar#machine-learning#research-seminar#online-seminar#online-learning