[Imperfect Information Learning Team Seminar] Prof. Micah Goldblum: Easy-to-Hard Generalization and Small Batch Training for LLMs
Columbia University's Micah Goldblum on easy-to-hard generalization and why small-batch SGD keeps pace with Adam for LLM pretraining.
- When
- Thu, August 27, 2026 · 16:00–17:00 JST
- Where
- Online
- Organizer
- RIKEN Center for Advanced Intelligence Project
- Language
- EN
- Source
- Doorkeeper
Summary
Micah Goldblum, assistant professor of electrical engineering at Columbia University, gives a one-hour research talk covering two threads of his work on neural networks. The first is easy-to-hard generalization: how to build models that solve problems far more complex than the ones they were trained on by letting them "think" for longer at inference time. The second is optimization stability for language models, where his recent results show that plain SGD without momentum is nearly as fast as Adam in the small-batch regime of LLM pretraining.
Goldblum was previously a postdoctoral researcher at New York University working with Yann LeCun and Andrew Gordon Wilson. His research spans applied and fundamental machine learning, including training, architecture, and inference strategies for large-scale models, AI safety, agents, and the theory of why complex AI systems work at all.
The seminar runs on Zoom and is open to everyone who registers; the meeting link goes out to registered participants only. An on-site viewing room at the RIKEN Nihonbashi Office is also available, but that option is restricted to RIKEN members.
About the community
This seminar series runs regular research talks from the imperfect information learning group, inviting machine learning researchers from Japan and abroad to present current work. Sessions are single-speaker talks with a technical abstract published in advance, aimed at researchers, graduate students, and engineers who follow the machine learning literature closely. Talks are held in English and streamed on Zoom, so anyone who registers can join remotely.
#llm#machine-learning#deep-learning#optimization#generalization#research-seminar#online