JapanTech

[Imperfect Information Learning Team Seminar] Prof. Micah Goldblum: Easy-to-Hard Generalization and Small Batch Training for LLMs

Columbia University's Micah Goldblum on easy-to-hard generalization and why small-batch SGD keeps pace with Adam for LLM pretraining.

When
Thu, August 27, 2026 · 16:00–17:00 JST
Where
Online
Organizer
RIKEN Center for Advanced Intelligence Project
Language
EN
Source
Doorkeeper
Summary
Micah Goldblum, assistant professor of electrical engineering at Columbia University, gives a one-hour research talk covering two threads of his work on neural networks. The first is easy-to-hard generalization: how to build models that solve problems far more complex than the ones they were trained on by letting them "think" for longer at inference time. The second is optimization stability for language models, where his recent results show that plain SGD without momentum is nearly as fast as Adam in the small-batch regime of LLM pretraining. Goldblum was previously a postdoctoral researcher at New York University working with Yann LeCun and Andrew Gordon Wilson. His research spans applied and fundamental machine learning, including training, architecture, and inference strategies for large-scale models, AI safety, agents, and the theory of why complex AI systems work at all. The seminar runs on Zoom and is open to everyone who registers; the meeting link goes out to registered participants only. An on-site viewing room at the RIKEN Nihonbashi Office is also available, but that option is restricted to RIKEN members.
About the community

This seminar series runs regular research talks from the imperfect information learning group, inviting machine learning researchers from Japan and abroad to present current work. Sessions are single-speaker talks with a technical abstract published in advance, aimed at researchers, graduate students, and engineers who follow the machine learning literature closely. Talks are held in English and streamed on Zoom, so anyone who registers can join remotely.

#llm#machine-learning#deep-learning#optimization#generalization#research-seminar#online