JapanTech

Constrained Autonomy: AI Models, Memory, and Retrieval at the Edge

Three talks on edge AI: local LLM inference economics, agentic memory for long-running agents, and embedded vector search for offline autonomy.

When
Fri, September 4, 2026 · 18:00–21:00 JST
Where
Bunkyo City, Japan · In person
Region
Kanto (Tokyo)
Organizer
Tokyo AI
Language
EN
Source
Luma
Summary
An evening of three technical talks on the engineering realities of running AI on constrained physical systems and edge devices. The program splits the on-device stack into three layers, models, memory, and retrieval, and works through each in turn: what it costs to run an LLM locally, how a long-running agent keeps track of its own state, and how vector search can happen inside the hardware rather than in a datacenter. Paul Willot (Senior MLE, Liquid AI) opens with the economics of inference, looking at how memory bandwidth, model size, hardware utilization, and concurrency set the cost and speed of a token, and at how quantization, sparsity, hybrid architectures, caching, and speculative decoding shift that trade-off. Stefania Druga (Staff Research Scientist, Sakana AI RSI Lab) follows with agentic memory, drawing on experiments with research agents that run for hundreds of turns to compare memory-off, deployed recall, gating, and active sub-goal ranking, and to argue that memory is a policy rather than storage. Ewa Szyszka (DevRel Engineer, Qdrant) closes with Qdrant Edge, an embedded vector search library with an install footprint around 11 MB, demonstrated in a live robot demo where hybrid search over dense vision vectors and BM25 captions returns in under a millisecond. Doors open at 18:00, talks run 18:30 to 20:00, and networking follows until 21:00. The material is aimed at engineers working in robotics, applied ML, and embedded systems who need agentic behavior that stays reliable without a cloud connection.
About the community

The largest international AI community in Japan, with more than 5,000 members based mainly in Tokyo: engineers, researchers, investors, product managers, and corporate innovation leaders. It runs over 80 events a year and has hosted 300+ speakers from startups, enterprises, and academia, with sessions typically built around several deep technical talks followed by open networking. Programming is in English and aimed at connecting people building AI in Japan with the wider global ecosystem.

#ai#edge-ai#llm#ai-agents#vector-search#robotics#embedded-systems#machine-learning#on-device-ai#agentic-ai