7 papers
Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions
Frank Wu, Mengye Ren
The Forward-Forward (FF) Algorithm is a recently proposed learning procedure for neural networks that employs two forward passes instead of the traditional forward and backward pas…
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
Kangsan Kim, Yanlai Yang, Suji Kim +4
As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human…
Opinion: Learning Intuitive Physics May Require More than Visual Data
Ellen Su, Solim Legris, Todd M. Gureckis +1
Humans expertly navigate the world by building rich internal models founded on an intuitive understanding of physics. Meanwhile, despite training on vast quantities of internet vid…
In-Context Clustering with Large Language Models
Ying Wang, Mengye Ren, Andrew Gordon Wilson
We propose In-Context Clustering (ICC), a flexible LLM-based procedure for clustering data from diverse distributions. Unlike traditional clustering algorithms constrained by prede…
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
Yanlai Yang, Zhuokai Zhao, Satya Narayan Shukla +4
Multimodal large language models (MLLMs) have made significant progress in visual-language reasoning, but their ability to efficiently handle long videos remains limited. Despite r…
Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos
Yanlai Yang, Mengye Ren
Self-supervised learning holds the promise of learning good representations from real-world continuous uncurated data streams. However, most existing works in visual self-supervise…