From the 1 of 5 linked papers with an AI index.
5 papers
The TIME Machine: On The Power of Motion for Efficient Perception
Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
The paper introduces TIME, a video representation learned from motion point-tracks using a masked autoencoder, enabling self‑supervised, language‑free training that requires far le…
It's a Matter of Time: Three Lessons on Long-Term Motion for Perception
Willem Davison, Xinyue Hao, Laura Sevilla-Lara
Temporal information has long been considered to be essential for perception. While there is extensive research on the role of image information for perceptual tasks, the role of t…
Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts
Aiden Yiliu Li, Xinyue Hao, Shilong Liu +1
Despite advances in multimodal large language models, autonomous web agents still struggle to reliably execute long-horizon tasks on complex and dynamic web interfaces. Existing ag…
Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
Shriram M Sathiyanarayanan, Xinyue Hao, Shihao Hou +4
The success of the machine learning field has reliably depended on training on large datasets. While effective, this trend comes at an extraordinary cost. This is due to two deeply…
Principles of Visual Tokens for Efficient Video Understanding
Xinyue Hao, Gen Li, Shreyank N Gowda +4
Video understanding has made huge strides in recent years, relying largely on the power of transformers. As this architecture is notoriously expensive and video data is highly redu…