189 citations · 740 across the 72 of their papers we have counts for
107 papers
Evidence-Backed Video Question Answering
Shijie Wang, Honglu Zhou, Ziyang Wang +5
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding.…
Artificial Intelligence Index Report 2026
Sha Sajadieh, Loredana Fattorini, Raymond Perrault +20
Welcome to the ninth edition of the AI Index report. As AI continues to advance rapidly, the question becomes whether the systems built around it can keep up. Governance frameworks…
Linear Scaling Video VLMs for Long Video Understanding
Cristobal Eyzaguirre, Jiajun Wu, Juan Carlos Niebles
Video vision-language models (VLMs) are increasingly used in long-horizon and streaming settings, yet most video encoders still rely on spatiotemporal self-attention, causing compu…
GPIC: A Giant Permissive Image Corpus for Visual Generation
Keshigeyan Chandrasegaran, Kyle Sargent, Suchir Agarwal +6
Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image Corpus of approximately 28 tri…
SurfPhase: 3D Interfacial Dynamics in Two-Phase Flows from Sparse Videos
Yue Gao, Hong-Xing Yu, Sanghyeon Chang +5
Interfacial dynamics in two-phase flows govern momentum, heat, and mass transfer, yet remain difficult to measure experimentally. Classical techniques face intrinsic limitations ne…
Future Optical Flow Prediction Improves Robot Control & Video Generation
Kanchana Ranasinghe, Honglu Zhou, Yu Fang +7
Future motion representations, such as optical flow, offer immense value for control and generative tasks. However, forecasting generalizable spatially dense motion representations…