11 papers
Spatial-Conditioned Reasoning in Long-Egocentric Videos
James Tribble, Hao Wang, Si-En Hong +4
Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-l…
Motion Focus Recognition in Fast-Moving Egocentric Video
Si-En Hong, James Tribble, Alexander Lake +8
From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of moti…
EZBlender: Efficient 3D Editing with Plan-and-ReAct Agent
Hao Wang, Wenhui Zhu, Shao Tang +8
As a cornerstone of the modern digital economy, 3D modeling and rendering demand substantial resources and manual effort when scene editing is performed in the traditional manner.…
RobustFormer: Noise-Robust Pre-training for images and videos
Ashish Bastola, Nishant Luitel, Hao Wang +3
While deep learning-based models like transformers, have revolutionized time-series and vision tasks, they remain highly susceptible to noise and often overfit on noisy patterns ra…
Fast 2DGS: Efficient Image Representation with Deep Gaussian Prior
Hao Wang, Ashish Bastola, Chaoyi Zhou +5
As generative models become increasingly capable of producing high-fidelity visual content, the demand for efficient, interpretable, and editable image representations has grown su…
Anomalous Decision Discovery using Inverse Reinforcement Learning
Ashish Bastola, Mert D. Pesé, Long Cheng +2
Anomaly detection plays a critical role in Autonomous Vehicles (AVs) by identifying unusual behaviors through perception systems that could compromise safety and lead to hazardous…