8 papers · 1 filter
Spatial-Conditioned Reasoning in Long-Egocentric Videos
James Tribble, Hao Wang, Si-En Hong +4
Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-l…
Motion Focus Recognition in Fast-Moving Egocentric Video
Si-En Hong, James Tribble, Alexander Lake +8
From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of moti…
RobustFormer: Noise-Robust Pre-training for images and videos
Ashish Bastola, Nishant Luitel, Hao Wang +3
While deep learning-based models like transformers, have revolutionized time-series and vision tasks, they remain highly susceptible to noise and often overfit on noisy patterns ra…
Fast 2DGS: Efficient Image Representation with Deep Gaussian Prior
Hao Wang, Ashish Bastola, Chaoyi Zhou +5
As generative models become increasingly capable of producing high-fidelity visual content, the demand for efficient, interpretable, and editable image representations has grown su…
AtomDiffuser: Time-Aware Degradation Modeling for Drift and Beam Damage in STEM Imaging
Hao Wang, Hongkui Zheng, Kai He +1
Scanning transmission electron microscopy (STEM) plays a critical role in modern materials science, enabling direct imaging of atomic structures and their evolution under external…
Diffusion Prism: Enhancing Diversity and Morphology Consistency in Mask-to-Image Diffusion
Hao Wang, Xiwen Chen, Ashish Bastola +2
The emergence of generative AI and controllable diffusion has made image-to-image synthesis increasingly practical and efficient. However, when input images exhibit low entropy and…