8 citations · 35 across the 35 of their papers we have counts for
19 papers · 1 filter
OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
Xiwen Chen, Wenhui Zhu, Gen Li +9
Multi-modal large language models (MLLMs) achieve strong visual-language reasoning but suffer from high inference cost due to redundant visual tokens. Recent work explores visual t…
Spatial-Conditioned Reasoning in Long-Egocentric Videos
James Tribble, Hao Wang, Si-En Hong +4
Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-l…
Motion Focus Recognition in Fast-Moving Egocentric Video
Si-En Hong, James Tribble, Alexander Lake +8
From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of moti…
Fast 2DGS: Efficient Image Representation with Deep Gaussian Prior
Hao Wang, Ashish Bastola, Chaoyi Zhou +5
As generative models become increasingly capable of producing high-fidelity visual content, the demand for efficient, interpretable, and editable image representations has grown su…
AtomDiffuser: Time-Aware Degradation Modeling for Drift and Beam Damage in STEM Imaging
Hao Wang, Hongkui Zheng, Kai He +1
Scanning transmission electron microscopy (STEM) plays a critical role in modern materials science, enabling direct imaging of atomic structures and their evolution under external…
How Effective Can Dropout Be in Multiple Instance Learning ?
Wenhui Zhu, Peijie Qiu, Xiwen Chen +4
Multiple Instance Learning (MIL) is a popular weakly-supervised method for various applications, with a particular interest in histological whole slide image (WSI) classification.…