4 papers
QVGGT: Post-Training Quantized Visual Geometry Grounded Transformer
Zhizhen Pan, Hesong Wang, Huan Wang
Estimating 3D attributes directly from images has advanced rapidly with the Visual Geometry Grounded Transformer (VGGT), which predicts camera parameters, depth maps, and point clo…
EarlyTom: Early Token Compression Completes Fast Video Understanding
Hesong Wang, Xin Jin, Lu Lu +4
Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficien…
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
Keda Tao, Yuhua Zheng, Jia Xu +13
Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily fo…
OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
Junhan Zhu, Hesong Wang, Mingluo Su +2
Large-scale text-to-image diffusion models, while powerful, suffer from prohibitive computational cost. Existing one-shot network pruning methods can hardly be directly applied to…