2 citations · 3 across the 9 of their papers we have counts for
11 papers · 1 filter
Dynamic Hub-and-Spoke Memory for Streaming Video Understanding
Xinru Jiang, Lin Zhao, Xi Xiao +7
Streaming video understanding requires answering questions at arbitrary times over a continuously growing visual stream. The central challenge is to compactly remember long-range h…
Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs
Xi Xiao, Chen Liu, Chih-Ting Liao +9
Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text. Despite inheriting strong reason…
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Juyi Lin, Arash Akbari, Yumei He +13
Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, ev…
Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation
Xi Xiao, Chenrui Ma, Yunbei Zhang +7
Low-Rank Adaptation (LoRA) has become a cornerstone of parameter-efficient fine-tuning (PEFT). Yet, its efficacy is hampered by two fundamental limitations: semantic drift, by trea…
HIERAMP: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
Lin Zhao, Xinru Jiang, Xi Xiao +7
Dataset distillation often prioritizes global semantic proximity when creating small surrogate datasets for original large-scale ones. However, object semantics are inherently hier…
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
Lin Zhao, Yushu Wu, Aleksei Lebedev +11
Diffusion Transformers (DiTs) have recently improved video generation quality. However, their heavy computational cost makes real-time or on-device generation infeasible. In this w…