12 papers
Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning
Yang Liu, Qianqian Xu, Peisong Wen +2
Recent advances in large vision-language models have expanded video retrieval from simple text-based search to more flexible scenarios, where users may specify the desired result t…
The Devil is in the Condition Numbers: Why is GLU Better than non-GLU Structure?
Xingyu Lyu, Qianqian Xu, Zhiyong Yang +2
Gated Linear Units (GLU) and their variants are widely adopted in modern open-source large language model architectures and consistently outperform their non-gated counterparts, ye…
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
Yang Liu, Qianqian Xu, Peisong Wen +3
Recent studies have made notable progress in video representation learning by transferring image-pretrained models to video tasks, typically with complex temporal modules and video…
HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
Zhiguang Lu, Qianqian Xu, Peisong Wen +2
Generative diffusion models show promise for data augmentation. However, applying them to fine-grained tasks presents a significant challenge: ensuring synthetic images accurately…
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
Yang Liu, Xilin Zhao, Peisong Wen +2
Recent progress in video generation has led to impressive visual quality, yet current models still struggle to produce results that align with real-world physical principles. To th…
TuckA: Hierarchical Compact Tensor Experts for Efficient Fine-Tuning
Qifeng Lei, Zhiyong Yang, Qianqian Xu +3
Efficiently fine-tuning pre-trained models for downstream tasks is a key challenge in the era of foundation models. Parameter-efficient fine-tuning (PEFT) presents a promising solu…