7 papers
LLM4Fluid: Large Language Models as Generalizable Neural Solvers for Fluid Dynamics
Qisong Xiao, Xinhai Chen, Qinglin Wang +10
Deep learning has emerged as a promising paradigm for spatio-temporal modeling of fluid dynamics. However, existing approaches often suffer from limited generalization to unseen fl…
Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
Liwei Luo, Shuaitengyuan Li, Dongwei Ren +3
Recently, remarkable progress has been made in large-scale pre-trained model tuning, and inference efficiency is becoming more crucial for practical deployment. Early exiting in co…
Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
Tianming Liang, Haichao Jiang, Yuting Yang +4
Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent y…
Learning a Neural Association Network for Self-supervised Multi-Object Tracking
Shuai Li, Michael Burke, Subramanian Ramamoorthy +1
This paper introduces a novel framework to learn data association for multi-object tracking in a self-supervised manner. Fully-supervised learning methods are known to achieve exce…
Exploring Dynamic Transformer for Efficient Object Tracking
Jiawen Zhu, Xin Chen, Haiwen Diao +6
The speed-precision trade-off is a critical problem for visual object tracking which usually requires low latency and deployment on constrained resources. Existing solutions for ef…
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Han Wang, Yuxiang Nie, Yongjie Ye +6
The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in…