12 papers
EEA: Exploration-Exploitation Agent for Long Video Understanding
Te Yang, Xiangyu Zhu, Bo Wang +3
Long-form video understanding requires efficient navigation of extensive visual data to pinpoint sparse yet critical information. Current approaches to longform video understanding…
UniField: Joint Multi-Domain Training for Universal Surface Pressure Modeling
Junhong Zou, Zhenxu Sun, Yueqing Wang +4
Accurate modeling of surface pressure fields around objects is fundamental to aerodynamic analysis and design. While neural networks have shown promise as efficient alternatives to…
Towards Realistic Hand-Object Interaction with Gravity-Field Based Diffusion Bridge
Miao Xu, Xiangyu Zhu, Xusheng Liang +3
Existing reconstruction or hand-object pose estimation methods are capable of producing coarse interaction states. However, due to the complex and diverse geometry of both human ha…
Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
Bao Li, Xiaomei Zhang, Miao Xu +3
Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal la…
RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution
Weisong Zhao, Jingkai Zhou, Xiangyu Zhu +4
Video Super-Resolution (VSR) has achieved significant progress through diffusion models, effectively addressing the over-smoothing issues inherent in GAN-based methods. Despite rec…
MLLM-Enhanced Face Forgery Detection: A Vision-Language Fusion Solution
Siran Peng, Zipei Wang, Li Gao +5
Reliable face forgery detection algorithms are crucial for countering the growing threat of deepfake-driven disinformation. Previous research has demonstrated the potential of Mult…