From the 1 of 11 linked papers with an AI index.
11 papers
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
Ting Lei, Jialin Liu, Zhu Xu +2
The paper introduces AgentHOI, a training‑free framework that leverages multimodal large language models to detect human‑object interactions in open‑world settings by using iterati…
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
Minghang Zheng, Zihao Yin, Yi Yang +2
Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causi…
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
Jiayi Gao, Changcheng Hua, Qingchao Chen +2
Identity-preserving text-to-video (IPT2V) generation creates videos faithful to both a reference subject image and a text prompt. While fine-tuning large pretrained video diffusion…
Investigating Domain Gaps for Indoor 3D Object Detection
Zijing Zhao, Zhu Xu, Qingchao Chen +2
As a fundamental task for indoor scene understanding, 3D object detection has been extensively studied, and the accuracy on indoor point cloud data has been substantially improved.…
Interact-Custom: Customized Human Object Interaction Image Generation
Zhu Xu, Zhaowen Wang, Yuxin Peng +1
Compositional Customized Image Generation aims to customize multiple target concepts within generation content, which has gained attention for its wild application. Existing approa…
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
Wentao Mo, Qingchao Chen, Yuxin Peng +2
The advancement of 3D vision-language (3D VL) learning is hindered by several limitations in existing 3D VL datasets: they rarely necessitate reasoning beyond a close range of obje…