From the 1 of 8 linked papers with an AI index.
8 papers
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
Ting Lei, Jialin Liu, Zhu Xu +2
The paper introduces AgentHOI, a training‑free framework that leverages multimodal large language models to detect human‑object interactions in open‑world settings by using iterati…
Investigating Domain Gaps for Indoor 3D Object Detection
Zijing Zhao, Zhu Xu, Qingchao Chen +2
As a fundamental task for indoor scene understanding, 3D object detection has been extensively studied, and the accuracy on indoor point cloud data has been substantially improved.…
Interact-Custom: Customized Human Object Interaction Image Generation
Zhu Xu, Zhaowen Wang, Yuxin Peng +1
Compositional Customized Image Generation aims to customize multiple target concepts within generation content, which has gained attention for its wild application. Existing approa…
TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-enhanced Relation-aware Knowledge Transferring
Zhu Xu, Ting Lei, Zhimin Li +4
Dynamic Scene Graph Generation (DSGG) aims to create a scene graph for each video frame by detecting objects and predicting their relationships. Weakly Supervised DSGG (WS-DSGG) re…
Hierarchical Sub-action Tree for Continuous Sign Language Recognition
Dejie Yang, Zhu Xu, Xinjie Gao +1
Continuous sign language recognition (CSLR) aims to transcribe untrimmed videos into glosses, which are typically textual words. Recent studies indicate that the lack of large data…
Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing
Zhuoying Li, Zhu Xu, Yuxin Peng +1
Instruction-based image editing, which aims to modify the image faithfully according to the instruction while preserving irrelevant content unchanged, has made significant progress…