1 citations · 2 across the 5 of their papers we have counts for
5 papers
Multimodal Road Network Generation Based on Large Language Model
Jiajing Chen, Weihang Xu, Haiming Cao +4
With the increasing popularity of ChatGPT, large language models (LLMs) have demonstrated their capabilities in communication and reasoning, promising for transportation sector int…
Cut-and-Paste: Subject-Driven Video Editing with Attention Control
Zhichao Zuo, Zhao Zhang, Yan Luo +4
This paper presents a novel framework termed Cut-and-Paste for real-word semantic video editing under the guidance of text prompt and additional reference image. While the text-dri…
Link-Context Learning for Multimodal LLMs
Yan Tai, Weichen Fan, Zhao Zhang +3
The ability to learn from context with novel concepts, and deliver appropriate responses are essential in human conversations. Despite current Multimodal Large Language Models (MLL…
Advancing Referring Expression Segmentation Beyond Single Image
Yixuan Wu, Zhao Zhang, Xie Chi +2
Referring Expression Segmentation (RES) is a widely explored multi-modal task, which endeavors to segment the pre-existing object within a single image with a given linguistic expr…
Co-Salient Object Detection with Co-Representation Purification
Ziyue Zhu, Zhao Zhang, Zheng Lin +2
Co-salient object detection (Co-SOD) aims at discovering the common objects in a group of relevant images. Mining a co-representation is essential for locating co-salient objects.…