2 citations · 6 across the 12 of their papers we have counts for
7 papers · 1 filter
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
Fan Yang, Shurong Zheng, Hongyin Zhao +5
Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominan…
CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension
Lihao Liu, Yan Wang, Biao Yang +10
Multimodal Large Language Models (MLLMs) have shown remarkable success in comprehension tasks such as visual description and visual question answering. However, their direct applic…
UniRef-Image-Edit: Towards Scalable and Consistent Multi-Reference Image Editing
Hongyang Wei, Bin Wen, Yancheng Long +22
We present UniRef-Image-Edit, a high-performance multi-modal generation system that unifies single-image editing and multi-image composition within a single framework. Existing dif…
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
Fan Yang, Xingping Dong, Xin Yu +3
Understanding high-resolution (HR) images remains a critical challenge for multimodal large language models (MLLMs). Recent approaches leverage vision-based retrieval-augmented gen…
FilmSceneDesigner: Chaining Set Design for Procedural Film Scene Generation
Zhifeng Xie, Keyi Zhang, Yiye Yan +4
Film set design plays a pivotal role in cinematic storytelling and shaping the visual atmosphere. However, the traditional process depends on expert-driven manual modeling, which i…
Logics-Parsing Technical Report
Xiangyang Chen, Shuzhao Li, Xiuwen Zhu +7
Recent advances in Large Vision-Language models (LVLM) have spurred significant progress in document parsing task. Compared to traditional pipeline-based methods, end-to-end paradi…