collaborators

10 papers

cs.CV2026

SAMTok: Representing Any Mask with Two Words

Yikang Zhou, Tao Zhang, Dengxian Gong +13

Pixel-wise capabilities are essential for building interactive intelligent systems. However, pixel-wise multi-modal LLMs (MLLMs) remain difficult to scale due to complex region-lev…

cs.CV2026

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

Jianzong Wu, Hao Lian, Jiongfan Yang +12

Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier field. Existing unified framewo…

cs.CV2026

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

Junjie Wang, Xinghua Lou, Jason Li +8

Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm…

cs.AI2026

Large Vision-Language Models Get Lost in Attention

Gongli Xi, Ye Tian, Mengyu Yang +5

Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer…

cs.AI2026

FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis

Woojin Lee, Pranav Mekkoth, Ye Tian +2

The widespread adoption of camera-equipped mobile devices and wearables has enabled convenient capture of meal images, making food recognition a key component for real time dietary…

cs.CV2026

Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

Haochen Wang, Yuhao Wang, Tao Zhang +13

While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle in capturing the dense world with complex scenes, requiring fine-grained analysis of i…