most citedRevisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

1 citations · 3 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Spatia: Video Generation with Updatable Spatial Memory

Jinjing Zhao, Fangyun Wei, Zhening Liu +3

Existing video generation models struggle to maintain long-term spatial and temporal consistency due to the dense, high-dimensional nature of video signals. To overcome this limita…

cs.CV2025

CustomX: Unified Character, Action, and Scene Customization in Video World Models

Yitong Wang, Fangyun Wei, Hongyang Zhang +2

Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, whic…

cs.CV2025

From Virtual Games to Real-World Play

Wenqiang Sun, Fangyun Wei, Jinjing Zhao +5

We introduce RealPlay, a neural network-based real-world game engine that enables interactive video generation from user control signals. Unlike prior works focused on game-style v…

cs.CV2024

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

Zhantao Yang, Ruili Feng, Keyu Yan +13

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…

cs.CV20241 cited

Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Jierun Chen, Fangyun Wei, Jinjing Zhao +5

Referring expression comprehension (REC) involves localizing a target instance based on a textual description. Recent advancements in REC have been driven by large multimodal model…