activity
20242026
most citedGriffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models

1 citations · 1 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CV2026

TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

Fan Yang, Shurong Zheng, Hongyin Zhao +5

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominan…

cs.CV2026

GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models

Shurong Zheng, Yousong Zhu, Hongyin Zhao +4

Multimodal Large Language Models (MLLMs) have demonstrated impressive progress in single-image grounding and general multi-image understanding. Recently, some methods begin to addr…

cs.CV2025

FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation

Fan Yang, Yousong Zhu, Xin Li +6

Recent Large Vision Language Models (LVLMs) demonstrate promising capabilities in unifying visual understanding and generative modeling, enabling both accurate content understandin…

cs.CV2025

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

Yufei Zhan, Hongyin Zhao, Yousong Zhu +4

Large Multimodal Models (LMMs) have recently demonstrated remarkable visual understanding performance on both vision-language and vision-centric tasks. However, they often fall sho…

cs.CV2025

Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Yufei Zhan, Yousong Zhu, Shurong Zheng +4

Large Vision-Language Models (LVLMs) typically follow a two-stage training paradigm-pretraining and supervised fine-tuning. Recently, preference optimization, derived from the lang…

cs.CV2024★ 1 cited

Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models

Yufei Zhan, Hongyin Zhao, Yousong Zhu +3

Large Multimodal Models (LMMs) have achieved significant breakthroughs in various vision-language and vision-centric tasks based on auto-regressive modeling. However, these models…