most citedFrom Data to Modeling: Fully Open-vocabulary Scene Graph Generation

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2025

MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites

Zhenxin Lei, Zhangwei Gao, Changyao Tian +12

Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains.…

cs.LG2025

PCaM: A Progressive Focus Attention-Based Information Fusion Method for Improving Vision Transformer Domain Adaptation

Zelin Zang, Fei Wang, Liangyu Li +4

Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Recent UDA methods based on Vision Transformers (ViTs) h…

cs.CV20251 cited

From Data to Modeling: Fully Open-vocabulary Scene Graph Generation

Zuyao Chen, Jinlin Wu, Zhen Lei +1

We present OvSGTR, a novel transformer-based framework for fully open-vocabulary scene graph generation that overcomes the limitations of traditional closed-set models. Conventiona…

cs.CV2025

Compile Scene Graphs with Reinforcement Learning

Zuyao Chen, Jinlin Wu, Zhen Lei +2

Next-token prediction is the fundamental principle for training large language models (LLMs), and reinforcement learning (RL) further enhances their reasoning performance. As an ef…

cs.CV2025

SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking

Yingjia Xu, Jinlin Wu, Daming Gao +5

Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven c…

cs.CV2024

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation

Zuyao Chen, Jinlin Wu, Zhen Lei +1

While text-to-image generation has been extensively studied, generating images from scene graphs remains relatively underexplored, primarily due to challenges in accurately modelin…