1 citations · 1 across the 4 of their papers we have counts for
6 papers
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
Zhenxin Lei, Zhangwei Gao, Changyao Tian +12
Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains.…
PCaM: A Progressive Focus Attention-Based Information Fusion Method for Improving Vision Transformer Domain Adaptation
Zelin Zang, Fei Wang, Liangyu Li +4
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Recent UDA methods based on Vision Transformers (ViTs) h…
From Data to Modeling: Fully Open-vocabulary Scene Graph Generation
Zuyao Chen, Jinlin Wu, Zhen Lei +1
We present OvSGTR, a novel transformer-based framework for fully open-vocabulary scene graph generation that overcomes the limitations of traditional closed-set models. Conventiona…
Compile Scene Graphs with Reinforcement Learning
Zuyao Chen, Jinlin Wu, Zhen Lei +2
Next-token prediction is the fundamental principle for training large language models (LLMs), and reinforcement learning (RL) further enhances their reasoning performance. As an ef…
SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking
Yingjia Xu, Jinlin Wu, Daming Gao +5
Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven c…
What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation
Zuyao Chen, Jinlin Wu, Zhen Lei +1
While text-to-image generation has been extensively studied, generating images from scene graphs remains relatively underexplored, primarily due to challenges in accurately modelin…