activity
20222025
most citedPrototype-based Embedding Network for Scene Graph Generation

5 citations · 14 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2025

FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models

Shengming Yuan, Xinyu Lyu, Shuailong Wang +3

Multimodal large language models (MLLMs) face an inherent trade-off between faithfulness and creativity, as different tasks require varying degrees of associative reasoning. Howeve…

cs.CV2025

Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation

Beitao Chen, Xinyu Lyu, Lianli Gao +2

Despite their success, Large Vision-Language Models (LVLMs) remain vulnerable to hallucinations. While existing studies attribute the cause of hallucinations to insufficient visual…

cs.CV2024★ 1 cited

Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization

Xinyu Lyu, Beitao Chen, Lianli Gao +2

Although Large Visual Language Models (LVLMs) have demonstrated exceptional abilities in understanding multimodal data, they invariably suffer from hallucinations, leading to a dis…

cs.CV2024★ 2 cited

Text-Video Retrieval with Global-Local Semantic Consistent Learning

Haonan Zhang, Pengpeng Zeng, Lianli Gao +4

Adapting large-scale image-text pre-training models, e.g., CLIP, to the video domain represents the current state-of-the-art for text-video retrieval. The primary approaches involv…

cs.CV2023

ALF: Adaptive Label Finetuning for Scene Graph Generation

Qishen Chen, Jianzhi Liu, Xinyu Lyu +3

Scene Graph Generation (SGG) endeavors to predict the relationships between subjects and objects in a given image. Nevertheless, the long-tail distribution of relations often leads…

cs.CV2023

Local-Global Information Interaction Debiasing for Dynamic Scene Graph Generation

Xinyu Lyu, Jingwei Liu, Yuyu Guo +1

The task of dynamic scene graph generation (DynSGG) aims to generate scene graphs for given videos, which involves modeling the spatial-temporal information in the video. However,…