activity
20212025
most citedMulti-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos

46 citations · 46 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CV2025

Robust Multimodal Large Language Models Against Modality Conflict

Zongmeng Zhang, Wengang Zhou, Jie Zhao +1

Despite the impressive capabilities of multimodal large language models (MLLMs) in vision-language tasks, they are prone to hallucinations in real-world scenarios. This paper inves…

cs.CV2025

Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering

Jianfeng Cai, Wengang Zhou, Zongmeng Zhang +3

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.However, hallucination, where the model generates plausible yet incorrect outputs,…

cs.IR2024

BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language?

Zongmeng Zhang, Jinhua Zhu, Wengang Zhou +3

Dense retrieval, which aims to encode the semantic information of arbitrary text into dense vector representations or embeddings, has emerged as an effective and efficient paradigm…

cs.CL2024

Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning

Zongmeng Zhang, Yufeng Shi, Jinhua Zhu +4

Trustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect…

cs.CV202146 cited

Multi-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos

Zongmeng Zhang, Xianjing Han, Xuemeng Song +2

This paper focuses on tackling the problem of temporal language localization in videos, which aims to identify the start and end points of a moment described by a natural language…