collaborators

7 papers

cs.CV2026

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

Yaoting Wang, Ziyi Zhang, Wenming Tu +10

Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI)…

cs.SI2026

Reducing Detail Hallucinations in Long-Context Regulatory Understanding via Targeted Preference Optimization

Yang Liu, Bin Chong, Yuhan Lin +7

Large language models (LLMs) frequently produce \emph{detail hallucinations} when processing long regulatory documents, including subtle errors in threshold values, units, scopes,…

cs.CV2025

"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments

Ziyi Zhang, Zhen Sun, Zongmin Zhang +6

The visually impaired population faces significant challenges in daily activities. While prior works employ vision language models for assistance, most focus on static content and…

cs.CV2025

Can VLMs Detect and Localize Fine-Grained AI-Edited Images?

Zhen Sun, Ziyi Zhang, Zeren Luo +10

Fine-grained detection and localization of localized image edits is crucial for assessing content authenticity, especially as modern diffusion models and image editors can produce…

cs.CV2025

FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts

Ziyi Zhang, Zhen Sun, Zongmin Zhang +2

Multimodal Large Language Models (MLLMs) have become powerful and widely adopted in some practical applications. However, recent research has revealed their vulnerability to multim…

cs.AI2025

Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media

Zhen Sun, Zongmin Zhang, Xinyue Shen +5

Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs). However, the misuse of AIGTs could have profound implications for public opinion, such as…