activity
20242026
most citedMiniCPM-V: A GPT-4V Level MLLM on Your Phone

26 citations · 49 across the 9 of their papers we have counts for

collaborators

10 papers

cs.CV2026

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs

Zhixiao Zheng, Zheren Fu, Zhiyuan Yao +3

Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual hallucinations, content fabrication, and unfait…

cs.CV2026

ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs

Zhiyuan Yao, Zheren Fu, Zhixiao Zheng +3

Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image. In this paper, we identify an internal s…

cs.AI2026

One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

Zhi Zheng, Ziqiao Meng, Hao Luan +2

External memory effectively grounds large language models (LLMs) and vision-language models (VLMs)-based question answering (QA) in relevant multimodal evidence. However, existing…

cs.LG2026

Multi-Gate Residuals

Zhizhan Zheng, Feiyun Zhang, Shuchun Liu +4

While Attention Residuals has shown some effectiveness in addressing the widespread issue of unbounded activation growth across deep residual layers, it inevitably incurs significa…

cs.AI2026

Data Science and Technology Towards AGI Part I: Tiered Data Management

Yudong Wang, Zixuan Fu, Hengyu Zhao +14

The development of artificial intelligence can be viewed as an evolution of data-driven learning paradigms, with successive shifts in data organization and utilization continuously…

cs.CL2025

Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data

Yudong Wang, Zixuan Fu, Jie Cai +9

Data quality has become a key factor in enhancing model performance with the rapid development of large language models (LLMs). Model-driven data filtering has increasingly become…