activity
20172025
most citedIntegrating both Visual and Audio Cues for Enhanced Video Caption

11 citations · 19 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CL2025

Improving Generalization in LLM Structured Pruning via Function-Aware Neuron Grouping

Tao Yu, Yongqi An, Kuan Zhu +3

Large Language Models (LLMs) demonstrate impressive performance across natural language tasks but incur substantial computational and storage costs due to their scale. Post-trainin…

cs.CV2025

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

Peng Xu, Shengwu Xiong, Jiajun Zhang +125

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…

cs.CV2025

AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection

Zhaopeng Gu, Bingke Zhu, Guibo Zhu +4

Anomaly detection is a critical task across numerous domains and modalities, yet existing methods are often highly specialized, limiting their generalizability. These specialized m…

cs.CV2024

Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive Manner

Pengxiang Cai, Zhiwei Liu, Guibo Zhu +2

Pixel-level fine-grained image editing remains an open challenge. Previous works fail to achieve an ideal trade-off between control granularity and inference speed. They either fai…

cs.CL2024

Recurrent Context Compression: Efficiently Expanding the Context Window of LLM

Chensen Huang, Guibo Zhu, Xuepeng Wang +5

To extend the context length of Transformer-based large language models (LLMs) and improve comprehension capabilities, we often face limitations due to computational resources and…

cs.CV2024

BFRFormer: Transformer-based generator for Real-World Blind Face Restoration

Guojing Ge, Qi Song, Guibo Zhu +5

Blind face restoration is a challenging task due to the unknown and complex degradation. Although face prior-based methods and reference-based methods have recently demonstrated hi…