collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback

Yang Chen, Yufan Shen, Wenxuan Huang +7

Multimodal Large Language Models (MLLMs) exhibit impressive performance across various visual tasks. Subsequent investigations into enhancing their visual reasoning abilities have…

cs.CV2025

FocusedAD: Character-centric Movie Audio Description

Xiaojun Ye, Chun Wang, Yiren Song +3

Movie Audio Description (AD) aims to narrate visual content during dialogue-free segments, particularly benefiting blind and visually impaired (BVI) audiences. Compared with genera…

cs.CV2025

MP-GUI: Modality Perception with MLLMs for GUI Understanding

Ziwei Wang, Weizhi Chen, Leyang Yang +7

Graphical user interface (GUI) has become integral to modern society, making it crucial to be understood for human-centric systems. However, unlike natural images or documents, GUI…

cs.CV2025

ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data

Yufan Shen, Chuwei Luo, Zhaoqing Zhu +5

Recently, large language models (LLMs) and multimodal large language models (MLLMs) have demonstrated promising results on document visual question answering (VQA) task, particular…

cs.CV2024

FGP: Feature-Gradient-Prune for Efficient Convolutional Layer Pruning

Qingsong Lv, Jiasheng Sun, Sheng Zhou +6

To reduce computational overhead while maintaining model performance, model pruning techniques have been proposed. Among these, structured pruning, which removes entire convolution…

cs.CV2024

WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation

Zirui Shao, Feiyu Gao, Hangdi Xing +5

In the era of content creation revolution propelled by advancements in generative models, the field of web design remains unexplored despite its critical role in modern digital com…