collaborators

12 papers

cs.AI2025

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

Zirui Shao, Feiyu Gao, Zhaoqing Zhu +6

Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand. As a mult…

cs.LG2025

Making Classic GNNs Strong Baselines Across Varying Homophily: A Smoothness-Generalization Perspective

Ming Gu, Zhuonan Zheng, Sheng Zhou +5

Graph Neural Networks (GNNs) have achieved great success but are often considered to be challenged by varying levels of homophily in graphs. Recent \textit{empirical} studies have…

cs.CL2025

BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks

Tianyuan Huang, Zepeng Zhu, Hangdi Xing +6

Braille plays a vital role in education and information accessibility for visually impaired individuals. However, Braille information processing faces challenges such as data scarc…

cs.CV2025

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback

Yang Chen, Yufan Shen, Wenxuan Huang +7

Multimodal Large Language Models (MLLMs) exhibit impressive performance across various visual tasks. Subsequent investigations into enhancing their visual reasoning abilities have…

cs.LG2025

Correlation-Aware Graph Convolutional Networks for Multi-Label Node Classification

Yuanchen Bei, Weizhi Chen, Hao Chen +5

Multi-label node classification is an important yet under-explored domain in graph mining as many real-world nodes belong to multiple categories rather than just a single one. Alth…

cs.LG2025

OpenGT: A Comprehensive Benchmark For Graph Transformers

Jiachen Tang, Zhonghao Wang, Sirui Chen +3

Graph Transformers (GTs) have recently demonstrated remarkable performance across diverse domains. By leveraging attention mechanisms, GTs are capable of modeling long-range depend…