most citedLingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

5 citations · 8 across the 6 of their papers we have counts for

collaborators

14 papers

cs.LG2025

ParaFormer: A Generalized PageRank Graph Transformer for Graph Representation Learning

Chaohao Yuan, Zhenjie Song, Ercan Engin Kuruoglu +5

Graph Transformers (GTs) have emerged as a promising graph learning tool, leveraging their all-pair connected property to effectively capture global information. To address the ove…

cs.CV2025

From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models

Zongzhao Li, Xiangzhe Kong, Jiahui Su +8

This paper introduces the concept of Microscopic Spatial Intelligence (MiSI), the capability to perceive and reason about the spatial relationships of invisible microscopic entitie…

cs.LG2025

Universally Invariant Learning in Equivariant GNNs

Jiacheng Cen, Anyi Li, Ning Lin +5

Equivariant Graph Neural Networks (GNNs) have demonstrated significant success across various applications. To achieve completeness -- that is, the universal approximation property…

cs.CL2025

GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning

Guizhen Chen, Weiwen Xu, Hao Zhang +4

Recent advancements in reinforcement learning (RL) have enhanced the reasoning abilities of large language models (LLMs), yet the impact on multimodal LLMs (MLLMs) is limited. Part…

cs.CV2025

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning

Ruifeng Yuan, Chenghao Xiao, Sicong Leng +9

Reinforcement learning has proven its effectiveness in enhancing the reasoning capabilities of large language models. Recent research efforts have progressively extended this parad…

cs.LG2025

DiffSpectra: Molecular Structure Elucidation from Spectra using Diffusion Models

Liang Wang, Yu Rong, Tingyang Xu +7

Molecular structure elucidation from spectra is a fundamental challenge in molecular science. Conventional approaches rely heavily on expert interpretation and lack scalability, wh…