13 papers
Symbolic and Abstractive Reasoning with Complex Visual Queries
Yichi Zhang, Jingdian Lu, Zhuo Chen +4
Understanding and reasoning over abstract visual content remains a challenge for current multi-modal large language models (MLLMs). In this paper, we explore a novel abstract data…
Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning
Jingzhi Chen, Landi He, Zhuo Chen +2
The processing of gigapixel whole slide images within vision language models faces a major difficulty due to an excessive number of visual tokens. Existing solutions typically rely…
Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models
Yang Zhou, Ranajoy Sadhukhan, Zhaofeng Sun +7
Despite being powerful, reinforcement learning with verifiable rewards (RLVR) induces extremely long COT, making it computationally expensive. Since RLVR per-step cost is dominated…
UniHR: Hierarchical Representation Learning for Unified Knowledge Graph Link Prediction
Zhiqiang Liu, Yin Hua, Mingyang Chen +4
Real-world knowledge graphs (KGs) contain not only standard triple-based facts, but also more complex, heterogeneous types of facts, such as hyper-relational facts with auxiliary k…
Every Little Helps: Building Knowledge Graph Foundation Model with Fine-grained Transferable Multi-modal Tokens
Yichi Zhang, Zhuo Chen, Lingbing Guo +2
Multi-modal knowledge graph reasoning (MMKGR) aims to predict the missing links by exploiting both graph structure information and multi-modal entity contents. Most existing works…
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
Yichi Zhang, Zhuo Chen, Lingbing Guo +4
Multi-modal large language models (MLLMs) incorporate heterogeneous modalities into LLMs, enabling a comprehensive understanding of diverse scenarios and objects. Despite the proli…