5 papers
S2Aligner: Pair-Efficient and Transferable Pre-Training for Sparse Text-Attributed Graphs
Yuhan Wang, Haopeng Zhang, Yibo Ding +6
Pre-training on text-attributed graphs (TAGs) is central to building transferable graph foundation models, where LLM-as-Aligner methods align graph and text representations through…
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
Mengqi Shi, Haopeng Zhang
While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over complex narratives remains…
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization
Yubo Jiang, Yitong An, Xin Yang +7
We introduce V-tableR1, a process-supervised reinforcement learning framework that elicits rigorous, verifiable reasoning from multimodal large language models (MLLMs). Current MLL…
MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
Zizhong Li, Haopeng Zhang, Jiawei Zhang
Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and lo…
Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
Zizhong Li, Haopeng Zhang, Jiawei Zhang
While large language models (LLMs) have achieved remarkable success in providing trustworthy responses for knowledge-intensive tasks, they still face critical limitations such as h…