works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.LG2026

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

Enjun Du, Hange Zhou, Chenxu Du +4

The paper introduces LedgerMind, a framework that records and constrains the evidence used by multimodal agents during visual question answering, ensuring that each reasoning step…

cs.CV2026

MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts

Hao Liang, Linzhuang Sun, Minxuan Zhou +7

With the rapid progress of Multimodal LLMs, evaluating their mathematical reasoning capabilities has become an increasingly important research direction. In particular, visual-text…

cs.LG2026

Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT

Jiacheng Wang, Ping Jian, Zhen Yang +3

Large language models (LLMs) have made rapid progress, yet adapting them to downstream scenarios still commonly relies on supervised fine-tuning (SFT). When downstream data exhibit…

cs.IR2026

What Should I Cite? A RAG Benchmark for Academic Citation Prediction

Leqi Zheng, Jiajun Zhang, Canzhi Chen +13

With the rapid growth of Web-based academic publications, more and more papers are being published annually, making it increasingly difficult to find relevant prior work. Citation…

cs.LG2025

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning

Zhen Hao Wong, Jingwen Deng, Runming He +7

Large language models (LLMs) excel at many supervised tasks but often struggle with structured reasoning in unfamiliar settings. This discrepancy suggests that standard fine-tuning…

cs.CV2025

EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation

Hao Liang, Zirong Chen, Hejun Dong +1

Video question-answering (QA) is a core task in video understanding. Evaluating the quality of video QA and video caption data quality for training video large language models (Vid…