10 papers
Code as Representation: A Compilable Parsing Paradigm for Academic Documents
Rihui Jin, Jun Wang, chengyuan zhu +11
Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For M…
EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
Xinbang Dai, Zheyu Xin, Huikang Hu +7
Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switch…
ChartSync: A Benchmark for Visuo-Logical Cascading Chart Editing
Jiakang Yu, Yixuan Chai, Tianci Wang +7
Generative image editing models struggle with structured statistical charts when data modifications require geometric synchronization. We formalize this task as Visuo-Logical Casca…
FollowTable: A Benchmark for Instruction-Following Table Retrieval
Rihui Jin, Yuchen Lu, Ting Zhang +7
Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic similarity. With the growing adop…
ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding
Xinbang Dai, Huikang Hu, Yongrui Chen +6
While large language models (LLMs) excel at many domain-specific tasks, their ability to deeply comprehend and reason about full-length academic papers remains underexplored. Exist…
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
Xinbang Dai, Huikang Hu, Yuncheng Hua +5
Retrieval-augmented generation (RAG) is a promising paradigm, yet its trustworthiness remains a critical concern. A major vulnerability arises prior to generation: models often fai…