activity
20242026
collaborators

18 papers

cs.CV2026

Code as Representation: A Compilable Parsing Paradigm for Academic Documents

Rihui Jin, Jun Wang, chengyuan zhu +11

Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For M…

cs.AI2026

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization

Xinbang Dai, Zheyu Xin, Huikang Hu +7

Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switch…

cs.LG2026

When Hard Negatives Hurt: Bridging the Generative-Discriminative Gap in Hard Negative Synthesis for Retrieval

Zhicheng Zhang, Jiwei Tang, Kuicai Dong +9

Hard negative mining has become the dominant strategy for training retrievers, yet it faces intrinsic limitations: negatives are bounded by corpus availability, selected by retriev…

cs.CL2026

Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding

Jianzhu Bao, Haozhen Zhang, Kuicai Dong +5

Vision-Language Models (VLMs) have demonstrated remarkable progress in chart understanding, largely driven by supervised fine-tuning (SFT) on increasingly large synthetic datasets.…

cs.IR2026

FollowTable: A Benchmark for Instruction-Following Table Retrieval

Rihui Jin, Yuchen Lu, Ting Zhang +7

Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic similarity. With the growing adop…

cs.CL2026

SRR-Judge: Step-Level Rating and Refinement for Enhancing Search-Integrated Reasoning in Search Agents

Chen Zhang, Kuicai Dong, Dexun Li +4

Recent deep search agents built on large reasoning models (LRMs) excel at complex question answering by iteratively planning, acting, and gathering evidence, a capability known as…