works on

From the 1 of 19 linked papers with an AI index.

collaborators

19 papers

cs.CV2026

MinerU.Chem: A High-Precision System for Optical Chemical Structure and Reaction Recognition

Haote Yang, Jiang Wu, Jingchao Wang +42

In organic chemistry papers and patents, molecular structures, reaction schemes, and experimental conditions are often presented as molecular structure depictions, reaction diagram…

cs.CL2026

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Rubin Wei, Jiaqi Cao, Jiarui Wang +4

The paper presents Memory Decoder at Scale, a pretrained parametric long‑term memory module for decoder‑only language models that is scaled up to 6.9 B parameters and shown to impr…

cs.LG2026

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

Jiarui Wang, Xiang Shi, Jiaqi Cao +8

Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantia…

cs.CL2026

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny

Chuanhao Yan, Fengdi Che, Xuhan Huang +12

Existing informal language-based (e.g., human language) Large Language Models (LLMs) trained with Reinforcement Learning (RL) face a significant challenge: their verification proce…

cs.CL2026

InternAgentHarness: A Scalable Synthetic Environment for Enhancing LLM Agentic Abilities

Peiji Li, Jiasheng Ye, Yongkang Chen +19

Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however, requires stable an…

cs.CV2026

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

Bin Wang, Tianyao He, Linke Ouyang +40

Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remains underexplored. Yet state-of-the-art…