activity
20242026
collaborators

10 papers

cs.SE2026

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement

Xin Wang, Liangtai Sun, Yaoming Zhu +8

Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output. However, real web development is different. Users seldom write a full spec a…

cs.IR2026

LLM-Oriented Information Retrieval: A Denoising-First Perspective

Lu Dai, Liang Sun, Fanpu Cao +4

Modern information retrieval (IR) is no longer consumed primarily by humans but increasingly by large language models (LLMs) via retrieval-augmented generation (RAG) and agentic se…

cs.CL2025

MULTI: Multimodal Understanding Leaderboard with Text and Images

Zichen Zhu, Yang Xu, Lu Chen +11

The rapid development of multimodal large language models (MLLMs) raises the question of how they compare to human performance. While existing datasets often feature synthetic or o…

cs.LG2025

Instance-level Randomization: Toward More Stable LLM Evaluations

Yiyang Li, Yonghuang Wu, Ying Luo +5

Evaluations of large language models (LLMs) suffer from instability, where small changes of random factors such as few-shot examples can lead to drastic fluctuations of scores and…

cs.CL2025

Developing ChemDFM as a large language foundation model for chemistry

Zihan Zhao, Da Ma, Lu Chen +11

Artificial intelligence (AI) has played an increasingly important role in chemical research. However, most models currently used in chemistry are specialist models that require tra…

cs.CL2025

NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question Answering

Ruisheng Cao, Hanchong Zhang, Tiancheng Huang +8

The increasing number of academic papers poses significant challenges for researchers to efficiently acquire key details. While retrieval augmented generation (RAG) shows great pro…