collaborators

10 papers

cs.CL2026

Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents

Ying He, Zhouhong Gu, Zhecheng Hu +8

Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language…

cs.LG2026

ADaPT: Token-Level Decoupling for Efficient Large Reasoning Models

Tingyun Li, Zishang Jiang, Jinyi Han +8

Large reasoning models rely on long chain-of-thought to achieve strong performance, but applying such reasoning uniformly incurs high computational cost. Existing efficiency-orient…

cs.AI2026

LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models

Qingyu Ren, Qianyu He, Jingwen Chang +9

Instruction following is critical for large language models, yet real-world instructions often involve multiple constraints with logical structures, such as parallel composition, s…

cs.AI2026

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

Sihang Jiang, Lipeng Ma, Zhonghua Hong +9

Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience…

cs.CL2026

GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)

Jiaqing Liang, Jinyi Han, Weijia Li +15

Long-horizon large language model (LLM) agents are fundamentally limited by context. As interactions become longer, tool descriptions, retrieved memories, and raw environmental fee…

cs.LG2026

ChemAmp: Amplified Chemistry Tools via Composable Agents

Zhucong Li, Powei Chang, Jin Xiao +6

Although LLM-based agents are proven to master tool orchestration in scientific fields, particularly chemistry, their single-task performance remains limited by underlying tool con…