activity
20242026
collaborators

7 papers

cs.CL2026

Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks

Yuqing Yang, Tengxiao Liu, Wang Bill Zhu +3

As LLM-based assistants become persistent and personalized, they must extract and retain useful information from past conversations as memory. However, the types of information wor…

cs.AI2026

WildSci: Advancing Scientific Reasoning from In-the-Wild Literature

Tengxiao Liu, Deepak Nathani, Zekun Li +2

Recent progress in large language model (LLM) reasoning has focused on domains like mathematics and coding, where abundant high-quality data and objective evaluation metrics are re…

cs.AI2025

Budget-Aware Tool Use Enables Effective Agent Scaling

Tengxiao Liu, Zifeng Wang, Jin Miao +12

Scaling test-time computation has been extended from language model reasoning to tool-augmented agents, where scaling involves not only thinking in tokens but also acting via tool…

cs.CL2024

Can Language Models Learn to Skip Steps?

Tengxiao Liu, Qipeng Guo, Xiangkun Hu +4

Trained on vast corpora of human language, language models demonstrate emergent human-like reasoning abilities. Yet they are still far from true intelligence, which opens up intrig…

cs.CL2024

ECon: On the Detection and Resolution of Evidence Conflicts

Cheng Jiayang, Chunkit Chan, Qianqian Zhuang +7

The rise of large language models (LLMs) has significantly influenced the quality of information in decision-making systems, leading to the prevalence of AI-generated content and c…

cs.CL2024

Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation

Qin Zhu, Qingyuan Cheng, Runyu Peng +5

The training process of large language models (LLMs) often involves varying degrees of test data contamination. Although current LLMs are achieving increasingly better performance…