12 papers
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?
Jingheng Ye, Huiqi Zou, Simon Yu +1
AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates…
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
Yibo Yan, Shen Wang, Jiahao Huo +7
Scientific reasoning, the process through which humans apply logic, evidence, and critical thinking to explore and interpret scientific phenomena, is essential in advancing knowled…
GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment
Jiwei Tang, Zhicheng Zhang, Shunlong Wu +8
Large Language Models (LLMs) have achieved remarkable performance across a wide range of Natural Language Processing (NLP) tasks. However, in long-context scenarios, they face two…
FEANEL: A Benchmark for Fine-Grained Error Analysis in K-12 English Writing
Jingheng Ye, Shen Wang, Jiaqi Chen +9
Large Language Models (LLMs) have transformed artificial intelligence, offering profound opportunities for educational applications. However, their ability to provide fine-grained…
CLGEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction
Shang Qin, Jingheng Ye, Yinghui Li +5
The growing demand for automated writing assistance in diverse academic domains highlights the need for robust Chinese Grammatical Error Correction (CGEC) systems that can adapt ac…
Position: LLMs Can be Good Tutors in English Education
Jingheng Ye, Shen Wang, Deqing Zou +8
While recent efforts have begun integrating large language models (LLMs) into English education, they often rely on traditional approaches to learning tasks without fully embracing…