activity
20242026
collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

Xianzhong Ding, Yangyang Yu, Changwei Liu +1

A frontier language model's acknowledged "helpful programming assistant" persona does not survive long agentic-coding sessions in the deployment regime that production products act…

cs.CL2026

MERMAID: Memory-Enhanced Retrieval and Reasoning with Multi-Agent Iterative Knowledge Grounding for Veracity Assessment

Yupeng Cao, Chengyang He, Yangyang Yu +2

Assessing the veracity of online content has become increasingly critical. Large language models (LLMs) have recently enabled substantial progress in automated veracity assessment,…

cs.CL2026

Exploring how EFL students talk to and through AI to develop texts

David James Woo, Yangyang Yu, Yilin Huang +3

Generative Artificial Intelligence (AI) introduces new considerations for English as a foreign language (EFL) writing pedagogy. This study explores how students talk to and through…

cs.CL2025

When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents

Lingfei Qian, Xueqing Peng, Yan Wang +14

Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies t…

cs.CL2025

MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application

Xueqing Peng, Lingfei Qian, Yan Wang +44

Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings. Yet most existing eval…

cs.CL2025

Truth Neurons

Haohang Li, Yupeng Cao, Yangyang Yu +2

Despite their remarkable success and deployment across diverse workflows, language models sometimes produce untruthful responses. Our limited understanding of how truthfulness is m…