5 papers
Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents
Baichuan Li, Junyi Yao, Zihao Zheng
Persistent memory can personalize an LLM agent, but an incorrect durable update can silently distort future behavior. We study the memory-clarification boundary: whether interactio…
Reliable Financial Named Entity Recognition under Domain Shift
Zihao Zheng, Baichuan Li, Junyi Yao +1
Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, while standard F1 scores do not in…
Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
Junyi Yao, Zihao Zheng, Baichuan Li
Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support. Motivated by recent c…
Ranking Abuse via Strategic Pairwise Data Perturbations
Junyi Yao, Zihao Zheng, Jiayu Long
Pairwise ranking systems based on Maximum Likelihood Estimation (MLE), such as the Bradley-Terry model, are widely used to aggregate preferences from pairwise comparisons. However,…
Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems
Junyi Yao, Zihao Zheng
Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary i…