works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CL2026

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

Ayoung Lee, Ryan Kwon, Yunxiang Zhang +3

The paper introduces a multilingual, culture-aware benchmark (MCLASH) and a two-step, theory‑grounded prompting method (MET) with a self‑distillation variant (MET‑D) to improve mor…

cs.AI2026

LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?

Kaijian Zou, Aaron Xiong, Yunxiang Zhang +6

Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, cu…

cs.AI2026

AdaMEM: Test-Time Adaptive Memory for Language Agents

Yunxiang Zhang, Yiheng Li, Ali Payani +1

A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions. While recent work demonstrates the promise of agentic memory mechanis…

cs.CL2026

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

Yunxiang Zhang, Muhammad Khalifa, Lechen Zhang +5

Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification. Yet, these capabilities typically require resourc…

cs.CL2026

Learning to Ideate for Machine Learning Engineering Agents

Yunxiang Zhang, Kang Zhou, Zhichao Xu +5

Existing machine learning engineering (MLE) agents struggle to iteratively optimize their implemented algorithms for effectiveness. To address this, we introduce MLE-Ideator, a dua…

cs.AI2026

Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation

Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim +6

Large language models (LLMs) are increasingly used as judges to evaluate agent performance, particularly in non-verifiable settings where judgments rely on agent trajectories inclu…