collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Investigating Advanced Reasoning of Large Language Models via Black-Box Environment Interaction

Congchi Yin, Tianyi Wu, Yankai Shu +5

Existing tasks fall short in evaluating reasoning ability of Large Language Models (LLMs) in an interactive, unknown environment. This deficiency leads to the isolated assessment o…

cs.AI2026

ESL-Bench: An Event-Driven Synthetic Longitudinal Benchmark for Health Agents

Chao Li, Cailiang Liu, Ang Gao +7

Longitudinal health agents must reason across multi-source trajectories that combine continuous device streams, sparse clinical exams, and episodic life events - yet evaluating the…

cs.AI2025

On Path to Multimodal Historical Reasoning: HistBench and HistAgent

Jiahao Qiu, Fulian Xiao, Yimin Wang +96

Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored…

cs.AI2025

AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes

Jiahao Qiu, Xinzhe Juan, Yimin Wang +11

While knowledge distillation has become a mature field for compressing large language models (LLMs) into smaller ones by aligning their outputs or internal representations, the dis…

cs.AI2025

Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution

Jiahao Qiu, Xuan Qi, Tongcheng Zhang +15

Recent advances in large language models (LLMs) have enabled agents to autonomously perform complex, open-ended tasks. However, many existing frameworks depend heavily on manually…

cs.AI2025

EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety

Jiahao Qiu, Yinghui He, Xinzhe Juan +7

The rise of LLM-driven AI characters raises safety concerns, particularly for vulnerable human users with psychological disorders. To address these risks, we propose EmoAgent, a mu…