collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI2026

Improving Generalization Robustness of Multimodal RLVR

Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng +11

Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing th…

cs.AI2026

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

Tong Bai, Zhenglin Wan, Pengfei Zhou +3

As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend on, conflict with, specializ…

cs.AI2026

Agent-as-a-Router: Agentic Model Routing for Coding Tasks

Pengfei Zhou, Zhiwei Tang, Yixing Ma +8

Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at distinct domains, yet none dominate all. Con…

cs.AI2026

Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents

Chubin Zhang, Zhenglin Wan, Xingrui Yu +5

Tool-augmented agents are typically evaluated by their gains under reliable external feedback. Yet these gains leave open a key counterfactual: when feedback is unreliable, would t…

cs.AI2026

Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention

Chubin Zhang, Zhenglin Wan, Xingrui Yu +5

Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a thre…

cs.AI2026

Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

Zhijie Ding, Weinan Hong, Zicheng Zhu +6

Multimodal large language models (MLLMs) have substantially advanced mobile agents, yet proactive mobile assistance remains challenging because agents must decide \emph{when} to in…