activity
20242026
most citedMulti-Task Model Merging via Adaptive Weight Disentanglement

1 citations · 1 across the 8 of their papers we have counts for

collaborators

10 papers

cs.AI2026

GameWAM: A World Action Model for Video Games

Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo +1

Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task contex…

cs.LG2026

Progressive Agent Skill Generation via Reinforcement Learning

Junhao Shen, Zhanqiu Zhang, Yiwen Guo +1

Recent large language model agents often use external skills as modular procedural units that condition inference and improve complex task solving. Thus, automatically generating h…

cs.CL2026

CONDUIT: A Unified Residual-Stream Restoration Framework for KV Cache Reuse in Vision-Language Models

Pengan Chen, Kaisheng Zheng, Liang Hong +10

Vision-language models (VLMs) often answer new questions about recurring visual content, where reusing the key-value (KV) cache can avoid re-encoding expensive visual prefixes. Exa…

cs.LG2026

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List

Zhanqi Zhang, Hua-Dong Xiong, Robert C. Wilson +3

Modern large language models (LLMs) can find a needle in a haystack (locating a single relevant fact buried among hundreds of thousands of irrelevant tokens) with near-saturated ac…

cs.AI2026

Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction

Xingwu Chen, Zhanqiu Zhang, Yiwen Guo +1

While LLMs demonstrate strong reasoning capabilities when provided with full information in a single turn, they exhibit substantial vulnerability in multi-turn interactions. Specif…

cs.RO2025

See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations

Guangyan Chen, Meiling Wang, Qi Shao +10

Developing robust and general-purpose manipulation policies represents a fundamental objective in robotics research. While Vision-Language-Action (VLA) models have demonstrated pro…