4 papers
When History Is Multimodal: Rethinking Context Management for Long-Horizon Agents
Jiaqi Su, Cong Pang, Jiawei Hong +4
Long-horizon agents need a context manager to compress growing interaction histories into a bounded working context, via passive strategies or active strategies that decide how mem…
ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories
Siyuan Luo, Nairong Zheng, Lin Zhou +6
Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded tool execution--properties absent from exist…
ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents
Cong Pang, Xuyu Feng, Yujie Yi +8
Long-horizon reinforcement learning for information seeking agents remains difficult because terminal rewards reveal whether the final answer is correct, but not which acquired inf…
Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
Cong Pang, Hongtao Yu, Zixuan Chen +2
Large Vision Language Models (LVLMs) have made remarkable progress, enabling sophisticated vision-language interaction and dialogue applications. However, existing benchmarks prima…