11 papers
StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
Wenyi Wu, Sibo Zhu, Kun Zhou +3
Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled increasingly capable digital agents for computer use. However, real-world tasks are o…
C-World: A Computer Use Agent Environment Creator
Ziqiao Xi, Shuang Liang, Qi Liu +9
To close the gap between LLM-based agents and humans in planning and reasoning, agents need large-scale, diverse environments for continuous learning -- yet building such environme…
Hybrid Self-evolving Structured Memory for GUI Agents
Sibo Zhu, Wenyi Wu, Kun Zhou +2
The remarkable progress of vision-language models (VLMs) has enabled GUI agents to interact with computers in a human-like manner. Yet real-world computer-use tasks remain difficul…
Causal Structure Learning in Hawkes Processes with Complex Latent Confounder Networks
Songyao Jin, Biwei Huang
Multivariate Hawkes process provides a powerful framework for modeling temporal dependencies and event-driven interactions in complex systems. While existing methods primarily focu…
Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings
Yifei Shao, Kun Zhou, Ziming Xu +5
We study how to extend chain-of-thought (CoT) beyond language to better handle multimodal reasoning. While CoT helps LLMs and VLMs articulate intermediate steps, its text-only form…
Ability Transfer and Recovery via Modularized Parameters Localization
Songyao Jin, Kun Zhou, Wenqi Li +2
Large language models can be continually pre-trained or fine-tuned to improve performance in specific domains, languages, or skills, but this specialization often degrades other ca…