activity
20242026
collaborators

7 papers

cs.AI2026

Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents

Xiaolong Sun, Qichao Wang, Hangyu Li +1

Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon interaction. Existing methods…

cs.CV2025

UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation

Zeyang Liu, Le Wang, Sanping Zhou +4

Content-aware layout generation is a critical task in graphic design automation, focused on creating visually appealing arrangements of elements that seamlessly blend with a given…

cs.CV2025

Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space

Jian Zhu, Zhengyu Jia, Tian Gao +6

Advanced end-to-end autonomous driving systems predict other vehicles' motions and plan ego vehicle's trajectory. The world model that can foresee the outcome of the trajectory has…

cs.CV2025

Length Matters: Length-Aware Transformer for Temporal Sentence Grounding

Yifan Wang, Ziyi Liu, Xiaolong Sun +2

Temporal sentence grounding (TSG) is a highly challenging task aiming to localize the temporal segment within an untrimmed video corresponding to a given natural language descripti…

cs.CL2025

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

Yuchang Zhu, Huazhen Zhong, Qunshu Lin +6

With the remarkable generative capabilities of large language models (LLMs), using LLM-generated data to train downstream models has emerged as a promising approach to mitigate dat…

cs.CV2025

Moment Quantization for Video Temporal Grounding

Xiaolong Sun, Le Wang, Sanping Zhou +5

Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishin…