6 papers
Trip+: Benchmarking Agents in Personalized Interactive Travel Planning
Junle Chen, Wei Chen, Yehong Xu +6
Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Su…
MExam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions
Zhengjun Huang, Wenxuan Liu, Zhoujin Tian +6
Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visuals and straightforward conten…
LifeSide: Benchmarking Agents as Lifelong Digital Companions
Yuqian Wu, Zhijie Deng, Wei Chen +8
Lifelong digital companions must integrate cross-session cues, continually update their understanding of users, and adapt to shifting privacy boundaries. Existing evaluations fail…
UrbanFM: Scaling Urban Spatio-Temporal Foundation Models
Wei Chen, Yuqian Wu, Junle Chen +2
Urban systems, as dynamic complex systems, continuously generate spatio-temporal data streams that encode the fundamental laws of human mobility and city evolution. While AI for Sc…
Learning from Complexity: Exploring Dynamic Sample Pruning of Spatio-Temporal Training
Wei Chen, Junle Chen, Yuqian Wu +2
Spatio-temporal forecasting is fundamental to intelligent systems in transportation, climate science, and urban planning. However, training deep learning models on the massive, oft…
Anchoring Trends: Mitigating Social Media Popularity Prediction Drift via Feature Clustering and Expansion
Chia-Ming Lee, Bo-Cheng Qiu, Cheng-Jun Kang +5
Predicting online video popularity faces a critical challenge: prediction drift, where models trained on historical data rapidly degrade due to evolving viral trends and user behav…