4 papers
Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
Tianyidan Xie, Shenyi Wang, Qiang Tang +7
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to…
EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory
Weitao Chen, Hu Jiaxin, Xie Tianyidan +15
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. Howev…
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
Xinyu Chen, Yuyi Qian, Jiang Lin +9
Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hin…
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
Dingbao Shao, Song Wu, Shenyi Wang +9
Due to the scarcity of large-scale in-the-wild triplet data and the improper use of masks, the performance of video virtual try-on models remains limited. In this paper, we first i…