5 papers
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
Zejun Xu, Taiyi Chen, Jin Li +13
Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven by benchmarks that evaluate…
Rethinking Air-Ground Collaboration: A Progressive Cross-Task Benchmark and Socialized Learning Framework
Zhoupeng Guo, Yunqi Zhu, Zhihe Fan +6
Air-ground collaborative perception is crucial for robust visual understanding in real-world dynamic environments. However, existing studies typically formulate collaboration as si…
Improving Partially Observed Trajectories Forecasting by Target-driven Self-Distillation
Peng Shu, Pengfei Zhu, Mengshi Qi +1
Accurate prediction of future trajectories of traffic agents is essential for ensuring safe autonomous driving. However, partially observed trajectories can significantly degrade t…
DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency
Mengshi Qi, Pengfei Zhu, Xiangtai Li +4
Given a single labeled example, in-context segmentation aims to segment corresponding objects. This setting, known as one-shot segmentation in few-shot learning, explores the segme…
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
Mengshi Qi, Xiaoyang Bi, Pengfei Zhu +1
Robustly predicting attention regions of interest for self-driving systems is crucial for driving safety but presents significant challenges due to the labor-intensive nature of ob…