3 papers
cs.CV2026
Dive Into the Implicit Biases of Low-rank Vision-language Alignment
Mingjia Shi, Shuo Wang, Xiaobo Wang +7
Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates.…
cs.CV2026
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
Wei Dong, Tianyu Fu, Zhe Yu +9
As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion…
cs.MM2025
Mano Technical Report
Tianyu Fu, Anyang Su, Chenxu Zhao +20
Graphical user interfaces (GUIs) are the primary medium for human-computer interaction, yet automating GUI interactions remains challenging due to the complexity of visual elements…