3 papers
cs.AI2026
StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
Wenyi Wu, Sibo Zhu, Kun Zhou +3
Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled increasingly capable digital agents for computer use. However, real-world tasks are o…
cs.CV2025
GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
Zixuan Song, Jing Zhang, Di Wang +5
Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradi…
cs.LG2025
Towards General Continuous Memory for Vision-Language Models
Wenyi Wu, Zixuan Song, Kun Zhou +3
Language models (LMs) and their extension, vision-language models (VLMs), have achieved remarkable performance across various tasks. However, they still struggle with complex reaso…