2 papers
cs.CV2025
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
Xinyu Wei, Guoli Yang, Jialu Zhou +4
Large Vision-Language Models (LVLMs) commonly follow a paradigm that projects visual features and then concatenates them with text tokens to form a unified sequence input for Large…
cs.HC2025
MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
Yi Kong, Dianxi Shi, Guoli Yang +4
The recent advancement of autonomous agents powered by Large Language Models (LLMs) has demonstrated significant potential for automating tasks on mobile devices through graphical…