2 papers
cs.LG2026
Zoom to Essence: Trainless GUI Grounding by Inferring upon Interface Elements
Ziwei Liu, Tao Feng, Borui Kang +2
Multimodal Large Language Model (MLLM)-based Graphical User Interface (GUI) agents develop rapidly, with visual grounding that maps natural language instructions to target UI eleme…
cs.CV2026
Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
Ziwei Liu, Borui Kang, Wei Li +6
Vision-Language Continual Learning (VLCL) has attracted significant research attention for its robust capabilities, and the adoption of Parameter-Efficient Fine-Tuning (PEFT) strat…