From the 2 of 4 linked papers with an AI index.
4 papers
VITAL-RAG: Invariance Race for Context Allocation in Coding Agents
Zijian Lu, Yonghua Lu, Mingcai Chen +4
The paper introduces VITAL-RAG, a method for coding agents that groups retrieved code fragments by their original code object and selectively includes only those that add new task-…
Prior Directions: Why GUI Grounding Gets Locked in the Past
Weile Gong, Zijian Lu, Mingcai Chen +3
The paper investigates how vision-language models can become locked onto outdated textual priors, causing incorrect visual grounding, and identifies recurring latent directions—cal…
ContractSkill: Repairable Contract-Based Skills for Multimodal Web Agents
Zijian Lu, Yiping Zuo, Yupeng Nie +4
Self-generated skills for web agents are often unstable and can even hurt performance relative to direct acting. We argue that the key bottleneck is not only skill generation quali…
Geometric Risk Control for Vision-Language Model OCR
Weile Gong, Yiping Zuo, Mingcai Chen +5
Vision-language models (VLMs) enable flexible generative optical character recognition (OCR), while their open-ended decoders can expose wrong but fluent text with weak visual supp…