3 papers
cs.CL2026
Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents
Sina Hajimiri, Masih Aminbeidokhti, Jose Dolz +4
Online web agents often augment a base actor with memory, workflow, or skill modules. These modules can improve performance, but they also consume test-time tokens, a cost rarely r…
cs.CV2026
ORION: ORthonormal Text Encoding for Universal VLM AdaptatION
Omprakash Chakraborty, Jose Dolz, Ismail Ben Ayed
Vision language models (VLMs) have demonstrated remarkable generalization across diverse tasks, yet their performance remains constrained by the quality and geometry of the textual…
cs.CV2026
Histopath-C: Towards Realistic Domain Shifts for Histopathology Vision-Language Adaptation
Mehrdad Noori, Gustavo Adolfo Vargas Hakim, David Osowiechi +6
Medical Vision-language models (VLMs) have shown remarkable performances in various medical imaging domains such as histo\-pathology by leveraging pre-trained, contrastive models t…