2 papers
cs.IR2026
NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference
Aurélien Lac, Tony Wu
Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language mode…
cs.AI2025
Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
Mathieu Andreux, Märt Bakler, Yanael Barbier +50
Building agents that generalize across web, desktop, and mobile environments remains an open challenge, as prior systems rely on environment-specific interfaces that limit cross-pl…