From the 1 of 17 linked papers with an AI index.
17 papers
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Mingqiao Ye, Zhaochong An, Zhitong Gao +11
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly…
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
Kaixin Ma, Di Feng, Alexander Metz +3
The paper introduces MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents that handle multi-image, multi-turn tasks across hundreds of too…
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
OÄuzhan Fatih Kar, Roman Bachmann, Yuanzheng Gong +2
The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to off…
(1D) Ordered Tokens Enable Efficient Test-Time Search
Zhitong Gao, Parham Rezaei, Ali Cy +7
Tokenization is a key component of autoregressive (AR) generative models, converting raw data into more manageable units for modeling. Commonly, tokens describe local information,…
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
Andrei Atanov, Jesse Allardice, Roman Bachmann +6
Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what information is preserved and…
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
Di Feng, Kaixin Ma, Feng Nan +9
Multimodal large language models (MLLMs) are increasingly deployed in real-world, agentic settings where outputs must not only be correct, but also conform to predefined data schem…