13 papers
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
Kaixin Ma, Di Feng, Alexander Metz +3
The paper introduces MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents that handle multi-image, multi-turn tasks across hundreds of too…
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
Di Feng, Kaixin Ma, Feng Nan +9
Multimodal large language models (MLLMs) are increasingly deployed in real-world, agentic settings where outputs must not only be correct, but also conform to predefined data schem…
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
Ce Zhang, Kaixin Ma, Tianqing Fang +5
Recent Large Vision-Language Models (LVLMs) have advanced multi-modal understanding by incorporating finer-grained visual perception and encoding. However, such methods incur signi…
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
Zhisong Zhang, Tianqing Fang, Kaixin Ma +4
With recent advancements in large language models, web agents have been greatly improved. However, dealing with complex and dynamic web environments requires more advanced planning…
VRoPE: Rotary Position Embedding for Video Large Language Models
Zikang Liu, Longteng Guo, Yepeng Tang +6
Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiot…
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
Hyunji Lee, Wenhao Yu, Hongming Zhang +4
Hybrid models that combine state space models (SSMs) with attention mechanisms have shown strong performance by leveraging the efficiency of SSMs and the high recall ability of att…