2 papers
cs.CV2026
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Yida Yin, Harish Krishnakumar, Chung Peng Lee +9
In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual…
cs.LG2025
Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
Zhengyu Tian, Anantha Padmanaban Krishna Kumar, Hemant Krishnakumar +1
As large language models (LLMs) and visual language models (VLMs) grow in scale and application, attention mechanisms have become a central computational bottleneck due to their hi…