1 citations · 1 across the 3 of their papers we have counts for
3 papers
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
Haoyu Zhao, Weizhong Ding, Yuhao Yang +4
Recent advances in Multimodal Large Language Models (MLLMs) have enabled their use as intelligent agents for smartphone operation. However, existing methods depend on the Android D…
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
Siting Wang, Minnan Pei, Luoyang Sun +6
Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual informat…
Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
Zhengyu Tian, Anantha Padmanaban Krishna Kumar, Hemant Krishnakumar +1
As large language models (LLMs) and visual language models (VLMs) grow in scale and application, attention mechanisms have become a central computational bottleneck due to their hi…