12 papers · 1 filter
Disentangling Semantic Attention from Structural Bias in the Attention Manifold
Pengkun Jiao, Bin Zhu, Jingjing Chen +1
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disprop…
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
Ziyao Tang, Pengkun Jiao, Bin Zhu +3
Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely…
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
Yian Li, Yang Jiao, Bin Zhu +4
Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language…
OSCBench: Benchmarking Object State Change in Text-to-Video Generation
Xianjing Han, Bin Zhu, Shiqi Hu +4
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on pe…
RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation
Shijie Zhou, Bin Zhu, Jiarui Yang +3
Recent advances in Vision-Language-Action (VLA) models have enabled robots to execute increasingly complex tasks. However, VLA models trained through imitation learning struggle to…
Benchmarking Gaslighting Negation Attacks Against Reasoning Models
Bin Zhu, Hailong Yin, Jingjing Chen +1
Recent advances in reasoning-centric models promise improved robustness through mechanisms such as chain-of-thought prompting and test-time scaling. However, their ability to withs…