7 papers
Test-Time Computing for Referring Multimodal Large Language Models
Mingrui Wu, Hao Chen, Jiayi Ji +5
We propose ControlMLLM++, a novel test-time adaptation framework that injects learnable visual prompts into frozen multimodal large language models (MLLMs) to enable fine-grained r…
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
Mingrui Wu, Hang Liu, Jiayi Ji +2
Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini-2.5-Flash-Image s…
From Atoms to Trees: Building a Structured Feature Forest with Hierarchical Sparse Autoencoders
Yifan Luo, Yang Zhan, Jiedong Jiang +4
Sparse autoencoders (SAEs) have proven effective for extracting monosemantic features from large language models (LLMs), yet these features are typically identified in isolation. H…
Vision Calorimeter for Anti-neutron Reconstruction: A Baseline
Hongtian Yu, Yangu Li, Mingrui Wu +8
In high-energy physics, anti-neutrons () are fundamental particles that frequently appear as final-state particles, and the reconstruction of their kinematic properties pr…
Do Images Speak Louder than Words? Investigating the Effect of Textual Misinformation in VLMs
Chi Zhang, Wenxuan Ding, Jiale Liu +3
Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities on Visual-Question-Answering (VQA) benchmarks. However, their robustness against textual misinform…
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
Mingrui Wu, Zhaozhi Wang, Fangjinhua Wang +3
While Multimodal Large Language Models (MLLMs) have achieved impressive performance on semantic tasks, their spatial intelligence--crucial for robust and grounded AI systems--remai…