6 papers
MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation
Zichun Wang, Hairong Shi, Bingzheng Wei +2
Volumetric Reasoning Segmentation (VRS) aims to segment a target region in a 3D medical scan from a free-form clinical query, where the referent is often implicit and requires both…
Open-Loop Planning, Closed-Loop Verification: Speculative Verification for VLA
Zihua Wang, Zhitao Lin, Ruibo Li +4
Vision-Language-Action (VLA) models, as large foundation models for embodied control, have shown strong performance in manipulation tasks. However, their performance comes at high…
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
Haiyang Xu, Xi Zhang, Haowei Liu +16
The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms…
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
Zihua Wang, Ruibo Li, Haozhe Du +3
Large language models and large multimodal models (LLMs and LMMs) deliver strong generative performance but suffer from slow decoding, a problem that becomes more severe when handl…
Efficient and Effective In-context Demonstration Selection with Coreset
Zihua Wang, Jiarui Wang, Haiyang Xu +6
In-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. Howeve…
Adaptively Clustering Neighbor Elements for Image-Text Generation
Zihua Wang, Xu Yang, Hanwang Zhang +4
We propose a novel Transformer-based image-to-text generation model termed as \textbf{ACF} that adaptively clusters vision patches into object regions and language words into phras…