25 citations · 36 across the 19 of their papers we have counts for
4 papers · 1 filter
MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
Hui Shen, Xin Wang, Ping Zhang +11
Vision-language models (VLMs) achieve strong performance on multimodal tasks but suffer from high inference latency due to large model sizes and long multimodal contexts. Speculati…
MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents
Peizhou Huang, Zixuan Zhong, Zhongwei Wan +12
Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA…
Famba-V: Fast Vision Mamba with Cross-Layer Token Fusion
Hui Shen, Zhongwei Wan, Xin Wang +1
Mamba and Vision Mamba (Vim) models have shown their potential as an alternative to methods based on Transformer architecture. This work introduces Fast Mamba for Vision (Famba-V),…
Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation
Che Liu, Zhongwei Wan, Yuqi Wang +5
Automatic radiology report generation holds significant potential to streamline the labor-intensive process of report writing by radiologists, particularly for 3D radiographs such…