25 citations · 41 across the 19 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents
Peizhou Huang, Zixuan Zhong, Zhongwei Wan +12
Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA…
cs.CV2024★ 3 cited
Autoregressive Models in Vision: A Survey
Jing Xiong, Gongye Liu, Lun Huang +17
Autoregressive modeling has been a huge success in the field of natural language processing (NLP). Recently, autoregressive models have emerged as a significant area of focus in co…
cs.CV2024
Famba-V: Fast Vision Mamba with Cross-Layer Token Fusion
Hui Shen, Zhongwei Wan, Xin Wang +1
Mamba and Vision Mamba (Vim) models have shown their potential as an alternative to methods based on Transformer architecture. This work introduces Fast Mamba for Vision (Famba-V),…