3 papers
cs.CV2026
Deformba: Vision State Space Model with Adaptive State Fusion
Hongyu Ke, Jack Morris, Yongkang Liu +4
State Space Models (SSMs) have emerged as a powerful and efficient alternative to Transformers, demonstrating linear-time complexity and exceptional sequence modeling capabilities.…
cs.CL2026
Can a Unimodal Language Agent Provide Preferences to Tune a Multimodal Vision-Language Model?
Sazia Tabasum Mim, Jack Morris, Manish Dhakal +3
To explore a more scalable path for adding multimodal capabilities to existing LLMs, this paper addresses a fundamental question: Can a unimodal LLM, relying solely on text, reason…
cs.CV2025
MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations
Hongyu Ke, Jack Morris, Kentaro Oguchi +4
3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally…