Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
Daichi Yashima, Shuhei Kurita, Yusuke Oda +3
In this study, we focus on video captioning by fully open multimodal large language models (MLLMs). The comprehension of visual sequences is challenging because of their intricate…
cs.CV2025
Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals
Shuntaro Suzuki, Shunya Nagashima, Komei Sugiura
Classification of electroencephalogram (EEG) signals obtained during motor imagery (MI) has substantial application potential, including communication assistance and rehabilitation…