3 papers
cs.CV2026
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
Daichi Yashima, Shuhei Kurita, Yusuke Oda +3
In this study, we focus on video captioning by fully open multimodal large language models (MLLMs). The comprehension of visual sequences is challenging because of their intricate…
q-bio.NC2025
MEGState: Phoneme Decoding from Magnetoencephalography Signals
Shuntaro Suzuki, Chia-Chun Dan Hsu, Yu Tsao +1
Decoding linguistically meaningful representations from non-invasive neural recordings remains a central challenge in neural speech decoding. Among available neuroimaging modalitie…
q-bio.NC2025
Condition-Invariant fMRI Decoding of Speech Intelligibility with Deep State Space Model
Ching-Chih Sung, Shuntaro Suzuki, Francis Pingfan Chien +2
Clarifying the neural basis of speech intelligibility is critical for computational neuroscience and digital speech processing. Recent neuroimaging studies have shown that intellig…