1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.SD2026★ 1 cited
Covo-Audio Technical Report
Wenfu Wang, Chenxing Li, Liqiang Zhang +23
In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture…
cs.CV2025
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
Chengfei Wu, Ronald Seoh, Bingxuan Li +3
Recent advances in large vision-language models have led to impressive performance in visual question answering and multimodal reasoning. However, it remains unclear whether these…