3 papers
cs.CV2026
SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search
Ming Dai, Zhihong Lu, Jinjie Gu +5
We present SimpleSearch-VL, an efficient, reliable, and practical framework for multimodal agentic search. Its core idea is to improve the agent's own search-and-verification proce…
cs.SD2026
DisSR: Disentangling Speech Representation for Degradation-Prior Guided Cross-Domain Speech Restoration
Ziqi Liang, Zhijun Jia, Chang Liu +3
Previous speech restoration (SR) primarily focuses on single-task speech restoration (SSR), which cannot address general speech restoration problems. Training specific SSR models f…
cs.CV2026
PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
Jiao Xu, Junwei Liu, Jiangwei Lao +9
Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of re…