3 papers
cs.CV2025
Med-LEGO: Editing and Adapting toward Generalist Medical Image Diagnosis
Yitao Zhu, Yuan Yin, Jiaming Li +5
The adoption of visual foundation models has become a common practice in computer-aided diagnosis (CAD). While these foundation models provide a viable solution for creating genera…
cs.CV2025
MITracker: Multi-View Integration for Visual Object Tracking
Mengjie Xu, Yitao Zhu, Haotian Jiang +8
Multi-view object tracking (MVOT) offers promising solutions to challenges such as occlusion and target loss, which are common in traditional single-view tracking. However, progres…
eess.AS2024
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
Xinyu Wang, Haotian Jiang, Haolin Huang +3
Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for…