14 papers
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
Yuetian Du, Yucheng Wang, Zhenyuan Chen +9
Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these mo…
HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM Inference
Hui Dong, Yanzhao Li, Jie Gao +5
We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on Ascend NPUs, while maintainin…
Training-free sparse attention based on cumulative energy filtering
Chunlu Li, Yixuan Pan, Bai Du +5
Sparse attention accelerates Diffusion Transformers (DiTs) for video generation by computing only the important tokens while skipping the rest. The token selection strategy is key…
RSEdit: Text-Guided Image Editing for Remote Sensing
Chen Zhenyuan, Zhang Zechuan, Zhang Feng
In this paper, we explore text-guided image editing in the remote sensing domain using generative modeling. We propose \rsedit, a collection of models from U-Net to DiT with variou…
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
Bo Gu, Zhikang Zhang, Zizhuang Wei +3
Recent multimodal large language models (MLLMs) have made remarkable progress in visual understanding and language-based reasoning, yet they lack a persistent world-centered repres…
EarthBridge: A Solution for 4th Multi-modal Aerial View Image Challenge Translation Track
Zhenyuan Chen, Guanyuan Shen, Feng Zhang
Cross-modal image-to-image translation among Electro-Optical (EO), Infrared (IR), and Synthetic Aperture Radar (SAR) sensors is essential for comprehensive multi-modal aerial-view…