2 papers
cs.CV2026
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
Kaixin zhang, Xiaohe Li, Jiahao Li +4
Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal cau…
cs.LG2025
SAFE: Self-Adjustment Federated Learning Framework for Remote Sensing Collaborative Perception
Xiaohe Li, Haohua Wu, Jiahao Li +5
The rapid increase in remote sensing satellites has led to the emergence of distributed space-based observation systems. However, existing distributed remote sensing models often r…