2 papers
cs.CL2026
To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection
Erfan Loweimi, Mengjie Qian, Kate Knill +7
When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may b…
cs.CV2026
Paving the Way for Point Cloud Video Representation Learning Using A PDE Model
Zhuoxu Huang, Zhenkun Fan, Jungong Han +1
Investigating spatial-temporal correlations, specifically how spatial points vary over time, is crucial for understanding point cloud videos. Traditional methods, particularly flow…