3 papers
cs.CV2026
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
Yunze Liu, Chi-Hao Wu, Enmin Zhou +1
Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) encoders extend this setting…
cs.CV2024
Gesture2Text: A Generalizable Decoder for Word-Gesture Keyboards in XR Through Trajectory Coarse Discretization and Pre-training
Junxiao Shen, Khadija Khaldi, Enmin Zhou +2
Text entry with word-gesture keyboards (WGK) is emerging as a popular method and becoming a key interaction for Extended Reality (XR). However, the diversity of interaction modes,…
cs.CV2024
Towards Open-World Gesture Recognition
Junxiao Shen, Matthias De Lange, Xuhai "Orson" Xu +7
Providing users with accurate gestural interfaces, such as gesture recognition based on wrist-worn devices, is a key challenge in mixed reality. However, static machine learning pr…