11 papers
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
Chaoqun He, Mingyang Xiang, Yingjing Xu +5
Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming inputs and respond at approp…
DisDop: Distillation with Domain Priors for Open-Vocabulary Aerial Object Detection
Ruihao Xu, Yong Liu, Yansong Tang +6
With the widespread application of drones in recent years, object detection of aerial images has attracted increasing attention, especially open-vocabulary aerial detection which i…
CoDA: Color Distribution Probing for Efficient and Generalizable AI-Generated Image Detection
Zexi Jia, Zhiqiang Yuan, Xiaoyue Duan +3
AI-generated image detection faces a persistent trade-off between generalization and efficiency: lightweight artifact-based methods often degrade on unseen generators or domains, w…
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
Wenxuan Guo, Xiuwei Xu, Yichen Liu +7
Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the…
Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
Deyi Zhu, Yuji Wang, Yong Liu +4
Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects and challenging scenarios with…
Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor
Yang Zhao, Jiahao Lu, Bin Huang +2
Narang et al. (2021) evaluated 40+ Transformer modifications at T5-base scale and concluded that most did not transfer. Five years later, the typical working regime has moved to 1-…