From the 1 of 9 linked papers with an AI index.
9 papers
Qwen-Audio-3.0-Gen-Preview Technical Report
Junyu Dai, Xiaoyue Duan, Xinyue Fan +14
The paper introduces Qwen-Audio-3.0-Gen-Preview, a unified non‑autoregressive model that uses a diffusion transformer and a shared VAE to generate complete mixed‑waveform audio fro…
Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues
Hanbo Bi, Zhiqiang Yuan, Chongyang Li +7
With the widespread adoption of multi-modal communication platforms, long-form dialogues interleaving text and images have become increasingly common. Users often need to retrieve…
SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling
Xiaoyue Duan, Nanxing Hu, Yutang Feng +4
Recent song generation systems can synthesize realistic audio, yet generating complete songs remains challenging for two reasons. First, explicit song-level arrangement planning re…
CoDA: Color Distribution Probing for Efficient and Generalizable AI-Generated Image Detection
Zexi Jia, Zhiqiang Yuan, Xiaoyue Duan +3
AI-generated image detection faces a persistent trade-off between generalization and efficiency: lightweight artifact-based methods often degrade on unseen generators or domains, w…
Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
Chongyang Li, Zhiqiang Yuan, Hanbo Bi +2
Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs) to develop effective walking a…
Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
Weihao Cao, Runqi Wang, Xiaoyue Duan +3
Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-training, existing OVOD methods achieve…