From the 1 of 4 linked papers with an AI index.
4 papers
Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
Haotian Mo, Jie Liu, Siqi Shen +8
The paper introduces a cross-domain audio deepfake detection method that uses a frozen Diffusion Transformer trained on real speech to generate reconstruction residuals at multiple…
MAGE: Multi-scale Autoregressive Generation for Offline Reinforcement Learning
Chenxing Lin, Xinhui Gao, Haipeng Zhang +7
Generative models have gained significant traction in offline reinforcement learning (RL) due to their ability to model complex trajectory distributions. However, existing generati…
Measuring the Unspoken: A Disentanglement Model and Benchmark for Psychological Analysis in the Wild
Yigui Feng, Qinglin Wang, Haotian Mo +7
Generative psychological analysis of in-the-wild conversations faces two fundamental challenges: (1) existing Vision-Language Models (VLMs) fail to resolve Articulatory-Affective A…
PlanU: Large Language Model Reasoning through Planning under Uncertainty
Ziwei Deng, Mian Deng, Chenjing Liang +7
Large Language Models (LLMs) are increasingly being explored across a range of reasoning tasks. However, LLMs sometimes struggle with reasoning tasks under uncertainty that are rel…