From the 2 of 8 linked papers with an AI index.
8 papers
Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
Haotian Mo, Jie Liu, Siqi Shen +8
The paper introduces a cross-domain audio deepfake detection method that uses a frozen Diffusion Transformer trained on real speech to generate reconstruction residuals at multiple…
HistoGPA: A Context-Conditioned Gene-Prior Attention Framework for Histology-Based Spatial Gene Expression Prediction
Ziang Liu, Xinhai Chen, Yigui Feng +3
Predicting spatial gene expression from routine hematoxylin and eosin (H&E) images provides a practical complement to experimental spatial transcriptomics. Existing approaches focu…
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
Yigui Feng, Qinglin Wang, Yang Liu +1
The paper introduces Fre-Res, a dual‑track video token compression method for video multimodal large language models that keeps a few high‑fidelity spatial anchor tokens while enco…
FedACT: Federated Adaptive Coordinate Trust Modulation for Robust Transformer Training under Data Heterogeneity
Shuai Li, Qinglin Wang, Ping Luo +8
Federated Transformer training increasingly relies on local AdamW, whose adaptive updates can provide much stronger local progress than SGD-based training. However, under heterogen…
LLM4Fluid: Large Language Models as Generalizable Neural Solvers for Fluid Dynamics
Qisong Xiao, Xinhai Chen, Qinglin Wang +10
Deep learning has emerged as a promising paradigm for spatio-temporal modeling of fluid dynamics. However, existing approaches often suffer from limited generalization to unseen fl…
Measuring the Unspoken: A Disentanglement Model and Benchmark for Psychological Analysis in the Wild
Yigui Feng, Qinglin Wang, Haotian Mo +7
Generative psychological analysis of in-the-wild conversations faces two fundamental challenges: (1) existing Vision-Language Models (VLMs) fail to resolve Articulatory-Affective A…