4 papers · 1 filter
LoTA-N2N: Local Trace Adaptation for Zero-Shot Self-Supervised Image Denoising
Jintong Hu, Bin Xia, Junlin Liu +2
Single-image self-supervised denoising replaces unavailable clean targets with surrogate targets constructed from noisy observations. Its effectiveness therefore depends on how clo…
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Shuimu Chen, Yuteng Chen, Yuanshen Guan +7
Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evide…
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
Jialiang Yang, Bin Xia, Ruihang Chu +6
Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactio…
MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation
Dongxia Liu, Jie Ma, Xiaochen Yang +7
The creation of cinematic-quality animal effects necessitates the precise modeling of muscle and fur dynamics, a process that remains both labor-intensive and computationally expen…