From the 1 of 5 linked papers with an AI index.
5 papers
Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
Menglin Han, Yang Ding, Yulei Lu +6
Real-time long-form avatar audio-video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained…
LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
Zhilin Wu, Zhangkai Ni, Chengmei Yang +4
The paper introduces LoMeVQA, a large benchmark of 206K longitudinal medical visual question answering pairs designed to evaluate temporal reasoning over sequential medical images,…
EntropyPrune: Matrix Entropy Guided Visual Token Pruning for Multimodal Large Language Models
Yahong Wang, Juncheng Wu, Zhangkai Ni +6
Multimodal large language models (MLLMs) incur substantial inference cost due to the processing of hundreds of visual tokens per image. Although token pruning has proven effective…
FlowBypass: Rectified Flow Trajectory Bypass for Training-Free Image Editing
Menglin Han, Zhangkai Ni
Training-free image editing has attracted increasing attention for its efficiency and independence from training data. However, existing approaches predominantly rely on inversion-…
Perceptual-GS: Scene-adaptive Perceptual Densification for Gaussian Splatting
Hongbi Zhou, Zhangkai Ni
3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis. However, existing methods struggle to adaptively optimize the distribution of Gaussian pr…