From the 3 of 17 linked papers with an AI index.
1 citations · 1 across the 14 of their papers we have counts for
14 papers · 1 filter
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Jiaxing Li, Kai Zou, Cindy Zhou +7
The paper studies autoregressive video distillation, showing that aligning the student model’s mode coverage with the teacher’s distribution improves generation quality and diversi…
ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID
Xulin Li, Yan Lu, Bin Liu +5
The paper proposes ANFI, an adaptive neighbor feature interaction method for person re-identification that jointly models affinity and discrepancy relations to mitigate the impact…
VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
Yupeng Zheng, Kai Zou, Bin Liu +1
VisCo introduces a training-efficient self-compression framework that reuses a pretrained vision-language model as an intrinsic autoencoder to compress visual tokens into a small s…
SGF-CDNet: A Consistency-Discrepancy Graph Network over Semantic-Geometric Fused Nodes for Face Forgery Detection
Jiayao Jiang, Bin Liu, Nenghai Yu
The rapid advancement of deepfakes necessitates robust face forgery detection. Although forged faces may lack obvious artifacts, they often contain subtle disharmony among differen…
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety
Shikai Qiu, Xiaowen Xu, Benlei Cui +55
General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI s…
ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers
Ruiliang Zhou, Xuecheng Wu, Kang He +6
While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computational bottleneck. Existing s…