works on

From the 3 of 17 linked papers with an AI index.

activity
20242026
most citedAdvancing Aesthetic Image Generation via Composition Transfer

1 citations · 1 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2026

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

Jiaxing Li, Kai Zou, Cindy Zhou +7

The paper studies autoregressive video distillation, showing that aligning the student model’s mode coverage with the teacher’s distribution improves generation quality and diversi…

cs.CV2026

ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID

Xulin Li, Yan Lu, Bin Liu +5

The paper proposes ANFI, an adaptive neighbor feature interaction method for person re-identification that jointly models affinity and discrepancy relations to mitigate the impact…

cs.CV2026

VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

Yupeng Zheng, Kai Zou, Bin Liu +1

VisCo introduces a training-efficient self-compression framework that reuses a pretrained vision-language model as an intrinsic autoencoder to compress visual tokens into a small s…

cs.CV2026

SGF-CDNet: A Consistency-Discrepancy Graph Network over Semantic-Geometric Fused Nodes for Face Forgery Detection

Jiayao Jiang, Bin Liu, Nenghai Yu

The rapid advancement of deepfakes necessitates robust face forgery detection. Although forged faces may lack obvious artifacts, they often contain subtle disharmony among differen…

cs.CV2026

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

Shikai Qiu, Xiaowen Xu, Benlei Cui +55

General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI s…

cs.CV2026

ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers

Ruiliang Zhou, Xuecheng Wu, Kang He +6

While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computational bottleneck. Existing s…