From the 1 of 7 linked papers with an AI index.
7 papers
PatchHead: Learning Spatial Patch Evidence for Generalizable AI-Generated Image Detection
Shengbo Qi, Hongyi Fang, Benjia Zhou +1
AI-generated image detectors generalize poorly when their training and test images originate from different generators or datasets. Despite the rich spatial representations produce…
Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective
Hongyi Fang, Chuwen Xie, Benjia Zhou +6
Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to-fine prediction. However, t…
Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolution
Hongyi Fang, Jiahui Wu, Yichen Yue +2
Hallucination remains a persistent challenge in generative super-resolution (GSR), where reconstructed results may contain visually plausible yet weakly supported content, structur…
Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction
Tianshun Han, Ziyu Shi, Lijian Liu +7
The paper introduces Human4K, a large-scale dataset of 4K multi-view images with precise SMPL-X motion capture annotations for whole-body 3D human reconstruction, aiming to improve…
RVLF: A Reinforcing Vision-Language Framework for Gloss-Free Sign Language Translation
Zhi Rao, Yucheng Zhou, Benjia Zhou +3
Gloss-free sign language translation (SLT) is hindered by two key challenges: **inadequate sign representation** that fails to capture nuanced visual cues, and **sentence-level sem…
PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
Tianshun Han, Benjia Zhou, Ajian Liu +4
PESTalk is a novel method for generating 3D facial animations with personalized emotional styles directly from speech. It overcomes key limitations of existing approaches by introd…