Publications (27)
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
Zhihao Xie, Junfeng Wu, Xinting Hu +2
VideoRAE is a representation autoencoder that leverages frozen video foundation model features to create compact, generation‑friendly video latents, supporting both continuous diff…
LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers
Jingze Zhu, Yongliang Wu, Wenbo Zhu +7
SoccerNet 2025 Challenges Results
Silvio Giancola, Anthony Cioppa, Marc Gutiérrez-Pérez +115
DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition
Yiyan Xu, Wuqiang Zheng, Wenjie Wang +5
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
Xinting Hu, Haoran Wang, Jan Eric Lenssen +1
Distilling Causal Effect of Data in Class-Incremental Learning
Xinting Hu, Kaihua Tang, Chunyan Miao +2
Learning to Segment the Tail
Xinting Hu, Yi Jiang, Kaihua Tang +3
SemanticNVS: Improving Semantic Scene Understanding in Generative Novel View Synthesis
Xinya Chen, Christopher Wewer, Jiahao Xie +2
Edit360: 2D Image Edits to 3D Assets from Any Angle
Junchao Huang, Xinting Hu, Shaoshuai Shi +2
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
Yongliang Wu, Yizhou Zhou, Zhou Ziheng +7
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
Bo Wang, Jiehong Lin, Chenzhi Liu +5
Do Instance Priors Help Weakly Supervised Semantic Segmentation?
Anurag Das, Anna Kukleva, Xinting Hu +2
SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
Ouxiang Li, Yuan Wang, Xinting Hu +3
Personalized Generation In Large Model Era: A Survey
Yiyan Xu, Jinghao Zhang, Alireza Salemi +6
On Non-Random Missing Labels in Semi-Supervised Learning
Xinting Hu, Yulei Niu, Chunyan Miao +2
Third Party Risk Modelling and Assessment for Safe UAV Path Planning in Metropolitan Environments
Bizhao Pang, Xinting Hu, Wei Dai +1
Mimic In-Context Learning for Multimodal Tasks
Yuchu Jiang, Jiale Fu, Chenduo Hao +4
Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
Junchao Huang, Xinting Hu, Boyao Han +4
LIVE: Learnable In-Context Vector for Visual Question Answering
Yingzhe Peng, Chenduo Hao, Xu Yang +3
Number it: Temporal Grounding Videos like Flipping Manga
Yongliang Wu, Xinting Hu, Yuyang Sun +5
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
Anurag Das, Xinting Hu, Li Jiang +1
LIVE: Long-horizon Interactive Video World Modeling
Junchao Huang, Ziyang Ye, Xinting Hu +5
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
Yongliang Wu, Zonghui Li, Xinting Hu +7
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
Yifei Yu, Xiaoshan Wu, Xinting Hu +8
Veda: Scalable Video Diffusion via Distilled Sparse Attention
Shihao Han, Hao Yang, Xinting Hu +3
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
Ouxiang Li, Yuan Wang, Xinting Hu +7
Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient
Yongliang Wu, Shiji Zhou, Mingzhuo Yang +6