9 papers
Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories
Merve Kocabas, Gege Gao, Bernhard Schölkopf +1
Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, wh…
Instance Data Condensation for Image Super-Resolution
Tianhao Peng, Ho Man Kwan, Yuxuan Jiang +5
Deep learning based Image Super-Resolution (ISR) relies on large training datasets to optimize model generalization; this requires substantial computational and storage resources d…
SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language Segmentation
Chengxi Zeng, Yuxuan Jiang, Ge Gao +6
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ende…
Ultra-lightweight Neural Video Representation Compression
Ho Man Kwan, Tianhao Peng, Ge Gao +4
Recent works have demonstrated the viability of utilizing over-fitted implicit neural representations (INRs) as alternatives to autoencoder-based models for neural video compressio…
NVRC: Neural Video Representation Compression
Ho Man Kwan, Ge Gao, Fan Zhang +2
Recent advances in implicit neural representation (INR)-based video coding have demonstrated its potential to compete with both conventional and other learning-based approaches. Wi…
View-Consistent Diffusion Representations for 3D-Consistent Video Generation
Duolikun Danier, Ge Gao, Steven McDonagh +3
Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated vid…