Publications (26)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
Andrew Z. Wang, Songwei Ge, Tero Karras +2
Coherent Zero-Shot Visual Instruction Generation
Quynh Phung, Songwei Ge, Jia-Bin Huang
Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
Songwei Ge, Seungjun Nah, Guilin Liu +7
Getting Topology and Point Cloud Generation to Mesh
Austin Dill, Chun-Liang Li, Songwei Ge +1
Hyperbolic Contrastive Learning for Visual Representations beyond Objects
Songwei Ge, Shlok Mishra, Simon Kornblith +2
Hallucinating Point Cloud into 3D Sculptural Object
Chun-Liang Li, Eunsu Kang, Songwei Ge +4
Creative Sketch Generation
Songwei Ge, Vedanuj Goswami, C. Lawrence Zitnick +1
Learning Robust Global Representations by Penalizing Local Predictive Power
Haohan Wang, Songwei Ge, Eric P. Xing +1
PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos
Ting-Hsuan Liao, Haowen Liu, Yiran Xu +3
Expressive Text-to-Image Generation with Rich Text
Songwei Ge, Taesung Park, Jun-Yan Zhu +1
Rethinking Score Distillation as a Bridge Between Image Distributions
David McAllister, Songwei Ge, Jia-Bin Huang +4
Cosmos World Foundation Model Platform for Physical AI
NVIDIA, :, Niket Agarwal +76
Illusion3D: 3D Multiview Illusion with 2D Diffusion Priors
Yue Feng, Vaibhav Sanjay, Spencer Lutz +3
Shift Invariance Can Reduce Adversarial Robustness
Songwei Ge, Vasu Singla, Ronen Basri +1
Flow Matching Policy Gradients
David McAllister, Songwei Ge, Brent Yi +5
Text-driven Visual Synthesis with Latent Diffusion Prior
Ting-Hsuan Liao, Songwei Ge, Yiran Xu +3
Learned Interpolation for 3D Generation
Austin Dill, Songwei Ge, Eunsu Kang +2
From Text to Sound: A Preliminary Study on Retrieving Sound Effects to Radio Stories
Songwei Ge, Curtis Xuan, Ruihua Song +3
Developing Creative AI to Generate Sculptural Objects
Songwei Ge, Austin Dill, Eunsu Kang +4
Personalizing Search Results Using Hierarchical RNN with Query-aware Attention
Songwei Ge, Zhicheng Dou, Zhengbao Jiang +2
MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration
Thomas Hayes, Songyang Zhang, Xi Yin +6
On the Content Bias in Fréchet Video Distance
Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar +2
Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
Songwei Ge, Thomas Hayes, Harry Yang +5
Visual Conceptual Blending with Large-scale Language and Vision Models
Songwei Ge, Devi Parikh
Robust Contrastive Learning Using Negative Samples with Diminished Semantics
Songwei Ge, Shlok Mishra, Haohan Wang +2
Grounded Text-to-Image Synthesis with Attention Refocusing
Quynh Phung, Songwei Ge, Jia-Bin Huang