8 papers · 1 filter
ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization
Yuanhe Guo, Linxi Xie, Zhuoran Chen +5
We introduce ImageGem, a dataset for studying generative models that understand fine-grained individual preferences. We posit that a key challenge hindering the development of such…
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
Jan Ackermann, Kiyohiro Nakayama, Guandao Yang +2
Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplore…
Animal Pose Labeling Using General-Purpose Point Trackers
Zhuoyang Pan, Boxiao Pan, Guandao Yang +2
Automatically estimating animal poses from videos is important for studying animal behaviors. Existing methods do not perform reliably since they are trained on datasets that are n…
AIpparel: A Multimodal Foundation Model for Digital Garments
Kiyohiro Nakayama, Jan Ackermann, Timur Levent Kesdogan +6
Apparel is essential to human life, offering protection, mirroring cultural identities, and showcasing personal style. Yet, the creation of garments remains a time-consuming proces…
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
Youming Deng, Wenqi Xian, Guandao Yang +4
In this paper, we present a self-calibrating framework that jointly optimizes camera parameters, lens distortion and 3D Gaussian representations, enabling accurate and efficient sc…
InfoGaussian: Structure-Aware Dynamic Gaussians through Lightweight Information Shaping
Yunchao Zhang, Guandao Yang, Leonidas Guibas +1
3D Gaussians, as a low-level scene representation, typically involve thousands to millions of Gaussians. This makes it difficult to control the scene in ways that reflect the under…