collaborators

8 papers

cs.RO2026

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Zishuo Li, Bowen Yang, Changtao Miao +29

The paper introduces Open-AoE, a large-scale egocentric video dataset of human manipulation together with a processing and downstream toolchain that provides hand pose, camera traj…

cs.CV2026

AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer

Lingting Zhu, Shengju Qian, Haidi Fan +6

The digital industry demands high-quality, diverse modular 3D assets, especially for user-generated content~(UGC). In this work, we introduce AssetFormer, an autoregressive Transfo…

cs.CV2025

Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation

Yi Wu, Shengju Qian, Lingting Zhu +5

Multimodal autoregressive (AR) models, based on next-token prediction and transformer architecture, have demonstrated remarkable capabilities in various multimodal tasks including…

cs.CV2025

Large Material Gaussian Model for Relightable 3D Generation

Jingrui Ye, Lingting Zhu, Runze Zhang +5

The increasing demand for 3D assets across various industries necessitates efficient and automated methods for 3D content creation. Leveraging 3D Gaussian Splatting, recent large r…

cs.CV2025

AssetDropper: Asset Extraction via Diffusion Models with Reward-Driven Optimization

Lanjiong Li, Guanhua Zhao, Lingting Zhu +4

Recent research on generative models has primarily focused on creating product-ready visual outputs; however, designers often favor access to standardized asset libraries, a domain…

cs.CV2025

StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation

Yi Wu, Lingting Zhu, Shengju Qian +4

In the current research landscape, multimodal autoregressive (AR) models have shown exceptional capabilities across various domains, including visual understanding and generation.…