12 papers
SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
Weihan Meng, Hongzhu Guo, Yi Jing +5
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on extern…
ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics
Zanwei Zhou, Jiazhong Cen, Jiemin Fang +7
Precise control over complex dynamics remains challenging for modern video generative models, as text prompts alone often cannot specify physically plausible, fine-grained motion a…
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
Lianghui Zhu, Yingyue Li, Jiemin Fang +4
Transformer has been very successful in various computer vision tasks and understanding the working mechanism of transformer is important. As touchstones, weakly-supervised semanti…
Text-Image Conditioned 3D Generation
Jiazhong Cen, Jiemin Fang, Sikuang Li +8
High-quality 3D assets are essential for VR/AR, industrial design, and entertainment, motivating growing interest in generative models that create 3D content from user prompts. Mos…
Dereflection Any Image with Diffusion Priors and Diversified Data
Jichen Hu, Chen Yang, Zanwei Zhou +4
Reflection removal of a single image remains a highly challenging task due to the complex entanglement between target scenes and unwanted reflections. Despite significant progress,…
WorldGrow: Generating Infinite 3D World
Sikuang Li, Chen Yang, Jiemin Fang +6
We tackle the challenge of generating the infinitely extendable 3D world -- large, continuous environments with coherent geometry and realistic appearance. Existing methods face ke…