12 papers
SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
Weihan Meng, Hongzhu Guo, Yi Jing +5
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on extern…
ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics
Zanwei Zhou, Jiazhong Cen, Jiemin Fang +7
Precise control over complex dynamics remains challenging for modern video generative models, as text prompts alone often cannot specify physically plausible, fine-grained motion a…
UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement
Jingwei Yang, Ruoxi Wu, Wei Shen +4
Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, leading to reference-content l…
Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset
Yuhai Deng, Huimin She, Wei Shen +4
Tone style transfer for photo retouching aims to adapt the stylistic tone of the reference image to a given content image. However, the lack of high-quality large-scale triplet dat…
RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
Yushuai Song, Weize Quan, Weining Wang +8
Recent advances in generative super-resolution (SR) have greatly improved visual realism, yet existing evaluation and optimization frameworks remain misaligned with human perceptio…
Text-Image Conditioned 3D Generation
Jiazhong Cen, Jiemin Fang, Sikuang Li +8
High-quality 3D assets are essential for VR/AR, industrial design, and entertainment, motivating growing interest in generative models that create 3D content from user prompts. Mos…