7 papers
Concept Heterogeneity-aware Representation Steering
Laziz U. Abdullaev, Noelle Y. L. Wong, Ryan T. Z. Lee +3
Representation steering offers a lightweight mechanism for controlling the behavior of large language models (LLMs) by intervening on internal activations at inference time. Most e…
MOVA: Towards Scalable and Synchronized Video-Audio Generation
OpenMOSS Team, Donghua Yu, Mingshu Chen +38
Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on casc…
MPJudge: Towards Perceptual Assessment of Music-Induced Paintings
Shiqi Jiang, Tianyi Liang, Huayuan Ye +2
Music induced painting is a unique artistic practice, where visual artworks are created under the influence of music. Evaluating whether a painting faithfully reflects the music th…
Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
Nan Xiang, Tianyi Liang, Haiwen Huang +6
Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While…
PPJudge: Towards Human-Aligned Assessment of Artistic Painting Process
Shiqi Jiang, Xinpeng Li, Xi Mao +2
Artistic image assessment has become a prominent research area in computer vision. In recent years, the field has witnessed a proliferation of datasets and methods designed to eval…
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation
Tianyi Liang, Jiangqi Liu, Yifei Huang +4
Text-to-image (T2I) generation has made remarkable progress in producing high-quality images, but a fundamental challenge remains: creating backgrounds that naturally accommodate t…