23 papers
DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation
Zijun Li, Yimin Zhou, Jia Sun +12
Diffusion-based generative AI has achieved remarkable success in e-commerce applications such as virtual try-on, poster generation, and product background synthesis. However, when…
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
Yankai Yang, Yancheng Long, Bin Wen +4
Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos sha…
SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing
Yankai Yang, Yancheng Long, Wei Chen +7
Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a single whole-image reward, which…
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
Yankai Yang, Yancheng Long, Hongyang Wei +12
Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of generative models. For complex tasks such as i…
OneReason Technical Report
OneRec Team, Biao Yang, Boyang Ding +81
Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. Howev…
CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
Jiyuan Wang, Huan Ouyang, Jiuzhou Lin +15
In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global tem…