5 papers
VibeFlow: Versatile Video Chroma-Lux Editing through Self-Supervised Learning
Yifan Li, Pei Cheng, Bin Fu +2
Video chroma-lux editing, which aims to modify illumination and color while preserving structural and temporal fidelity, remains a significant challenge. Existing methods typically…
Token Pruning for In-Context Generation in Diffusion Transformers
Junqing Lin, Xingyu Zheng, Pei Cheng +3
In-context generation significantly enhances Diffusion Transformers (DiTs) by enabling controllable image-to-image generation through reference examples. However, the resulting inp…
ReMix: Towards a Unified View of Consistent Character Generation and Editing
Benjia Zhou, Bin Fu, Pei Cheng +3
Recent advances in large-scale text-to-image diffusion models (e.g., FLUX.1) have greatly improved visual fidelity in consistent character generation and editing. However, existing…
AppAgent v2: Advanced Agent for Flexible Mobile Interactions
Yanda Li, Chi Zhang, Wenjia Jiang +6
With the advancement of Multimodal Large Language Models (MLLM), LLM-driven visual agents are increasingly impacting software interfaces, particularly those with graphical user int…
X-Intelligence 3.0: Training and Evaluating Reasoning LLM for Semiconductor Display
Xiaolin Yan, Yangxing Liu, Jiazhang Zheng +53
Large language models (LLMs) have recently achieved significant advances in reasoning and demonstrated their advantages in solving challenging problems. Yet, their effectiveness in…