collaborators

6 papers

cs.CV2026

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer

Ying Shen, Zhiyang Xu, Jiuhai Chen +6

Recent advances in multimodal foundation models unifying image understanding and generation have opened exciting avenues for tackling a wide range of vision-language tasks within a…

cs.CV2026

SuperFlow: Training Flow Matching Models with RL on the Fly

Kaijie Chen, Zhiyang Xu, Ying Shen +3

Recent progress in flow-based generative models and reinforcement learning (RL) has improved text-image alignment and visual quality. However, current RL training for flow models s…

cs.CL2025

LLM Braces: Straightening Out LLM Predictions with Relevant Sub-Updates

Ying Shen, Lifu Huang

Recent findings reveal that much of the knowledge in a Transformer-based Large Language Model (LLM) is encoded in its feed-forward (FFN) layers, where each FNN layer can be interpr…

cs.CL2025

Modality-Specialized Synergizers for Interleaved Vision-Language Generalists

Zhiyang Xu, Minqian Liu, Ying Shen +5

Recent advancements in Vision-Language Models (VLMs) have led to the emergence of Vision-Language Generalists (VLGs) capable of understanding and generating both text and images. H…

cs.CV2025

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation

Kaijie Chen, Zihao Lin, Zhiyang Xu +5

Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a w…

cs.CV2025

SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models

Yue Zhang, Zhiyang Xu, Ying Shen +2

Integrating the 3D world into large language models (3D-based LLMs) has been a promising research direction for 3D scene understanding. However, current 3D-based LLMs fall short in…