collaborators

5 papers

cs.CV2026

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

Utkarsh A. Mishra, Yongxin Chen, Danfei Xu +3

Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on ro…

cs.CL2026

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

Baochang Ren, Xinjie Liu, Xi Chen +15

Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach. AI can help read lit…

cs.CV2026

Factuality Matters: When Image Generation and Editing Meet Structured Visuals

Le Zhuo, Songhao Han, Yuandong Pu +8

While modern visual generation models excel at creating aesthetically pleasing natural images, they struggle with producing or editing structured visuals like charts, diagrams, and…

cs.CV2026

PICABench: How Far Are We from Physically Realistic Image Editing?

Yuandong Pu, Le Zhuo, Songhao Han +10

Image editing has achieved remarkable progress recently. Modern editing models could already follow complex instructions to manipulate the original content. However, beyond complet…

cs.CL2025

GenieBlue: Integrating both Linguistic and Multimodal Capabilities for Large Language Models on Mobile Devices

Xudong Lu, Yinghao Chen, Renshou Wu +10

Recent advancements in Multimodal Large Language Models (MLLMs) have enabled their deployment on mobile devices. However, challenges persist in maintaining strong language capabili…