collaborators

6 papers

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…

cs.CV2025

Unleashing Perception-Time Scaling to Multimodal Reasoning Models

Yifan Li, Zhenghao Chen, Ziheng Wu +7

Recent advances in inference-time scaling, particularly those leveraging reinforcement learning with verifiable rewards, have substantially enhanced the reasoning capabilities of L…

cs.CV2025

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Yufei Zhan, Ziheng Wu, Yousong Zhu +10

Despite notable advancements in multimodal reasoning, leading Multimodal Large Language Models (MLLMs) still underperform on vision-centric multimodal reasoning tasks in general sc…

cs.CV2025

Valley: Video Assistant with Large Language model Enhanced abilitY

Ruipu Luo, Ziwang Zhao, Min Yang +6

Large Language Models (LLMs), with remarkable conversational capability, have emerged as AI assistants that can handle both visual and textual modalities. However, their effectiven…

cs.CV2025

VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Zejun Li, Ruipu Luo, Jiwen Zhang +3

While large multi-modal models (LMMs) have exhibited impressive capabilities across diverse tasks, their effectiveness in handling complex tasks has been limited by the prevailing…

cs.CV2025

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Ziheng Wu, Zhenghao Chen, Ruipu Luo +6

Recently, vision-language models have made remarkable progress, demonstrating outstanding capabilities in various tasks such as image captioning and video understanding. We introdu…