From the 1 of 8 linked papers with an AI index.
8 papers
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
Shufan Li, Konstantinos Kallidromitis, Hritik Bansal +7
LaViDa introduces a diffusion-based vision-language model that combines a vision encoder with discrete diffusion to enable fast parallel decoding and controllable multimodal genera…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization
Shufan Li, Konstantinos Kallidromitis, Akash Gokul +2
Group-advantage-based reinforcement learning methods, such as GRPO and DAPO, have demonstrated strong performance across diverse domains, including mathematical reasoning and text-…
MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
Shufan Li, Konstantinos Kallidromitis, Akash Gokul +3
World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face prac…
Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
Hiroaki Hayashi, Bo Pang, Wenting Zhao +6
Large language model (LLM) based agents are increasingly used to tackle software engineering tasks that require multi-step reasoning and code modification, demonstrating promising…
xGen-small Technical Report
Erik Nijkamp, Bo Pang, Egor Pakhomov +5
We introduce xGen-small, a family of 4B and 9B Transformer decoder models optimized for long-context applications. Our vertically integrated pipeline unites domain-balanced, freque…