19 papers
TeleGapper: On the (un)reliability of Privacy Policies in Telegram Mini apps
Luca Ferrari, Mariano Ceccato, Luca Verderame
Telegram Mini Apps are Web applications embedded within the Telegram client, forming an ecosystem of third-party services within one of the world's most widely used messaging platf…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation
NVIDIA, :, Aarti Basant +32
As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving pol…
APE: Agentic Prompt Enhancer for Image Generation and Editing
Zijian Huang, Jay Zhangjie Wu, Zian Wang +5
Natural language has become a powerful interface for image generation and editing, yet text-guided visual systems remain highly sensitive to prompt formulation. Semantically simila…
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
Fangfu Liu, Kai He, Tianchang Shen +7
World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many gen…
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
Yifan Lu, Qi Wu, Jay Zhangjie Wu +4
Most practical high-resolution text-to-image systems, including latent diffusion and autoregressive models, perform generation in a compact latent space, and a decoder maps the gen…