activity
20232026
most citedSAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding

13 citations · 40 across the 21 of their papers we have counts for

collaborators
Showing 2026Show all

6 papers · 1 filter

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…

cs.AI2026

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

Haoxiang Wang, Da Yu, Huishuai Zhang

Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor effects that mask capability gaps…

cs.CV2026

Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

Zhengxian Yang, Shengqi Wang, Shi Pan +8

Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such experiences can be achieved thr…

cs.GR2026

Real-time Neural Six-way Lightmaps

Wei Li, Hanxiao Sun, Tao Huang +4

Participating media are a pervasive and intriguing visual effect in virtual environments. Unfortunately, rendering such phenomena in real-time is notoriously difficult due to the c…

cs.GT2026

Interbank Lending Games

Jinyun Tong, Bart de Keijzer, Haoxiang Wang +1

We define and study a lending game to model the interbank money market, in which lending banks strategically allocate their cash to borrowing banks. The interest rate offered by ea…

cs.CV2026

DuoGen: Towards General Purpose Interleaved Multimodal Generation

Min Shi, Xiaohui Zeng, Jiannan Huang +13

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts f…