activity
20192026
most citedFoundational Models Defining a New Era in Vision: A Survey and Outlook

68 citations · 201 across the 60 of their papers we have counts for

collaborators
Showing cs.CVShow all

65 papers · 1 filter

cs.CV2026

Gaussian-JEPA: Joint-Embedding Predictive Learning for 3D Gaussian Splats

Bin Ren, Qi Ma, Yue Li +7

3D Gaussian Splatting (3DGS) represents 3D content with anisotropic primitives that jointly encode geometry and appearance. Fixed-budget encoders consume sampled observations of Ga…

cs.CV2026

Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM

Amol Harsh, Zongyan Han, Jean Lahoud +5

Natural-language queries about 3D environments become actionable when responses are verifiable and metric. Verifiability requires explicit grounding to the referred 3D region, whil…

cs.CV2026

MAviS: A Multimodal Conversational Assistant For Avian Species

Yevheniia Kryklyvets, Mohammed Irfan Kurpath, Sahal Shaji Mullappilly +5

Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monitoring. However, existing multimo…

cs.CV2026

Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation

Jinxing Zhou, Yanghao Zhou, Yaoting Wang +5

Language-referred audio-visual segmentation (Ref-AVS) aims to segment target objects described by natural language by jointly reasoning over video, audio, and text. Beyond generati…

cs.CV2026

DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding

Shubham Patle, Sara Ghaboura, Hania Tariq +4

Arabic calligraphy represents one of the richest visual traditions of the Arabic language, blending linguistic meaning with artistic form. Although multimodal models have advanced…

cs.CV2025

A Benchmark for Omni-Modal Reasoning in Long Videos

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou +12

Long-form omni-modal video understanding requires integrating vision, speech, and ambient audio with coherent long-context reasoning. Existing video benchmarks often trade off temp…