works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

Renjie Liang, Zijian Xu, Jinqian Pan +6

A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed befor…

cs.CV2026

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Shawn Li, Wei Yang, Jike Zhong +11

The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…

cs.AI2026

Geometry over Density: Few-Shot Cross-Domain OOD Detection

Shawn Li, You Qin, Jiate Li +4

Out-of-distribution (OOD) detection identifies test samples that fall outside a model's training distribution, a capability critical for safe deployment in high-stakes applications…

cs.CV2026

Audio-Visual Intelligence in Large Foundation Models

You Qin, Kai Liu, Shengqiong Wu +12

Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate…

cs.CV2024

Grounding is All You Need? Dual Temporal Grounding for Video Dialog

You Qin, Wei Ji, Xinze Lan +5

In the realm of video dialog response generation, the understanding of video content and the temporal nuances of conversation history are paramount. While a segment of current rese…

cs.CV2024

Described Spatial-Temporal Video Detection

Wei Ji, Xiangyan Liu, Yingfei Sun +6

Detecting visual content on language expression has become an emerging topic in the community. However, in the video domain, the existing setting, i.e., spatial-temporal video grou…