From the 1 of 17 linked papers with an AI index.
17 papers
Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure
Mingyuan Zhang
The per-instance Jaccard score, or intersection over union (IoU), is standard in multi-label classification and binary segmentation. With labels, its loss matrix has outc…
Mental World Modeling
Hao Fei, Yiran Zhao
The paper introduces Mental World Modeling (MWM), a framework that integrates agents' mental states (beliefs, desires, intentions) into world models, and presents a training‑free b…
Audio-Visual Intelligence in Large Foundation Models
You Qin, Kai Liu, Shengqiong Wu +12
Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate…
SCP: Spatial Causal Prediction in Video
Yanguang Zhao, Jie Yang, Shengqiong Wu +9
Spatial reasoning, the ability to understand spatial relations, causality, and dynamic evolution, is central to human intelligence and essential for real-world applications such as…
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
Shengqiong Wu, Bobo Li, Xinkai Wang +6
Unified Vision-Language Models (UVLMs) aim to advance multimodal learning by supporting both understanding and generation within a single framework. However, existing approaches la…
Modeling Cross-vision Synergy for Unified Large Vision Model
Shengqiong Wu, Lanhu Wu, Mingyang Bao +5
Recent advances in large vision models (LVMs) have shifted from modality-specific designs toward unified architectures that jointly process images, videos, and 3D data. However, ex…