3 papers
cs.CV2026
Structured Relational Reasoning for Group Activity Assessment
Thinesh Thiyakesan Ponbagavathi, Chengzheng Yang, Alina Roitberg
Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DINOv2, offer excellent features b…
cs.CL2026
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
Cheng Yang, Chufan Shi, Bo Shui +7
Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent their representations…
cs.CV2026
Optical Context Compression Is Just (Bad) Autoencoding
Ivan Yee Lee, Cheng Yang, Taylor Berg-Kirkpatrick
DeepSeek-OCR shows that rendered text can be reconstructed from a small number of vision tokens, sparking excitement about using vision as a compression medium for long textual con…