3 papers
cs.CL2026
UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models
Cheng Yang, Chufan Shi, Bo Shui +7
Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent textual and visual mo…
cs.CV2025
Optical Context Compression Is Just (Bad) Autoencoding
Ivan Yee Lee, Cheng Yang, Taylor Berg-Kirkpatrick
DeepSeek-OCR shows that rendered text can be reconstructed from a small number of vision tokens, sparking excitement about using vision as a compression medium for long textual con…
cs.CV2025
Structured Relational Reasoning for Group Activity Assessment
Thinesh Thiyakesan Ponbagavathi, Chengzheng Yang, Alina Roitberg
Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DINOv2, offer excellent features b…