1 citations · 1 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
Jitai Hao, Ke Yang, Qiang Huang +1
Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, autonomous driving, industrial monitoring, surveillance and ear…
cs.CV2025
Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
Jitai Hao, Hao Liu, Xinyan Xiao +2
Unified Multimodal Models (UMMs) built on shared autoregressive (AR) transformers are attractive for their architectural simplicity. However, we identify a critical limitation: whe…