Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection
Boyu Han, Qianqian Xu, Shilong Bao +3
In this report, we address the problem of determining whether a user performs an action incorrectly from egocentric video data. To this end, we propose an Understanding-Enhanced Mo…
cs.CV2026
Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
Boyu Han, Qianqian Xu, Shilong Bao +4
The limited understanding capacity of the visual encoder in Contrastive Language-Image Pre-training (CLIP) has become a key bottleneck for downstream performance. This capacity inc…