Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable
Tianyuan Zou, Liang Yue, Yang Liu +2
With the growing prevalence of always-on hardware such as smart glasses, body cameras, and home security systems, life-logging visual sensing is becoming inevitable, forming the ba…
cs.CV2026
Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
Yuelin Zhang, Sijie Cheng, Chen Li +4
Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models…
cs.CV2024
Instruction-Guided Visual Masking
Jinliang Zheng, Jianxiong Li, Sijie Cheng +6
Instruction following is crucial in contemporary LLM. However, when extended to multimodal setting, it often suffers from misalignment between specific textual instruction and targ…