4 papers
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
Chen Li, Sijie Cheng, Yuelin Zhang +4
Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose Gra…
Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable
Tianyuan Zou, Liang Yue, Yang Liu +2
With the growing prevalence of always-on hardware such as smart glasses, body cameras, and home security systems, life-logging visual sensing is becoming inevitable, forming the ba…
Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
Yuelin Zhang, Sijie Cheng, Chen Li +4
Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models…
Instruction-Guided Visual Masking
Jinliang Zheng, Jianxiong Li, Sijie Cheng +6
Instruction following is crucial in contemporary LLM. However, when extended to multimodal setting, it often suffers from misalignment between specific textual instruction and targ…