activity
20242026
collaborators

5 papers

cs.AI2026

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

Chen Li, Sijie Cheng, Yuelin Zhang +4

Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose Gra…

cs.CV2026

Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress

Yuelin Zhang, Sijie Cheng, Chen Li +4

Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models…

cs.CV2025

VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management

Hongbo Jin, Qingyuan Wang, Wenhao Zhang +2

Ultra long video understanding remains an open challenge, as existing vision language models (VLMs) falter on such content due to limited context length and inefficient long term m…

cs.CL2025

StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs

Zhicheng Guo, Sijie Cheng, Yuchen Niu +4

The rapid advancement of large language models (LLMs) has spurred significant interest in tool learning, where LLMs are augmented with external tools to tackle complex tasks. Howev…

cs.CV2024

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Sijie Cheng, Kechen Fang, Yangyang Yu +6

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoTh…