4 papers
SGS: Structured Sparse Gaussian Streaming for Efficient Free-Viewpoint Video Reconstruction on Edge-IoT Devices
Yiwei Li, Jiannong Cao, Weixun Gao +4
Streaming reconstruction of Free-Viewpoint Videos (FVVs) supports immersive Internet of Things (IoT) services, such as telepresence and digital twin visualization. Existing methods…
Demystifying Agent Skills: Why They Work-Until They Don't
Zhiyuan Jiang, Fangrui Huang, Hanwen Xing +6
Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structured packages of knowledge. However, existing evaluations largely…
PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration
Han Wang, Zijun Wang, Shuoshuo Xue +5
Action-conditioned world models are a key component of embodied AI, serving as scalable policy evaluators that reduce reliance on expensive real-world rollouts. To accurately captu…
FactCheck: Feasibility-aware Long-term Action Anticipation with Multi-agent Collaboration
Rui Cao, Jiannong Cao, Bo Yuan +2
Long-term action anticipation (LTA) aims to predict an ordered sequence of future verb-noun actions from a partially observed video. While this task serves as the foundation for em…