From the 1 of 9 linked papers with an AI index.
9 papers
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
Weiquan Lin, Yu Deng, Shiyang Liu +6
This survey systematically reviews how large foundation models provide geometric, semantic, and visual priors for hand‑object interaction tasks such as reconstruction and generatio…
FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
Ruicheng Li, Qixiu Li, Ruichun Ma +8
Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by condi…
GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs
Yu Deng
Activation steering controls model behavior by modifying intermediate hidden states at inference time without retraining. Existing methods handle only single-direction injection; w…
CogniFold: Always-On Proactive Memory via Cognitive Folding
Suli Wang, Yiqun Duan, Yu Deng +6
Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genui…
Kintsugi: Learning Policies by Repairing Executable Knowledge Bases
Teng Cao, Yu Deng, Hikaru Shindo +6
Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory, making individual policy kn…
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
Huizhi Liang, Yichao Shen, Yu Deng +5
Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D…