From the 1 of 9 linked papers with an AI index.
9 papers
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
Jiawei Wang, Ke Rui, Yushen Zuo +2
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guar…
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Simple AI, :, Yuteng Wei +16
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale;…
Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition
Ke Rui, Yushen Zuo, Jiawei Wang +4
The paper investigates why robots fail when chaining language‑conditioned skills for long‑horizon household tasks, introducing a vision‑language‑action harness that checks skill ch…
Emotion Diffusion Classifier with Adaptive Margin Discrepancy Training for Facial Expression Recognition
Rongkang Dong, Cuixin Yang, Cong Zhang +2
Facial Expression Recognition (FER) is essential for human-machine interaction, as it enables machines to interpret human emotions and internal states from facial affective behavio…
SANet: Scale-Adaptive Structure-Affinity Transformation for Spine Segmentation from Ultrasound Volume Projection Imaging
Hao Xie, Zixun Huang, Yushen Zuo +6
Spine segmentation, based on ultrasound volume projection imaging (VPI), plays a vital role for intelligent scoliosis diagnosis in clinical applications. However, this task faces s…
Enhancing Novel View Synthesis from extremely sparse views with SfM-free 3D Gaussian Splatting Framework
Zongqi He, Hanmin Li, Kin-Chung Chan +5
3D Gaussian Splatting (3DGS) has demonstrated remarkable real-time performance in novel view synthesis, yet its effectiveness relies heavily on dense multi-view inputs with precise…