works on

From the 1 of 9 linked papers with an AI index.

activity
20242026
collaborators

9 papers

cs.LG2026

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

Jiawei Wang, Ke Rui, Yushen Zuo +2

JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guar…

cs.RO2026

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Simple AI, :, Yuteng Wei +16

Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale;…

cs.RO2026

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition

Ke Rui, Yushen Zuo, Jiawei Wang +4

The paper investigates why robots fail when chaining language‑conditioned skills for long‑horizon household tasks, introducing a vision‑language‑action harness that checks skill ch…

cs.CV2026

Emotion Diffusion Classifier with Adaptive Margin Discrepancy Training for Facial Expression Recognition

Rongkang Dong, Cuixin Yang, Cong Zhang +2

Facial Expression Recognition (FER) is essential for human-machine interaction, as it enables machines to interpret human emotions and internal states from facial affective behavio…

cs.CV2025

SANet: Scale-Adaptive Structure-Affinity Transformation for Spine Segmentation from Ultrasound Volume Projection Imaging

Hao Xie, Zixun Huang, Yushen Zuo +6

Spine segmentation, based on ultrasound volume projection imaging (VPI), plays a vital role for intelligent scoliosis diagnosis in clinical applications. However, this task faces s…

cs.CV2025

Enhancing Novel View Synthesis from extremely sparse views with SfM-free 3D Gaussian Splatting Framework

Zongqi He, Hanmin Li, Kin-Chung Chan +5

3D Gaussian Splatting (3DGS) has demonstrated remarkable real-time performance in novel view synthesis, yet its effectiveness relies heavily on dense multi-view inputs with precise…