Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
Tianyi Xiong, Zhengyuan Yang, Xiaofei Wang +10
Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identify a fundamental obstacle, ter…
cs.CV2025
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
Heyu Guo, Shanmu Wang, Ruichun Ma +5
Vision-language-action (VLA) models have shown strong generalization for robotic action prediction through large-scale vision-language pretraining. However, most existing models re…