Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Causal World Modeling for Robot Control
Lin Li, Qihang Zhang, Yiming Luo +9
This work highlights that video world modeling, alongside vision-language pre-training, establishes a fresh and independent foundation for robot learning. Intuitively, video world…
cs.CV2024
OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
Ming Hu, Peng Xia, Lin Wang +11
Surgical scene perception via videos is critical for advancing robotic surgery, telesurgery, and AI-assisted surgery, particularly in ophthalmology. However, the scarcity of divers…