2 papers
cs.CV2026
TopoAgent: A Structure-Aware Perception-to-Reasoning Framework for Diagram-to-Graph Topology Extraction with Large Vision-Language Models
Bangwei Guo, Xujiang Zhao, Yanchi Liu +7
Diagram-to-graph topology extraction aims to extract a graph of entities and their connections from a structural diagram. This task remains challenging for current vision-language…
cs.CV2026
CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction
Wanhao Liu, Jinsong Lin, Rulin Zhou +11
Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observe…