3 papers
cs.CV2026
BookNet: Book Image Rectification via Cross-Page Attention Network
Shaokai Liu, Hao Feng, Bozhi Luan +3
Book image rectification presents unique challenges in document image processing due to complex geometric distortions from binding constraints, where left and right pages exhibit d…
cs.CV2025
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
Ganlin Yang, Tianyi Zhang, Haoran Hao +15
While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (V…
cs.LG2025
Evaluating Loss Functions for Graph Neural Networks: Towards Pretraining and Generalization
Khushnood Abbas, Ruizhe Hou, Zhou Wengang +4
Graph Neural Networks (GNNs) became useful for learning on non-Euclidean data. However, their best performance depends on choosing the right model architecture and the training obj…