Showing 2026Show all
3 papers · 1 filter
cs.CV2026
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
Bo Gu, Zhikang Zhang, Zizhuang Wei +3
Recent multimodal large language models (MLLMs) have made remarkable progress in visual understanding and language-based reasoning, yet they lack a persistent world-centered repres…
cs.CV2026
Clinical-Prior Guided Multi-Modal Learning with Latent Attention Pooling for Gait-Based Scoliosis Screening
Dong Chen, Zizhuang Wei, Jialei Xu +6
Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity whose progression can be mitigated through early detection. Conventional screening methods are often subjectiv…
cs.CV2026
Token Entropy Regularization for Multi-modal Antenna Affiliation Identification
Dong Chen, Ruoyu Li, Xinyan Zhang +5
Accurate antenna affiliation identification is crucial for optimizing and maintaining communication networks. Current practice, however, relies on the cumbersome and error-prone pr…