2 papers
cs.CV2026
Spatial Temporal Synergy: Balancing Change and Invariance in Text Driven 3D Human Motion Editing
Shaohui Lin, Zhenwu Shi, Jingyu Gong +5
Text-driven human motion editing aims to modify existing motion sequences according to natural language instructions while maintaining the structural consistency of the original mo…
cs.CV2026
SA-GEM: Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning for Efficient Remote Sensing Large Vision-Language Models
Kexin Ma, Jing Xiao, Bowen Xing +2
RS-LVLMs have advanced multimodal understanding of Earth observation imagery, yet their performance is fundamentally constrained by high-resolution processing, as visual token coun…