2 papers
cs.CV2026
RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing
Hao Li, Ju Dai, Feng Zhou +6
Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geometry. Existing methods prima…
cs.CV2026
ApET: Approximation-Error Guided Token Compression for Efficient VLMs
Qiankun Ma, Ziyao Zhang, Haofei Wang +3
Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead an…