3 papers
cs.CV2026
USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning
Siru Jiang, Jian Liang, Ran He +1
Test-time adaptation (TTA) has emerged as a popular paradigm for improving the performance of vision-language models (e.g., CLIP) on downstream tasks. Among existing CLIP-based TTA…
cs.AI2026
WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis
Shuo Lu, Yinuo Xu, Kecheng Yu +8
Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural language. Browser-native 3D, co…
cs.AI2026
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
Shuo Lu, Jianjie Cheng, Yinuo Xu +16
Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the…