3 papers
cs.AI2026
Structured Self-Consistency:A Multi-Task Evaluation of LLMs on VirtualHome
Jiaqi Xu, Tao Huang, Kai Zhang
Embodied AI requires agents to understand goals, plan actions, and execute tasks in simulated environments. We present a comprehensive evaluation of Large Language Models (LLMs) on…
cs.CV2025
Learning to Expand Images for Efficient Visual Autoregressive Modeling
Ruiqing Yang, Kaixin Zhang, Zheng Zhang +2
Autoregressive models have recently shown great promise in visual generation by leveraging discrete token sequences akin to language modeling. However, existing approaches often su…
cs.CV2025
ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
Kaixin Zhang, Ruiqing Yang, Yuan Zhang +2
Visual Autoregressive (VAR) models enable efficient image generation via next-scale prediction but face escalating computational costs as sequence length grows. Existing static pru…