3 papers
cs.CV2026
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time
Ruizhi Zhang, Ye Huang, Yuangang Pan +6
While artificial intelligence has mastered structured games like chess and Go, vision-language agents still struggle in visually-driven 3D games without access to game states. Exis…
cs.SE2026
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
Zhilin Liu, Ye Huang, Ting Xie +3
Recent advances in Large Language Model (LLM)-based agents have shown remarkable progress in code generation. However, current agent methods mainly rely on text-output-based feedba…
cs.CV2025
VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer
Xikai Tang, Ye Huang, Guangqiang Yin +1
We present VPNeXt, a new and simple model for the Plain Vision Transformer (ViT). Unlike the many related studies that share the same homogeneous paradigms, VPNeXt offers a fresh p…