3 papers
cs.CV2026
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
Chaoda Zheng, Sean Li, Jinhao Deng +9
Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams…
cs.CL2025
Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
Ziheng Chi, Yifan Hou, Chenxi Pang +3
Diagrams convey symbolic information in a visual format rather than a linear stream of words, making them especially challenging for AI models to process. While recent evaluations…
cs.LG2024
Could Chemical LLMs benefit from Message Passing
Jiaqing Xie, Ziheng Chi
Pretrained language models (LMs) showcase significant capabilities in processing molecular text, while concurrently, message passing neural networks (MPNNs) demonstrate resilience…