7 papers
Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models
Xiangjie Sui, Songyang Li, Hanwei Zhu +3
Visual corruptions can change vision--language model (VLM) behavior in ways that top-1 accuracy does not capture. A model may keep the same answer while losing distributional suppo…
Reinforcement Inference: Leveraging Uncertainty for Self-Correcting Language Model Reasoning
Xinhai Sun
Modern large language models (LLMs) are often evaluated and deployed under a one-shot, greedy inference protocol, especially in professional settings that require deterministic beh…
Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
Jiancheng Dong, Pengyue Jia, Jingyu Peng +7
Carefully engineered system prompts play a critical role in guiding the behavior of LLM agents, but their considerable length introduces significant drawbacks, including increased…
Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
Jiahuan Pei, Fanghua Ye, Xin Sun +3
Large language models (LLMs) have advanced virtual educators and learners, bridging NLP with AI4Education. Existing work often lacks scalability and fails to leverage diverse, larg…
Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft
Xin Sun, Lei Wang, Yue Li +5
With large language models (LLMs) on the rise, in-game interactions are shifting from rigid commands to natural conversations. However, the impacts of LLMs on player performance an…
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
Haochen Huang, Jiahuan Pei, Yue Su +11
Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precis…