3 papers
cs.LG2026
A-LLM: An End-to-end Conversational Audio Avatar Large Language Model
Xiaolin Hu, Hang Yuan, Xinzhu Sang +4
Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significa…
cs.RO2026
TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
Bin Yu, Shijie Lian, Xiaopeng Lin +8
The fundamental premise of Vision-Language-Action (VLA) models is to harness the extensive general capabilities of pre-trained Vision-Language Models (VLMs) for generalized embodie…
cs.RO2025
PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
Xiaopeng Lin, Shijie Lian, Bin Yu +10
Robotic generalization relies on physical intelligence: the ability to reason about state changes, contact-rich interactions, and long-horizon planning under egocentric perception…