4 papers
ManGo: Manga Active Narrative Grounding Optimization
Hao Qiu, Junyan Wang, Zheyuan Liu +4
Manga visual question answering requires models to answer questions over panel-based visual narratives, where relevant evidence is distributed across ordered panels, embedded text,…
LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding
Zhewei Zhang, Puyue Wang, Guanren Qiao +10
Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM featu…
HoRD: Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation
Puyue Wang, Jiawei Hu, Yan Gao +7
Humanoid robots can suffer significant performance drops under small changes in dynamics, task specifications, or environment setup. We propose HoRD, a two-stage learning framework…
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
Letian Cheng, Junyan Wang, Yan Gao +3
Perplexity is a widely adopted metric for assessing the predictive quality of large language models (LLMs) and often serves as a reference metric for downstream evaluations. Howeve…