6 papers
DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions
Weicheng Zheng, Yixin Huang, Qiao Sun +2
Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise. We propose DriveMA, a Driv…
Is Your Trajectory Displacement Safe in Long-tail?
Qiao Sun, Weicheng Zheng, Yixin Huang +1
Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing evaluation pipelines are rarely human-aligne…
DriveMA: Rethinking Language Interfaces in Driving VLAs with One-Step Meta-Actions
Weicheng Zheng, Yixin Huang, Qiao Sun +2
Driving Vision-Language-Action Models (Driving VLAs) commonly introduce natural-language reasoning as an intermediate interface for end-to-end planning, but reasoning-centric inter…
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
Lanxiang Hu, Abhilash Shankarampeta, Yixin Huang +7
The next frontier for video generation lies in developing models capable of zero-shot reasoning, where understanding real-world scientific laws is crucial for accurate physical out…
RUIE: Retrieval-based Unified Information Extraction using Large Language Model
Xincheng Liao, Junwen Duan, Yixi Huang +1
Unified information extraction (UIE) aims to extract diverse structured information from unstructured text. While large language models (LLMs) have shown promise for UIE, they requ…
TrustLLM: Trustworthiness in Large Language Models
Yue Huang, Lichao Sun, Haoran Wang +67
Large language models (LLMs), exemplified by ChatGPT, have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs prese…