10 papers
VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation
Jiajun Xu, Yanghao Zhou, Jingyun Liao +6
Natural-language-driven "vibe coding" enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace.…
Diverse Thinking Schemata Elicit Better Reasoning in Large Language Models
Xinyue Liang, Yizhe Yang, Yu Bai +3
Large reasoning models (LRMs) have attracted increasing attention for their ability to solve complex mathematical problems by generating extended reasoning chains. In this work, we…
TrustMargin: Training-Free Arbitration between Parametric Memory and Retrieved Evidence in Large Language Models
Jingyan Xu, Hong Shi, Yi Shan +4
Large language models answer knowledge-intensive questions using both parametric memory and retrieved evidence, but neither source is uniformly reliable. Retrieval can fill knowled…
SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization
Huashan Sun, Shengyi Liao, Yansen Han +8
Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-world long-context information, primar…
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
Haitian Li, Yanghao Zhou, Heyan Huang +15
In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-visual alignment. However, th…
ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control
Yanghao Zhou, Jingyu Ma, Yibo Peng +3
Humanoid control systems have made significant progress in recent years, yet modeling fluent interaction-rich behavior between a robot, its surrounding environment, and task-releva…