7 papers
MOVA: Towards Scalable and Synchronized Video-Audio Generation
OpenMOSS Team, Donghua Yu, Mingshu Chen +38
Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on casc…
Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal Intervention
Zhe Xu, Zhicai Wang, Junkang Wu +2
Large Vision-Language Models (LVLMs) often suffer from object hallucination, making erroneous judgments about the presence of objects in images. We propose this primar- ily stems f…
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
Xingjian Zhao, Zhe Xu, Qinyuan Cheng +20
Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits exp…
ProCut: LLM Prompt Compression via Attribution Estimation
Zhentao Xu, Fengyi Li, Albert Chen +1
In large-scale industrial LLM systems, prompt templates often expand to thousands of tokens as teams iteratively incorporate sections such as task instructions, few-shot examples,…
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Shiduo Zhang, Zhe Xu, Peiju Liu +8
General-purposed embodied agents are designed to understand the users' natural instructions or intentions and act precisely to complete universal tasks. Recently, methods based on…
LongSafety: Enhance Safety for Long-Context LLMs
Mianqiu Huang, Xiaoran Liu, Shaojun Zhou +11
Recent advancements in model architectures and length extrapolation techniques have significantly extended the context length of large language models (LLMs), paving the way for th…