10 papers
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
Wentao Mo, Yang Liu
Current 3D spatial reasoning methods face a fundamental trade-off: neuro-symbolic 3D (NS3D) concept learners achieve interpretable reasoning through compositional programs but are…
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
Weixing Chen, Zhuoqian Feng, Yang Liu +6
Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense objec…
Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning
Yang Liu, Qianqian Xu, Peisong Wen +2
Recent advances in large vision-language models have expanded video retrieval from simple text-based search to more flexible scenarios, where users may specify the desired result t…
Benign Inputs, Harmful Outputs: Cross-Modal Jailbreaking via Distributed Semantic Recomposition
Yani Wang, Yilong Yang, Yang Liu +3
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in content synthesis and autonomous reasoning. Previous safety guardrails are primarily…
Enhancing LLM Metacognition via Cognitive Pairwise Training
Weitao Li, Hao Zhou, Xuanyu Lei +11
Reinforcement learning with verifiable rewards (RLVR) has become central to LLM reasoning, but its outcome-level rewards can make models more willing to give confident answers when…
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Yicheng Zou, Dongsheng Zhu, Lin Zhu +174
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…