13 papers
Rosetta: Composable Native Multimodal Pretraining
Xiangyue Liu, Zijian Zhang, Miles Yang +3
Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuou…
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
Xiangyue Liu, Zijian Zhang, Miles Yang +3
Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts. While existing parad…
OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning
Wenxuan Jiang, Zining Fan, Zijian Zhang +6
Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creati…
GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation
Zijian Zhang, Yuqing Jiang, Qian Cheng +6
Vision-language-action (VLA) policies have advanced language-conditioned robotic manipulation by transferring semantic priors from pretrained vision-language models to action gener…
QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding
Shuxiang Cao, Zijian Zhang, Abhishek Agarwal +29
Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for this task, yet no system…
Conformal Margin Risk Minimization: An Envelope Framework for Robust Learning under Label Noise
Yuanjie Shi, Peihong Li, Zijian Zhang +2
Most methods for learning with noisy labels require privileged knowledge such as noise transition matrices, clean subsets or pretrained feature extractors, resources typically unav…