From the 1 of 6 linked papers with an AI index.
6 papers
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
Jiahao Zhao, Junyi Liu, Lifeng Xu +19
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-s…
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
Yigui Feng, Qinglin Wang, Yang Liu +1
The paper introduces Fre-Res, a dual‑track video token compression method for video multimodal large language models that keeps a few high‑fidelity spatial anchor tokens while enco…
S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
Qingxiao Li, Zikai Wang, Qingli Wang +1
We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image generation models, scien…
S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images
Qingxiao Li, Lifeng Xu, QingLi Wang +4
We present S1-VL, a multimodal reasoning model for scientific domains that natively supports two complementary reasoning paradigms: Scientific Reasoning, which relies on structured…
SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated Code
Qinglin Wang, Zhihong Sun, Ruyun Wang +4
Large Language Models (LLMs) can translate natural language requirements into code, yet empirical analyses of representative models reveal that semantic errors-programs that compil…
ChartReasoner: Code-Driven Modality Bridging for Long-Chain Reasoning in Chart Question Answering
Caijun Jia, Nan Xu, Jingxuan Wei +4
Recently, large language models have shown remarkable reasoning capabilities through long-chain reasoning before responding. However, how to extend this capability to visual reason…