5 papers
PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
Yantao Li, Qiang Hui, Chenyang Yan +8
Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness an…
Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization via Operator Trees
Xiaoyang Liu, Zineng Dong, Yifan Bai +3
Statement autoformalization acts as a critical bridge between human mathematics and formal mathematics by translating natural language problems into formal language. While prior wo…
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
Yifan Bai, Xiaoyang Liu, Zihao Mou +7
As large language models (LLMs) are increasingly deployed for software engineering, constructing high-quality benchmarks is crucial for evaluating not just the functional correctne…
MediaClaw: Multimodal Intelligent-Agent Platform Technical Report
Shaoan Zhao, Huanlin Gao, Qiang Hui +9
MediaClaw is a multimodal agent platform built on the OpenClaw ecosystem. Its core design follows a three-layer architecture of unified abstraction, pluginized extension, and workf…
Vision-Language Models Can Self-Improve Reasoning via Reflection
Kanzhi Cheng, Yantao Li, Fangzhi Xu +3
Chain-of-thought (CoT) has proven to improve the reasoning capability of large language models (LLMs). However, due to the complexity of multimodal scenarios and the difficulty in…