13 papers
SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery
SciForge Team, Zhangyang Gao, Minghao Fang +10
Scientific work increasingly spans heterogeneous artifacts -- papers, code, datasets, scientific file formats, model outputs, figures, manuscripts, and team decisions -- yet genera…
Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement
Tingyu Li, Le Zhou, Siyuan Li +5
Interleaved thinking, where a unified multimodal model alternates between textual reasoning and visual generation, has shown promise on spatial and physical tasks. However, in comp…
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
Bo Yin, Xiaobin Hu, Chengming Xu +6
Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evi…
PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control
Jingxuan Wei, Xi Bai, Shan Liu +8
Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a f…
NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation
Jinhang Xu, Qiyuan Zhu, Yujun Wu +11
LLM-powered multi-agent systems can now automate the full research pipeline from ideation to paper writing, but a fundamental question remains: automation for whom? Researchers ope…
PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents
Bihui Yu, Xinglong Xu, Junjie Jiang +6
A LaTeX manuscript that compiles without error is not necessarily publication-ready. The resulting PDFs frequently suffer from misplaced floats, overflowing equations, inconsistent…