55 citations · 121 across the 63 of their papers we have counts for
93 papers
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Shenxi Wu, Yuhong Liu, Haosong Zhang +8
Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchm…
WorldReward: Reward Modeling for Camera-Conditioned World Models
Yibin Wang, Zehan Wang, Junshu Tang +13
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics re…
Intern-S2-Preview: Scientific Agentic Foundation Model
Lei Bai, Jiaqi Cao, Chiyu Chen +121
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sus…
HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models
Jiazi Bu, Pengyang Ling, Yujie Zhou +10
Text-Image-to-Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text-to-video (T2V) and image-to-video (I2V) generation. Given…
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games
Shengyuan Ding, Xilin Wei, Xinyu Fang +4
Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observations that are no longer visible. However, existing benchmarks e…
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning
Penghui Yang, Long Xing, Xiaoyi Dong +10
Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Curren…