activity
20182026
most citedUnified Vision and Language Prompt Learning

55 citations · 121 across the 63 of their papers we have counts for

collaborators

93 papers

cs.AI2026

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

Shenxi Wu, Yuhong Liu, Haosong Zhang +8

Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchm…

cs.CV2026

WorldReward: Reward Modeling for Camera-Conditioned World Models

Yibin Wang, Zehan Wang, Junshu Tang +13

Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics re…

cs.LG2026

Intern-S2-Preview: Scientific Agentic Foundation Model

Lei Bai, Jiaqi Cao, Chiyu Chen +121

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sus…

cs.CV2026

HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models

Jiazi Bu, Pengyang Ling, Yujie Zhou +10

Text-Image-to-Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text-to-video (T2V) and image-to-video (I2V) generation. Given…

cs.CV2026

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games

Shengyuan Ding, Xilin Wei, Xinyu Fang +4

Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observations that are no longer visible. However, existing benchmarks e…

cs.CV2026

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

Penghui Yang, Long Xing, Xiaoyi Dong +10

Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Curren…