collaborators

8 papers

cs.CV2026

OvisOCR2 Technical Report

Shiyin Lu, Yinglun Li, Yu Xia +10

OvisOCR2 is a 0.8 B parameter end‑to‑end model that converts document page images into Markdown, handling text, formulas, tables, and visual regions, and achieves state‑of‑the‑art…

cs.CL2026

TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question Answering

An-Yang Ji, Jun-Peng Jiang, De-Chuan Zhan +1

Large Language Models (LLMs) have advanced Table Question Answering, where most queries can be answered by extracting information or simple aggregation. However, a common class of…

cs.LG2026

Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation

Guo Yu, Wenlin Liu, Yulan Hu +3

On-policy distillation (OPD) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student-generated trajectories and dense token-l…

cs.CV2026

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs

Zhi-Kai Chen, Jun-Peng Jiang, Jun-Jie Tao +2

Users increasingly expect image generation models to quickly adapt to highly diverse and personalized requirements, such as producing images with distinctive styles or characterist…

cs.CV2026

PhysInOne: Visual Physics Learning and Reasoning in One Suite

Siyuan Zhou, Hejun Wang, Hu Cheng +36

We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to mere…

cs.CV2025

Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation

Zhi-Kai Chen, Jun-Peng Jiang, Han-Jia Ye +1

Autoregressive (AR) image generation models are capable of producing high-fidelity images but often suffer from slow inference due to their inherently sequential, token-by-token de…