7 papers
Allegory of the Cave: Measurement-Grounded Vision-Language Learning
Kepeng Xu, Li Xu, Gang He +1
Vision-language models typically reason over post-ISP RGB images, although RGB rendering can clip, suppress, or quantize sensor evidence before inference. We study whether groundin…
FireRed-OCR Technical Report
Hao Wu, Haoran Lou, Xinyue Li +19
We present FireRed-OCR, a systematic framework to specialize general VLMs into high-performance OCR models. Large Vision-Language Models (VLMs) have demonstrated impressive general…
Nipping the Drift in the Bud: Retrospective Rectification for Robust Vision-Language Navigation
Gang He, Zhenyang Liu, Kepeng Xu +5
Vision-Language Navigation (VLN) requires embodied agents to interpret natural language instructions and navigate through complex continuous 3D environments. However, the dominant…
Beyond Feature Mapping GAP: Integrating Real HDRTV Priors for Superior SDRTV-to-HDRTV Conversion
Gang He, Kepeng Xu, Li Xu +3
The rise of HDR-WCG display devices has highlighted the need to convert SDRTV to HDRTV, as most video sources are still in SDR. Existing methods primarily focus on designing neural…
An End-to-End Real-World Camera Imaging Pipeline
Kepeng Xu, Zijia Ma, Li Xu +5
Recent advances in neural camera imaging pipelines have demonstrated notable progress. Nevertheless, the real-world imaging pipeline still faces challenges including the lack of jo…
Dual Inverse Degradation Network for Real-World SDRTV-to-HDRTV Conversion
Kepeng Xu, Li Xu, Gang He +4
In this study, we address the emerging necessity of converting Standard Dynamic Range Television (SDRTV) content into High Dynamic Range Television (HDRTV) in light of the limited…