4 papers
Logics-Parsing-Omni Technical Report
Xin An, Jingyi Cai, Xiangyang Chen +22
Addressing the challenges of fragmented task definitions and the heterogeneity of unstructured data in multimodal parsing, this paper proposes the Omni Parsing framework. This fram…
Diffusion Models are Efficient Data Generators for Human Mesh Recovery
Yongtao Ge, Wenjia Wang, Yongfan Chen +4
Despite remarkable progress having been made on the problem of 3D human pose and shape estimation (HPS), current state-of-the-art methods rely heavily on either confined indoor moc…
Logics-Parsing Technical Report
Xiangyang Chen, Shuzhao Li, Xiuwen Zhu +7
Recent advances in Large Vision-Language models (LVLM) have spurred significant progress in document parsing task. Compared to traditional pipeline-based methods, end-to-end paradi…
A Physical Coherence Benchmark for Evaluating Video Generation Models via Optical Flow-guided Frame Prediction
Yongfan Chen, Xiuwen Zhu, Tianyu Li
Recent advances in video generation models demonstrate their potential as world simulators, but they often struggle with videos deviating from physical laws, a key concern overlook…