5 papers
Reinforcing Structured Chain-of-Thought for Video Understanding
Peiyao Wang, Haotian Xu, Noranart Vesdapunt +6
Multi-modal Large Language Models (MLLMs) show promise in video understanding. However, their reasoning often suffers from thinking drift and weak temporal comprehension, even when…
PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
Peiyao Wang, Weining Wang, Qi Li
Recent advances in text-to-video generation have achieved impressive perceptual quality, yet generated content often violates fundamental principles of physical plausibility - mani…
XDIP: A Curated X-ray Absorption Spectrum Dataset for Iron-Containing Proteins
Yufeng Wang, Peiyao Wang, Lu Wei +6
Earth-abundant iron is an essential metal in regulating the structure and function of proteins. This study presents the development of a comprehensive X-ray Absorption Spectroscopy…
Spectra-to-Structure and Structure-to-Spectra Inference Across the Periodic Table
Yufeng Wang, Peiyao Wang, Lu Wei +4
X-ray Absorption Spectroscopy (XAS) is a powerful technique for probing local atomic environments, yet its interpretation remains limited by the need for expert-driven analysis, co…
SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
Peiyao Wang, Haibin Ling
Spatial reasoning remains a critical yet underdeveloped capability in existing vision-language models (VLMs), especially for Spatial Visual Question Answering (Spatial VQA) tasks t…