4 papers
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
Encheng Su, Jianyu Wu, Chen Tang +9
As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of…
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
Han Deng, Yuan Meng, Shixiang Tang +2
Competitive programming benchmarks are widely used in scenarios such as programming contests and large language model assessments. However, the growing presence of duplicate or hig…
Point2Primitive: CAD Reconstruction from Point Cloud by Direct Primitive Prediction
Xinzhu Ma, Cheng Wang, Chen Tang +5
Recovering CAD models from point clouds requires reconstructing their topology and sketch-based extrusion primitives. A dominant paradigm for representing sketches involves implici…
GMAI-VL-R1: Harnessing Reinforcement Learning for Multimodal Medical Reasoning
Yanzhou Su, Tianbin Li, Jiyao Liu +15
Recent advances in general medical AI have made significant strides, but existing models often lack the reasoning capabilities needed for complex medical decision-making. This pape…