4 papers
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
Encheng Su, Jianyu Wu, Chen Tang +9
As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of…
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
Han Deng, Yuan Meng, Shixiang Tang +2
Competitive programming benchmarks are widely used in scenarios such as programming contests and large language model assessments. However, the growing presence of duplicate or hig…
Point2Primitive: CAD Reconstruction from Point Cloud by Direct Primitive Prediction
Xinzhu Ma, Cheng Wang, Chen Tang +5
Recovering CAD models from point clouds requires reconstructing their topology and sketch-based extrusion primitives. A dominant paradigm for representing sketches involves implici…
3DAxisPrompt: Promoting the 3D Grounding and Reasoning in GPT-4o
Dingning Liu, Cheng Wang, Peng Gao +4
Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existi…