4 papers
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution
Terry Yue Zhuo, Xiaolong Jin, Hange Liu +37
Crowdsourced model evaluation platforms, such as Chatbot Arena, enable real-time evaluation from human perspectives to assess the quality of model responses. In the coding domain,…
FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis
Jan Ondras, Marek Å uppa
Mathematical reasoning requires abstracting symbolic rules from visual patterns -- inferring the infinite from the finite. We investigate whether multimodal AI systems possess this…
LCDC: Bridging Science and Machine Learning for Light Curve Analysis
Daniel Kyselica, Tomáš Hrobár, JiÅà Šilha +2
The characterization and analysis of light curves are vital for understanding the physical and rotational properties of artificial space objects such as satellites, rocket stages,…
RoBo6: Standardized MMT Light Curve Dataset for Rocket Body Classification
Daniel Kyselica, Marek Å uppa, JiÅÃ Å ilha +1
Space debris presents a critical challenge for the sustainability of future space missions, emphasizing the need for robust and standardized identification methods. However, a comp…