13 papers
When Graph-JEPA Learns the Wrong Thing: Diagnosing and Repairing Category-Conditional Collapse
Gollam Rabby, Sören Auer
Joint-embedding predictive architectures are selected almost universally by linear probing and effective rank. We report a case where both read healthily while the representation c…
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
Tawsif Ahmed, Andrej Radonjic, Gollam Rabby
We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song. To the best of our knowledge, there are no open-source high-quality dataset representing popula…
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
Sasi Kiran Gaddipati, Diyana Muhammed, Farhana Keya +2
Autonomous research systems capable of generating complete scientific manuscripts have advanced rapidly, yet robust and realistic evaluation frameworks have failed to keep pace. To…
From Knowledge to Action: Outcomes of the 2025 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry
Aritra Roy, Kevin Shen, Andrew MacBride +350
Large language models (LLMs) are rapidly changing how researchers in materials science and chemistry discover, organize, and act on scientific knowledge. This paper analyzes a broa…
AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
Sasi Kiran Gaddipati, Farhana Keya, Gollam Rabby +1
High-quality scientific review and perspective papers require substantial time and effort, limiting researchers' ability to synthesize emerging knowledge. While Large Language Mode…
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby +6
Speech emotion recognition (SER) systems are constrained by existing datasets that typically cover only 6-10 basic emotions, lack scale and diversity, and face ethical challenges w…