9 papers
SpikeDecoder: Realizing the GPT Architecture with Spiking Neural Networks
Claas Beger, Florian Walter, Alois Knoll
The Transformer architecture is widely regarded as the most powerful tool for natural language processing, but due to a high number of complex operations, it inherently faces the i…
OmniCode: A Benchmark for Evaluating Software Engineering Agents
Atharv Sonwane, Eng-Shen Tu, Wei-Chung Lu +11
LLM-powered coding agents are redefining how real-world software is developed. To drive the research towards better coding agents, we require challenging benchmarks that can rigoro…
-CoT: Prolog-Initialized Chain-of-Thought Prompting for Multi-Hop Question-Answering
Chao Wan, Albert Gong, Mihir Mishra +3
Chain-of-Thought (CoT) prompting significantly enhances large language models' (LLMs) problem-solving capabilities, but still struggles with complex multi-hop questions, often fall…
Bongards at the Boundary of Perception and Reasoning: Programs or Language?
Cassidy Langenfeld, Claas Beger, Gloria Geng +4
Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans…
Do AI Models Perform Human-like Abstract Reasoning Across Modalities?
Claas Beger, Ryan Yi, Shuhao Fu +5
OpenAI's o3-preview reasoning model exceeded human accuracy on the ARC-AGI-1 benchmark, but does that mean state-of-the-art models recognize and reason with the abstractions the be…
Citegeist: Automated Generation of Related Work Analysis on the arXiv Corpus
Claas Beger, Carl-Leander Henneking
Large Language Models provide significant new opportunities for the generation of high-quality written works. However, their employment in the research community is inhibited by th…