activity
20232026
most citedAre Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions

3 citations · 13 across the 81 of their papers we have counts for

collaborators
Showing cs.CLShow all

43 papers · 1 filter

cs.CL2026

MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang +5

This paper presents Murano, an open source framework for designing, running, and reproducing mechanistic interpretability studies of large language models, intended for researchers…

cs.CL2026

To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias

Federico Marcuzzi, Xuefei Ning, Roy Schwartz +1

As Large Language Models are increasingly deployed in critical applications, robustly evaluating their social biases is paramount. However, the current literature suffers from wide…

cs.CL2026

Judgment-Grounded Expansion for Peer Review Generation

Sheng Lu, Lizhen Qu, Iryna Gurevych

Automatic review generation is a promising direction for accelerating scientific progress. While most work adopts an end-to-end setup, its fully automated nature makes it less suit…

cs.CL2026

From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent

Haishuo Fang, Yue Feng, Iryna Gurevych

Large language models (LLMs) have shown promise in automating scientific peer review. However, existing approaches often struggle to generate in-depth reviews supported by concrete…

cs.CL2026

ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning

Vladislav Smirnov, Chieu Nguyen, Sergey Senichev +14

Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language model (LLM) reasoning by allocating additional compute during inference, e.g., via m…

cs.CL2026

Contextualized Prompting For Stance Detection On Social Media

Tilman Beck, Shakib Yazdani, Simon Kruschinski +2

Stance detection on social media is challenging due to short, noisy, and context-dependent language. While large language models (LLMs) show zero-shot generalization, they are typi…