activity
20202026
most citedStarCoder 2 and The Stack v2: The Next Generation

67 citations · 97 across the 39 of their papers we have counts for

collaborators

46 papers

cs.CL2026

Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering

Li Zheng, Xin Zhang, Shuyi He +5

Accurate comprehension and controllable generation of emotion and rhetoric are pivotal for enhancing the reasoning capabilities of large language models (LLMs). Existing studies mo…

cs.CL2026

How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts

Minh-Vuong Nguyen, Fatemeh Shiri, Zhuang Li +1

Large Language Models (LLMs) are increasingly being explored for clinical question answering and decision support, yet safe deployment critically requires reliable handling of pati…

cs.LG2026

Evidence-based Distributional Alignment for Large Language Models

Viet-Thanh Pham, Lizhen Qu, Zhuang Li +1

Distributional alignment enables large language models (LLMs) to predict how a target population distributes its responses across answer options, rather than collapsing disagreemen…

cs.CV2025

AutoPP: Towards Automated Product Poster Generation and Optimization

Jiahao Fan, Yuxin Qin, Wei Feng +10

Product posters blend striking visuals with informative text to highlight the product and capture customer attention. However, crafting appealing posters and manually optimizing th…

cs.CL2025

ARQUSUMM: Argument-aware Quantitative Summarization of Online Conversations

An Quang Tang, Xiuzhen Zhang, Minh Ngoc Dinh +1

Online conversations have become more prevalent on public discussion platforms (e.g. Reddit). With growing controversial topics, it is desirable to summarize not only diverse argum…

cs.CV2025

Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans

Hongbin Huang, Junwei Li, Tianxin Xie +8

High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present…