activity
20242026
collaborators

8 papers

cs.CL2026

Clustered Self-Assessment: A Simple yet Effective Method for Uncertainty Quantification in Large Language Models

Qi Cao, Takeshi Kojima, Andrew Gambardella +3

Large language models (LLMs) demonstrate remarkable performance across diverse tasks, but they often generate responses that appear plausible while being factually incorrect. This…

cs.CL2026

Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models

Qi Cao, Andrew Gambardella, Takeshi Kojima +2

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks. However, the truthfulness of their outputs is not guaranteed, and their tendency toward…

cs.LG2026

Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment

Jingyuan Feng, Andrew Gambardella, Gouki Minegishi +3

Current safety alignment methods encode safe behavior implicitly within model parameters, creating a fundamental opacity: we cannot easily inspect why a model refuses a request, no…

cs.CL2025

Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?

Fumiya Uchiyama, Takeshi Kojima, Andrew Gambardella +3

Recent large language models (LLMs) have demonstrated remarkable generalization abilities in mathematics and logical reasoning tasks. Prior research indicates that LLMs pre-trained…

cs.CL2025

Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning

Shota Takashiro, Takeshi Kojima, Andrew Gambardella +3

As large language models (LLMs) are applied across diverse domains, the ability to selectively unlearn specific information is becoming increasingly essential. For instance, LLMs a…

cs.CL2025

Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar

Andrew Gambardella, Takeshi Kojima, Yusuke Iwasawa +1

Typical methods for evaluating the performance of language models evaluate their ability to answer questions accurately. These evaluation metrics are acceptable for determining the…