1 citations · 1 across the 24 of their papers we have counts for
36 papers
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
Ku Onoda, Paavo Parmas, Manato Yaguchi +1
In policy gradient reinforcement learning, access to a differentiable model enables 1st-order gradient estimation that accelerates learning compared to relying solely on derivative…
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
Kenji Kubo, Shunsuke Kamiya, Masanori Koyama +3
Neural network models with latent recurrent processing, where identical layers are recursively applied to the latent state, have gained attention as promising models for performing…
Understanding Emergent Misalignment via Feature Superposition Geometry
Gouki Minegishi, Hiroki Furuta, Takeshi Kojima +2
Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs. Despite growing empirical evidence, it…
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
Qi Cao, Andrew Gambardella, Takeshi Kojima +2
Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks. However, the truthfulness of their outputs is not guaranteed, and their tendency toward…
Towards High-resolution and Disentangled Reference-based Sketch Colorization
Dingkun Yan, Xinrui Wang, Ru Wang +5
Sketch colorization is a critical task for automating and assisting in the creation of animations and digital illustrations. Previous research identified the primary difficulty as…
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
Jingyuan Feng, Andrew Gambardella, Gouki Minegishi +3
Current safety alignment methods encode safe behavior implicitly within model parameters, creating a fundamental opacity: we cannot easily inspect why a model refuses a request, no…