collaborators

8 papers

quant-ph2026

Circuit Depth Compression via Spectral Gap Amplification in Quantum Phase Estimation

Sk Mujaffar Hossain, Satadeep Bhattacharjee

We show that quantum phase estimation (QPE) circuits can be significantly compressed in depth by preprocessing the input operator with a sigmoid spectral filter before estimation.…

cs.LG2026

Exploration of Fast-Slow Latent Recurrence for Train-Short, Test-Long Generalization

Shota Takashiro, Masanori Koyama, Takeru Miyato +3

We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.…

cs.LG2026

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation

Paavo Parmas, Yongmin Kim, Kohsei Matsutani +5

Policy-gradient methods usually optimize expected return, but many real world applications care about distributional properties of returns: tail risk, outlier robustness, or best-o…

cs.LG2026

On Advantage Estimates for Max@K Policy Gradients

Shota Takashiro, Soichiro Nishimori, Paavo Parmas +6

Reinforcement learning with verifiable rewards is widely used for post-training reasoning models, but sparse outcome rewards make exploration difficult. A complementary approach is…

cs.CV2026

CLIP-like Model as a Foundational Density Ratio Estimator

Fumiya Uchiyama, Rintaro Yanagi, Shohei Taniguchi +5

Density ratio estimation is a core concept in statistical machine learning because it provides a unified mechanism for tasks such as importance weighting, divergence estimation, an…

cs.AI2026

RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs

Kohsei Matsutani, Shota Takashiro, Gouki Minegishi +3

Large language models (LLMs) are typically trained by reinforcement learning (RL) with verifiable rewards (RLVR) and supervised fine-tuning (SFT) on reasoning traces to improve the…