◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Soichiro Nishimori

3 papers hereh-index 489 citations19 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author1

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG3
same name
  • Soichiro Nishimori — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2026

Mitigating Reward Hacking in RLHF via Advantage Sign Robustness

Shinnosuke Ono, Johannes Ackermann, Soichiro Nishimori +2

Reward models (RMs) used in reinforcement learning from human feedback (RLHF) are vulnerable to reward hacking: as the policy maximizes a learned proxy reward, true quality plateau…

cs.LG2025

Recursive Reward Aggregation

Yuting Tang, Yivan Zhang, Johannes Ackermann +3

In reinforcement learning (RL), aligning agent behavior with specific objectives typically requires careful design of the reward function, which can be challenging when the desired…

cs.LG2025

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences

Soichiro Nishimori, Yu-Jie Zhang, Thanawat Lodkaew +1

Optimizing policies based on human preferences is key to aligning language models with human intent. This work focuses on reward modeling, a core component in reinforcement learnin…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.