◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

S. Aphale

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • last author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG3

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.LGShow all

3 papers · 1 filter

cs.LG2026

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

Ayushman Singh, Siddharth Aphale

Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of-K selection, planning, and critic-guid…

cs.LG2026

SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning

Siddharth Aphale, Ayushman Singh

Sparse-reward reinforcement learning often fails because rollouts from the unassisted evaluation start rarely reach later task stages. Reset curricula address this by starting some…

cs.LG2026

SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR

Siddharth Aphale, Kelly Liu

The standard heuristic of selecting the SFT checkpoint with the highest pass@1 for GRPO can fail when SFT compresses the rollout distribution. For binary rewards, the expected with…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.