works on

From the 1 of 46 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.LGShow all

17 papers · 1 filter

cs.LG2026

Vector Symbolic Policy Gradient

Ryozo Masukawa, Sanggeon Yun, SungHeon Jeong +6

We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to t…

cs.LG2026

Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization

Zhiqi Gao, Albert Ge, Alexander Berenbeim +2

The paper investigates why text‑to‑optimization models struggle to correctly ground problem data, introduces a benchmark (Text2Opt‑Bench) to study this, and proposes a binding‑focu…

cs.LG2026

Interactive Critique-Revision Training for Reliable Structured LLM Generation

Fei Xu Yu, Zuyuan Zhang, Mahdi Imani +2

In structured decision-making workflows such as form filling, compliance checking, and maintenance reporting, LLM outputs must be locally correct, globally consistent, and auditabl…

cs.LG2026

HIPO: Instruction Hierarchy via Constrained Reinforcement Learning

Keru Chen, Jun Luo, Sen Lin +4

Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO…

cs.LG2026

MissionHD: Hyperdimensional Refinement of Distribution-Deficient Reasoning Graphs for Video Anomaly Detection

Sanggeon Yun, Raheeb Hassan, Ryozo Masukawa +2

LLM-generated reasoning graphs, referred to as mission-specific graphs (MSGs), are increasingly used for video anomaly detection (VAD) and recognition (VAR). However, they are typi…

cs.LG2026

-Musketeers: Reinforcement Learning Shapes Collaboration Among Language Models

Ryozo Masukawa, Sanggeon Yun, Hyunwoo Oh +8

Recent progress in reinforcement learning with verifiable rewards (RLVR) shows that small, specialized language models (SLMs) can exhibit structured reasoning without relying on la…