From the 1 of 46 linked papers with an AI index.
17 papers · 1 filter
Vector Symbolic Policy Gradient
Ryozo Masukawa, Sanggeon Yun, SungHeon Jeong +6
We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to t…
Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization
Zhiqi Gao, Albert Ge, Alexander Berenbeim +2
The paper investigates why text‑to‑optimization models struggle to correctly ground problem data, introduces a benchmark (Text2Opt‑Bench) to study this, and proposes a binding‑focu…
Interactive Critique-Revision Training for Reliable Structured LLM Generation
Fei Xu Yu, Zuyuan Zhang, Mahdi Imani +2
In structured decision-making workflows such as form filling, compliance checking, and maintenance reporting, LLM outputs must be locally correct, globally consistent, and auditabl…
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
Keru Chen, Jun Luo, Sen Lin +4
Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO…
MissionHD: Hyperdimensional Refinement of Distribution-Deficient Reasoning Graphs for Video Anomaly Detection
Sanggeon Yun, Raheeb Hassan, Ryozo Masukawa +2
LLM-generated reasoning graphs, referred to as mission-specific graphs (MSGs), are increasingly used for video anomaly detection (VAD) and recognition (VAR). However, they are typi…
-Musketeers: Reinforcement Learning Shapes Collaboration Among Language Models
Ryozo Masukawa, Sanggeon Yun, Hyunwoo Oh +8
Recent progress in reinforcement learning with verifiable rewards (RLVR) shows that small, specialized language models (SLMs) can exhibit structured reasoning without relying on la…