4 papers
Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
Sangoh Lee, Sangwoo Mo, Wook-Shin Han
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which m…
CADENZA: Compiling Natural-Language Intent into Task-Specific Operator DAGs for Semantic Query Processing
Jaehyun Ha, Yongjoo Park, Wook-Shin Han
Semantic query processing engines (SQPEs) extend relational query processing with semantic operators that are executed via model inference over unstructured data. Optimizing such q…
CADENZA in Action: Breaking the Monolith with Intent-Dependent Plan Spaces for Semantic Queries
Jaehyun Ha, Yongjoo Park, Wook-Shin Han
Semantic query processing engines execute semantic operators, whose behavior is specified by natural-language intents, via model inference over multimodal data. Most existing optim…
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
Sangoh Lee, Sangwoo Mo, Wook-Shin Han
While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring my cup," where the robot must act on one…