4 papers
Adaptive Prompt Elicitation for Text-to-Image Generation
Xinyi Wen, Lena Hegemann, Xiaofu Jin +2
Aligning text-to-image generation with user intent remains challenging, as users frequently provide ambiguous inputs and struggle with model idiosyncrasies. We propose Adaptive Pro…
Hierarchical Resource Rationality Explains Human Reading Behavior
Yunpeng Bai, Xiaofu Jin, Shengdong Zhao +1
Reading is a pervasive and cognitively demanding activity that underpins modern human culture. It is a prime instance of a class of tasks where eye movements are coordinated for th…
Improving Procedural Skill Explanations via Constrained Generation: A Symbolic-LLM Hybrid Architecture
Rahul Dass, Thomas Bowlin, Zebing Li +2
In procedural skill learning, instructional explanations must convey not just steps, but the causal, goal-directed, and compositional logic behind them. Large language models (LLMs…
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
Chendong Wang, Donglin Bai, Yifan Yang +11
We present \emph{Video-in-the-Loop} (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first \emph{localizing} question-relevant interval(s) with a…