4 papers
Semantic Glitch: Agency and Artistry in an Autonomous Pixel Cloud
Qing Zhang, Jing Huang, Mingyang Xu +1
While mainstream robotics pursues metric precision and flawless performance, this paper explores the creative potential of a deliberately "lo-fi" approach. We present the "Semantic…
Panel-by-Panel Souls: A Performative Workflow for Expressive Faces in AI-Assisted Manga Creation
Qing Zhang, Jing Huang, Yifei Huang +1
Current text-to-image models struggle to render the nuanced facial expressions required for compelling manga narratives, largely due to the ambiguity of language itself. To bridge…
Look and Talk: Seamless AI Assistant Interaction with Gaze-Triggered Activation
Zhang Qing, Rekimoto Jun
Engaging with AI assistants to gather essential information in a timely manner is becoming increasingly common. Traditional activation methods, like wake words such as Hey Siri, Ok…
GazeLLM: Multimodal LLMs incorporating Human Visual Attention
Jun Rekimoto
Large Language Models (LLMs) are advancing into Multimodal LLMs (MLLMs), capable of processing image, audio, and video as well as text. Combining first-person video, MLLMs show pro…