activity
20242026
collaborators

7 papers

cs.CV2026

Multi-Object Advertisement Creative Generation

Jialu Gao, Mithun Das Gupta, Qun Li +4

Lifestyle images are photographs that capture environments and objects in everyday settings. In furniture product marketing, advertisers often create lifestyle images containing pr…

cs.HC2026

Proscenium: Exploring Design Spaces of Layered Information Experience on a Large Dual-Layer Transparent Display

Chen Chen, Michel Pahud, David Brown +5

Layering information spaces is a promising strategy to design intuitive and engaging interactive experiences. Although multi-layer displays enable promising interaction techniques…

cs.AI2026

MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents

Tamil Sudaravan Mohan Doss, Michael Xu, Sudha Rao +2

We present MineNPC-Task, a user-authored benchmark and evaluation harness for testing memory-aware, mixed-initiative LLM agents in open-world Minecraft. Rather than relying on synt…

cs.HC2025

Doc To The Future: Infomorphs for Interactive, Multimodal Document Transformation and Generation

Balasaravanan Thoravi Kumaravel

Creating new documents by synthesizing information from existing sources is an important part of knowledge work in many domains. This process often involves gathering content from…

cs.CV2025

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames

Sahithya Ravi, Gabriel Sarch, Vibhav Vineet +2

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relati…

cs.CV2025

Grounding Task Assistance with Multimodal Cues from a Single Demonstration

Gabriel Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi +2

A person's demonstration often serves as a key reference for others learning the same task. However, RGB video, the dominant medium for representing these demonstrations, often fai…