activity
20242026
collaborators

7 papers

cs.CV2026

DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding

Yilin Wang, Haochen Shi, Guanyu Chen +4

Food segmentation is essential for applications such as intelligent catering, dietary assessment, and recommendation. However, existing benchmarks fail to capture the complexity of…

cs.HC2026

Regulation, Power, and the Compliance 1 Paradox: A Longitudinal Study of Smart Homes

Wael Albayaydh, Ivan Flechais, Rui Zhao

Smart home technologies are becoming increasingly embedded in domestic environments, yet their implications for power, privacy, and inequality remain insufficiently understood, par…

cs.HC2025

Modeling Object Attention in Mobile AR for Intrinsic Cognitive Security

Shane Dirksen, Radha Kumaran, You-Jin Kim +2

We study attention in mobile Augmented Reality (AR) using object recall as a proxy outcome. We observe that the ability to recall an object (physical or virtual) that was encounter…

cs.CV2024

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Xi Chen, Zhifei Zhang, He Zhang +10

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles:…

cs.CV2024

FilterPrompt: A Simple yet Efficient Approach to Guide Image Appearance Transfer in Diffusion Models

Xi Wang, Yichen Peng, Heng Fang +4

In controllable generation tasks, flexibly manipulating the generated images to attain a desired appearance or structure based on a single input image cue remains a critical and lo…

cs.CV2024

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

Hang Hua, Qing Liu, Lingzhi Zhang +5

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, inclu…