7 papers
DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding
Yilin Wang, Haochen Shi, Guanyu Chen +4
Food segmentation is essential for applications such as intelligent catering, dietary assessment, and recommendation. However, existing benchmarks fail to capture the complexity of…
Regulation, Power, and the Compliance 1 Paradox: A Longitudinal Study of Smart Homes
Wael Albayaydh, Ivan Flechais, Rui Zhao
Smart home technologies are becoming increasingly embedded in domestic environments, yet their implications for power, privacy, and inequality remain insufficiently understood, par…
Modeling Object Attention in Mobile AR for Intrinsic Cognitive Security
Shane Dirksen, Radha Kumaran, You-Jin Kim +2
We study attention in mobile Augmented Reality (AR) using object recall as a proxy outcome. We observe that the ability to recall an object (physical or virtual) that was encounter…
UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics
Xi Chen, Zhifei Zhang, He Zhang +10
We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles:…
FilterPrompt: A Simple yet Efficient Approach to Guide Image Appearance Transfer in Diffusion Models
Xi Wang, Yichen Peng, Heng Fang +4
In controllable generation tasks, flexibly manipulating the generated images to attain a desired appearance or structure based on a single input image cue remains a critical and lo…
FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
Hang Hua, Qing Liu, Lingzhi Zhang +5
The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, inclu…