1 citations · 2 across the 5 of their papers we have counts for
5 papers
Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding
Zhiyuan Zhu, Yixuan Chen, Yiwen Shao +13
Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial rel…
Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular Video
Seonghwa Choi, Moonkyeong Choi, Mingyu Jang +4
Modeling relightable and animatable human avatars from monocular video is a long-standing and challenging task. Recently, Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3D…
CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
Ruoxuan Zhang, Bin Wen, Hongxia Xie +5
Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusi…
MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering
Xinqi Fan, Jingting Li, John See +6
Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial ex…
SMP Challenge: An Overview and Analysis of Social Media Prediction Challenge
Bo Wu, Peiye Liu, Wen-Huang Cheng +5
Social Media Popularity Prediction (SMPP) is a crucial task that involves automatically predicting future popularity values of online posts, leveraging vast amounts of multimodal d…