9 papers
Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients
Yan Jiao, Pin-Han Ho, Limei Peng
Learned representations in intelligent sensing systems are often evaluated by reconstruction fidelity or downstream prediction accuracy, but these criteria do not specify which lat…
Adaptive Inference-Time Scaling via Early-Step Latent Verification for Image Editing
Yue Yu, Yang Jiao, Jiayu Wang +2
Instruction-based image editing has made notable progress with recent advances in generative models. However, the quality of the edited result is still influenced by the randomly s…
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
Yian Li, Yang Jiao, Bin Zhu +4
Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language…
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
Feng Han, Yang Jiao, Shaoxiang Chen +3
The field of controllable image generation has seen significant advancements, with various architectures improving generation layout consistency with control signals. However, cont…
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Jiacheng Zhang, Yang Jiao, Shaoxiang Chen +5
Recently, Multimodal Large Language Models (MLLMs) have made significant progress in the video comprehension field. Despite remarkable content reasoning and instruction following c…
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
Jiayu Wang, Yang Jiao, Yue Yu +4
Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for…