8 papers
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
Tuo Liang, Zhe Hu, Disheng Liu +2
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communic…
Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
Disheng Liu, Tuo Liang, Chaoda Song +1
Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existing approaches to exploiting this potentia…
Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion
Yiran Qiao, Yiren Lu, Yunlai Zhou +5
3D asset generation plays a pivotal role in fields such as gaming and virtual reality, enabling the rapid synthesis of high-fidelity 3D objects from a single or multiple images. Bu…
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
Tuo Liang, Zhe Hu, Jing Li +8
Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant challenge for large vision-language mod…
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
Yiren Lu, Yi Du, Disheng Liu +3
Effective embodied exploration requires agents to accumulate and retain spatial knowledge over time. However, existing scene representations, such as discrete scene graphs or stati…
CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data
Disheng Liu, Yiran Qiao, Wuche Liu +5
True intelligence hinges on the ability to uncover and leverage hidden causal relations. Despite significant progress in AI and computer vision (CV), there remains a lack of benchm…