7 papers
A Forward Simulation-Based Hierarchy of Linearizable Concurrent Objects
Chao Wang, Ruijia Li, Yang Zhou +5
In this paper, we systematically investigate the connection between linearizable objects and forward simulation. We prove that the sets of linearizable objects satisfying wait-free…
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework
Chao Wang, Chunbai Zhang, Yongxiao Tian +2
Visual reasoning refers to the task of solving questions about visual information. Current visual reasoning methods typically employ pre-trained vision-language model (VLM) strateg…
LAYOUTDREAMER: Physics-guided Layout for Text-to-3D Compositional Scene Generation
Yang Zhou, Zongjin He, Qixuan Li +1
Recently, the field of text-guided 3D scene generation has garnered significant attention. High-quality generation that aligns with physical realism and high controllability is cru…
Can Large Language Models Unveil the Mysteries? An Exploration of Their Ability to Unlock Information in Complex Scenarios
Chao Wang, Luning Zhang, Zheng Wang +1
Combining multiple perceptual inputs and performing combinatorial reasoning in complex scenarios is a sophisticated cognitive function in humans. With advancements in multi-modal l…
TPC: Cross-Temporal Prediction Connection for Vision-Language Model Hallucination Reduction
Chao Wang, Weiwei Fu, Yang Zhou
Vision-language models (VLMs) have achieved remarkable advancements, capitalizing on the impressive capabilities of large language models (LLMs) across diverse tasks. Despite this,…
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
Chao Wang, Xuancheng Zhou, Weiwei Fu +1
Large Visual Language Models (LVLMs) integrate visual and linguistic modalities, exhibiting exceptional performance across various multimodal tasks. Nevertheless, LVLMs remain vuln…