most citedInteractive Interior Design Recommendation via Coarse-to-fine Multimodal Reinforcement Learning

12 citations · 20 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2025

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

Yuefei Chen, Jiang Liu, Xiaodong Lin +1

Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reasoning, and instruction-following tasks…

cs.MM202312 cited

Interactive Interior Design Recommendation via Coarse-to-fine Multimodal Reinforcement Learning

He Zhang, Ying Sun, Weiyu Guo +4

Personalized interior decoration design often incurs high labor costs. Recent efforts in developing intelligent interior design systems have focused on generating textual requireme…

cs.CL2023

Prompt Space Optimizing Few-shot Reasoning Success with Large Language Models

Fobo Shi, Peijun Qing, Dong Yang +5

Prompt engineering is an essential technique for enhancing the abilities of large language models (LLMs) by providing explicit and specific instructions. It enables LLMs to excel i…

cs.CV20231 cited

Towards Language-guided Interactive 3D Generation: LLMs as Layout Interpreter with Generative Feedback

Yiqi Lin, Hao Wu, Ruichen Wang +4

Generating and editing a 3D scene guided by natural language poses a challenge, primarily due to the complexity of specifying the positional relations and volumetric changes within…

cs.CV2023

Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models

Ruichen Wang, Zekang Chen, Chen Chen +3

Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align…

cs.CV20237 cited

Edit Everything: A Text-Guided Generative System for Images Editing

Defeng Xie, Ruichen Wang, Jian Ma +5

We introduce a new generative system called Edit Everything, which can take image and text inputs and produce image outputs. Edit Everything allows users to edit images using simpl…