activity
20242026
most citedSceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code

2 citations · 3 across the 11 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

5 papers · 2 filters

cs.CV2024

Visual Lexicon: Rich Image Features in Language Space

XuDong Wang, Xingyi Zhou, Alireza Fathi +2

We present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intricate visual details that are of…

cs.CV2024

Language-Guided Image Tokenization for Generation

Kaiwen Zha, Lijun Yu, Alireza Fathi +4

Image tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generatio…

cs.CV2024

Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach

Mathilde Caron, Alireza Fathi, Cordelia Schmid +1

Web-scale visual entity recognition, the task of associating images with their corresponding entities within vast knowledge bases like Wikipedia, presents significant challenges du…

cs.CV2024★ 1 cited

A Generative Approach for Wikipedia-Scale Visual Entity Recognition

Mathilde Caron, Ahmet Iscen, Alireza Fathi +1

In this paper, we address web-scale visual entity recognition, specifically the task of mapping a given query image to one of the 6 million existing entities in Wikipedia. One way…

cs.CV2024★ 2 cited

SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code

Ziniu Hu, Ahmet Iscen, Aashi Jain +5

This paper introduces SceneCraft, a Large Language Model (LLM) Agent converting text descriptions into Blender-executable Python scripts which render complex scenes with up to a hu…