papers

Publications (29)

cs.CV2023

3D-LLM: Injecting the 3D World into Large Language Models

Yining Hong, Haoyu Zhen, Peihao Chen +4

Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are…

cs.SE2026

Towards Verifiably Safe Tool Use for LLM Agents

Aarya Doshi, Yining Hong, Congying Xu +3

Large language model (LLM)-based AI agents extend LLM capabilities by enabling access to tools such as data sources, APIs, search engines, code sandboxes, and even other agents. Wh…

cs.CV2023

GENOME: GenerativE Neuro-symbOlic visual reasoning by growing and reusing ModulEs

Zhenfang Chen, Rui Sun, Wenjun Liu +2

Recent works have shown that Large Language Models (LLMs) could empower traditional neuro-symbolic models via programming capabilities to translate language into module description…

cs.LG2023

A Minimalist Dataset for Systematic Generalization of Perception, Syntax, and Semantics

Qing Li, Siyuan Huang, Yining Hong +3

Inspired by humans' exceptional ability to master arithmetic and generalize to new problems, we present a new dataset, Handwritten arithmetic with INTegers (HINT), to examine machi…

cs.CV2023

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Junyan Li, Delin Chen, Yining Hong +4

A remarkable ability of human beings resides in compositional reasoning, i.e., the capacity to make "infinite use of finite means". However, current large vision-language foundatio…

cs.CV2021

PTR: A Benchmark for Part-based Conceptual, Relational, and Physical Reasoning

Yining Hong, Li Yi, Joshua B. Tenenbaum +2

A critical aspect of human visual perception is the ability to parse visual scenes into individual objects and further into object parts, forming part-whole hierarchies. Such compo…