activity
20212024
most citedVAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer

28 citations · 47 across the 12 of their papers we have counts for

collaborators

12 papers

cs.CV2024

MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning

Yifan Jiang, Jiarui Zhang, Kexuan Sun +5

While multi-modal large language models (MLLMs) have shown significant progress on many popular visual reasoning benchmarks, whether they possess abstract visual reasoning abilitie…

cs.AI2024

SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense

Yifan Jiang, Filip Ilievski, Kaixin Ma

While vertical thinking relies on logical and commonsense reasoning, lateral thinking requires systems to defy commonsense associations and overwrite them through unconventional th…

cs.CV2024

VASE: Object-Centric Appearance and Shape Manipulation of Real Videos

Elia Peruzzo, Vidit Goel, Dejia Xu +5

Recently, several works tackled the video editing task fostered by the success of large-scale text-to-image generative models. However, most of these methods holistically edit the…

cs.CL2023

BRAINTEASER: Lateral Thinking Puzzles for Large Language Models

Yifan Jiang, Filip Ilievski, Kaixin Ma +1

The success of language models has inspired the NLP community to attend to tasks that require implicit and complex reasoning, relying on human-like commonsense mechanisms. While su…

cs.CV20232 cited

Efficient-3DiM: Learning a Generalizable Single-image Novel-view Synthesizer in One Day

Yifan Jiang, Hao Tang, Jen-Hao Rick Chang +3

The task of novel view synthesis aims to generate unseen perspectives of an object or scene from a limited set of input images. Nevertheless, synthesizing novel views from a single…

cs.CV20221 cited

AligNeRF: High-Fidelity Neural Radiance Fields via Alignment-Aware Training

Yifan Jiang, Peter Hedman, Ben Mildenhall +4

Neural Radiance Fields (NeRFs) are a powerful representation for modeling a 3D scene as a continuous function. Though NeRF is able to render complex 3D scenes with view-dependent e…