activity
20242026
collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects

Danrui Li, Jiahao Zhang, Bernhard Egger +4

Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly ste…

cs.CV2025

Auto-Vocabulary 3D Object Detection

Haomeng Zhang, Kuan-Chuan Peng, Suhas Lohit +1

Open-vocabulary 3D object detection methods are able to localize 3D boxes of classes unseen during training. Despite the name, existing methods rely on user-specified classes both…

cs.CV2025

WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate

Anoop Cherian, River Doyle, Eyal Ben-Dov +2

Recent large language models (LLMs) are trained on diverse corpora and tasks, leading them to develop complementary strengths. Multi-agent debate (MAD) has emerged as a popular way…

cs.CV2025

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

Xinhao Xiang, Kuan-Chuan Peng, Suhas Lohit +2

3D object detection plays a crucial role in autonomous systems, yet existing methods are limited by closed-set assumptions and struggle to recognize novel objects and their attribu…

cs.CV2025

Programmatic Video Prediction Using Large Language Models

Hao Tang, Kevin Ellis, Suhas Lohit +2

The task of estimating the world model describing the dynamics of a real world process assumes immense importance for anticipating and preparing for future outcomes. For applicatio…

cs.CV2025

FreBIS: Frequency-Based Stratification for Neural Implicit Surface Representations

Naoko Sawada, Pedro Miraldo, Suhas Lohit +2

Neural implicit surface representation techniques are in high demand for advancing technologies in augmented reality/virtual reality, digital twins, autonomous navigation, and many…