papers

Publications (103)

cs.CV2022

Learning Online Multi-Sensor Depth Fusion

Erik Sandström, Martin R. Oswald, Suryansh Kumar +4

Many hand-held or mixed reality devices are used with a single sensor for 3D reconstruction, although they often comprise multiple sensors. Multi-sensor depth fusion is able to sub…

cs.RO2023

TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction

Zhejun Zhang, Alexander Liniger, Dengxin Dai +2

Data-driven simulation has become a favorable way to train and test autonomous driving algorithms. The idea of replacing the actual environment with a learned simulator has also be…

cs.CV2022

Spatio-Temporal Action Detection Under Large Motion

Gurkirt Singh, Vasileios Choutas, Suman Saha +2

Current methods for spatiotemporal action tube detection often extend a bounding box proposal at a given keyframe into a 3D temporal cuboid and pool features from nearby frames. Ho…

cs.CV2023

Segment Anything in High Quality

Lei Ke, Mingqiao Ye, Martin Danelljan +4

The recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being…

cs.AI2019

Deep Object-Centric Policies for Autonomous Driving

Dequan Wang, Coline Devin, Qi-Zhi Cai +2

While learning visuomotor skills in an end-to-end manner is appealing, deep neural networks are often uninterpretable and fail in surprising ways. For robotics tasks, such as auton…

cs.CV2019

Disentangling Propagation and Generation for Video Prediction

Hang Gao, Huazhe Xu, Qi-Zhi Cai +3

A dynamic scene has two types of elements: those that move fluidly and can be predicted from previous frames, and those which are disoccluded (exposed) and cannot be extrapolated.…

cs.CR2022

Babylon: Reusing Bitcoin Mining to Enhance Proof-of-Stake Security

Ertem Nusret Tas, David Tse, Fisher Yu +1

Bitcoin is the most secure blockchain in the world, supported by the immense hash power of its Proof-of-Work miners, but consumes huge amount of energy. Proof-of-Stake chains are e…

cs.CV2023

3DPPE: 3D Point Positional Encoding for Multi-Camera 3D Object Detection Transformers

Changyong Shu, JIajun Deng, Fisher Yu +1

Transformer-based methods have swept the benchmarks on 2D and 3D detection on images. Because tokenization before the attention mechanism drops the spatial information, positional…

cs.CV2018

IDK Cascades: Fast Deep Learning by Learning not to Overthink

Xin Wang, Yujia Luo, Daniel Crankshaw +3

Advances in deep learning have led to substantial increases in prediction accuracy but have been accompanied by increases in the cost of rendering predictions. We conjecture that f…

cs.CV2022

Fast Hierarchical Learning for Few-Shot Object Detection

Yihang She, Goutam Bhat, Martin Danelljan +1

Transfer learning based approaches have recently achieved promising results on the few-shot detection task. These approaches however suffer from ``catastrophic forgetting'' issue d…

cs.CV2023

Video Task Decathlon: Unifying Image and Video Tasks in Autonomous Driving

Thomas E. Huang, Yifan Liu, Luc Van Gool +1

Performing multiple heterogeneous visual tasks in dynamic scenes is a hallmark of human perception capability. Despite remarkable progress in image and video recognition via repres…

cs.CV2023

OVTrack: Open-Vocabulary Multiple Object Tracking

Siyuan Li, Tobias Fischer, Lei Ke +3

The ability to recognize, localize and track dynamic objects in a scene is fundamental to many real-world applications, such as self-driving and robotic systems. Yet, traditional m…

cs.CV2023

Scaling Vision Transformers to 22 Billion Parameters

Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39

The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…

cs.CV2022

CC-3DT: Panoramic 3D Object Tracking via Cross-Camera Fusion

Tobias Fischer, Yung-Hsu Yang, Suryansh Kumar +2

To track the 3D locations and trajectories of the other traffic participants at any given time, modern autonomous vehicles are equipped with multiple cameras that cover the vehicle…

cs.CV2020

BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

Fisher Yu, Haofeng Chen, Xin Wang +5

Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving. Re…

cs.CV2024

Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Appearance Graphs

Mattia Segu, Luigi Piccinelli, Siyuan Li +3

The supervision of state-of-the-art multiple object tracking (MOT) methods requires enormous annotation efforts to provide bounding boxes for all frames of all videos, and instance…

cs.CV2016

LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop

Fisher Yu, Ari Seff, Yinda Zhang +3

While there has been remarkable progress in the performance of visual recognition algorithms, the state-of-the-art models tend to be exceptionally data-hungry. Large labeled traini…

cs.CR2025

Bitcoin-Enhanced Proof-of-Stake Security: Possibilities and Impossibilities

Ertem Nusret Tas, David Tse, Fangyu Gai +3

Bitcoin is the most secure blockchain in the world, supported by the immense hash power of its Proof-of-Work miners. Proof-of-Stake chains are energy-efficient, have fast finality…

cs.CL2025

Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs

Yehui Tang, Yichun Yin, Yaoyuan Wang +71

Sparse large language models (LLMs) with Mixture of Experts (MoE) and close to a trillion parameters are dominating the realm of most capable language models. However, the massive…

cs.RO2024

DexDribbler: Learning Dexterous Soccer Manipulation via Dynamic Supervision

Yutong Hu, Kehan Wen, Fisher Yu

Learning dexterous locomotion policy for legged robots is becoming increasingly popular due to its ability to handle diverse terrains and resemble intelligent behaviors. However, j…

cs.CV2023

QDTrack: Quasi-Dense Similarity Learning for Appearance-Only Multiple Object Tracking

Tobias Fischer, Thomas E. Huang, Jiangmiao Pang +4

Similarity learning has been recognized as a crucial step for object tracking. However, existing multiple object tracking methods only use sparse ground truth matching as the train…

cs.CV2018

SkipNet: Learning Dynamic Routing in Convolutional Networks

Xin Wang, Fisher Yu, Zi-Yi Dou +2

While deeper convolutional networks are needed to achieve maximum accuracy in visual perception tasks, for many inputs shallower networks are sufficient. We exploit this observatio…

cs.CV2022

Transforming Model Prediction for Tracking

Christoph Mayer, Martin Danelljan, Goutam Bhat +4

Optimization based tracking methods have been widely successful by integrating a target model prediction module, providing effective global reasoning by minimizing an objective fun…

cs.CV2024

Strategic Preys Make Acute Predators: Enhancing Camouflaged Object Detectors by Generating Camouflaged Objects

Chunming He, Kai Li, Yachao Zhang +5

Camouflaged object detection (COD) is the challenging task of identifying camouflaged objects visually blended into surroundings. Albeit achieving remarkable success, existing COD…

cs.CV2021

Prototypical Cross-Attention Networks for Multiple Object Tracking and Segmentation

Lei Ke, Xia Li, Martin Danelljan +3

Multiple object tracking and segmentation requires detecting, tracking, and segmenting objects belonging to a set of given classes. Most approaches only exploit the temporal dimens…

cs.CV2023

Uncertainty-Driven Dense Two-View Structure from Motion

Weirong Chen, Suryansh Kumar, Fisher Yu

This work introduces an effective and practical solution to the dense two-view structure from motion (SfM) problem. One vital question addressed is how to mindfully use per-pixel o…

cs.CV2021

Warp Consistency for Unsupervised Learning of Dense Correspondences

Prune Truong, Martin Danelljan, Fisher Yu +1

The key challenge in learning dense correspondences lies in the lack of ground-truth matches for real image pairs. While photometric consistency losses provide unsupervised alterna…

cs.CV2016

Scribbler: Controlling Deep Image Synthesis with Sketch and Color

Patsorn Sangkloy, Jingwan Lu, Chen Fang +2

Recently, there have been several promising methods to generate realistic imagery from deep convolutional networks. These methods sidestep the traditional computer graphics renderi…

cs.CV2015

3D ShapeNets: A Deep Representation for Volumetric Shapes

Zhirong Wu, Shuran Song, Aditya Khosla +4

3D shape is a crucial but heavily underutilized cue in today's computer vision systems, mostly due to the lack of a good generic shape representation. With the recent availability…

cs.CV2016

FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation

Judy Hoffman, Dequan Wang, Fisher Yu +1

Fully convolutional models for dense prediction have proven successful for a wide range of visual tasks. Such models perform well in a supervised setting, but performance can be su…

cs.CV2023

COOLer: Class-Incremental Learning for Appearance-Based Multiple Object Tracking

Zhizheng Liu, Mattia Segu, Fisher Yu

Continual learning allows a model to learn multiple tasks sequentially while retaining the old knowledge without the training data of the preceding tasks. This paper extends the sc…

cs.CV2023

Cascade-DETR: Delving into High-Quality Universal Object Detection

Mingqiao Ye, Lei Ke, Siyuan Li +4

Object localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not comp…

cs.CV2022

Generative Cooperative Learning for Unsupervised Video Anomaly Detection

Muhammad Zaigham Zaheer, Arif Mahmood, Muhammad Haris Khan +3

Video anomaly detection is well investigated in weakly-supervised and one-class classification (OCC) settings. However, unsupervised video anomaly detection methods are quite spars…

cs.CV2024

SM-Net: Joint Learning of Semantic Segmentation and Stereo Matching for Autonomous Driving

Zhiyuan Wu, Yi Feng, Chuang-Wei Liu +3

Semantic segmentation and stereo matching are two essential components of 3D environmental perception systems for autonomous driving. Nevertheless, conventional approaches often ad…

cs.CV2022

TACS: Taxonomy Adaptive Cross-Domain Semantic Segmentation

Rui Gong, Martin Danelljan, Dengxin Dai +4

Traditional domain adaptive semantic segmentation addresses the task of adapting a model to a novel target domain under limited or no additional supervision. While tackling the inp…

cs.CV2024

Multi-modal NeRF Self-Supervision for LiDAR Semantic Segmentation

Xavier Timoneda, Markus Herb, Fabian Duerr +2

LiDAR Semantic Segmentation is a fundamental task in autonomous driving perception consisting of associating each LiDAR point to a semantic label. Fully-supervised models have wide…

cs.CV2020

Frustratingly Simple Few-Shot Object Detection

Xin Wang, Thomas E. Huang, Trevor Darrell +2

Detecting rare objects from a few examples is an emerging problem. Prior works show meta-learning is a promising approach. But, fine-tuning techniques have drawn scant attention. W…

cs.CV2023

Dense Prediction with Attentive Feature Aggregation

Yung-Hsu Yang, Thomas E. Huang, Min Sun +3

Aggregating information from features across different layers is an essential operation for dense prediction models. Despite its limited expressiveness, feature concatenation domin…

cs.CV2020

Task-Aware Feature Generation for Zero-Shot Compositional Learning

Xin Wang, Fisher Yu, Trevor Darrell +1

Visual concepts (e.g., red apple, big elephant) are often semantically compositional and each element of the compositions can be reused to construct novel concepts (e.g., red eleph…

cs.CV2023

BiBench: Benchmarking and Analyzing Network Binarization

Haotong Qin, Mingyuan Zhang, Yifu Ding +5

Network binarization emerges as one of the most promising compression approaches offering extraordinary computation and memory savings by minimizing the bit-width. However, recent…

cs.CV2022

Composite Learning for Robust and Effective Dense Predictions

Menelaos Kanakis, Thomas E. Huang, David Bruggemann +2

Multi-task learning promises better model generalization on a target task by jointly optimizing it with an auxiliary task. However, the current practice requires additional labelin…

cs.CV2022

Video Mask Transfiner for High-Quality Video Instance Segmentation

Lei Ke, Henghui Ding, Martin Danelljan +3

While Video Instance Segmentation (VIS) has seen rapid progress, current approaches struggle to predict high-quality masks with accurate boundary details. Moreover, the predicted s…

cs.CV2023

Unifying Flow, Stereo and Depth Estimation

Haofei Xu, Jing Zhang, Jianfei Cai +4

We present a unified formulation and model for three motion and 3D perception tasks: optical flow, rectified stereo matching and unrectified stereo depth estimation from posed imag…

cs.CV2019

Deep Mixture of Experts via Shallow Embedding

Xin Wang, Fisher Yu, Lisa Dunlap +5

Larger networks generally have greater representational power at the cost of increased computational complexity. Sparsifying such networks has been an active area of research but h…

cs.CV2021

Robust Object Detection via Instance-Level Temporal Cycle Confusion

Xin Wang, Thomas E. Huang, Benlin Liu +4

Building reliable object detectors that are robust to domain shifts, such as various changes in context, viewpoint, and object appearances, is critical for real-world applications.…

cs.CV2018

TextureGAN: Controlling Deep Image Synthesis with Texture Patches

Wenqi Xian, Patsorn Sangkloy, Varun Agrawal +5

In this paper, we investigate deep image synthesis guided by sketch, color, and texture. Previous image synthesis methods can be controlled by sketch and color strokes but we are t…

cs.CV2023

DARTH: Holistic Test-time Adaptation for Multiple Object Tracking

Mattia Segu, Bernt Schiele, Fisher Yu

Multiple object tracking (MOT) is a fundamental component of perception systems for autonomous driving, and its robustness to unseen conditions is a requirement to avoid life-criti…

cs.CV2024

Distilling ODE Solvers of Diffusion Models into Smaller Steps

Sanghwan Kim, Hao Tang, Fisher Yu

Abstract Diffusion models have recently gained prominence as a novel category of generative models. Despite their success, these models face a notable drawback in terms of slow sam…

cs.CV2016

Multi-Scale Context Aggregation by Dilated Convolutions

Fisher Yu, Vladlen Koltun

State-of-the-art models for semantic segmentation are based on adaptations of convolutional networks that had originally been designed for image classification. However, dense pred…

cs.CV2021

End-to-End Urban Driving by Imitating a Reinforcement Learning Coach

Zhejun Zhang, Alexander Liniger, Dengxin Dai +2

End-to-end approaches to autonomous driving commonly rely on expert demonstrations. Although humans are good drivers, they are not good coaches for end-to-end algorithms that deman…

cs.GR2015

ShapeNet: An Information-Rich 3D Model Repository

Angel X. Chang, Thomas Funkhouser, Leonidas Guibas +10

We present ShapeNet: a richly-annotated, large-scale repository of shapes represented by 3D CAD models of objects. ShapeNet contains 3D models from a multitude of semantic categori…

cs.CV2023

Real-Time Motion Prediction via Heterogeneous Polyline Transformer with Relative Pose Encoding

Zhejun Zhang, Alexander Liniger, Christos Sakaridis +2

The real-world deployment of an autonomous driving system requires its components to run on-board and in real-time, including the motion prediction module that predicts the future…

cs.CV2024

Matching Anything by Segmenting Anything

Siyuan Li, Lei Ke, Martin Danelljan +4

The robust association of the same objects across video frames in complex scenes is crucial for many applications, especially Multiple Object Tracking (MOT). Current methods predom…

cs.CV2023

Dual Aggregation Transformer for Image Super-Resolution

Zheng Chen, Yulun Zhang, Jinjin Gu +3

Transformer has recently gained considerable popularity in low-level vision tasks, including image super-resolution (SR). These networks utilize self-attention along different dime…

cs.CV2022

SAGA: Stochastic Whole-Body Grasping with Contact

Yan Wu, Jiahao Wang, Yan Zhang +4

The synthesis of human grasping has numerous applications including AR/VR, video games and robotics. While methods have been proposed to generate realistic hand-object interaction…

cs.CV2022

SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation

Tao Sun, Mattia Segu, Janis Postels +5

Adapting to a continuously evolving environment is a safety-critical challenge inevitably faced by all autonomous driving systems. Existing image and video driving datasets, howeve…

cs.CV2025

Condition-Invariant Semantic Segmentation

Christos Sakaridis, David Bruggemann, Fisher Yu +1

Adaptation of semantic segmentation networks to different visual conditions is vital for robust perception in autonomous cars and robots. However, previous work has shown that most…

cs.CV2022

Uncertainty Guided Policy for Active Robotic 3D Reconstruction using Neural Radiance Fields

Soomin Lee, Le Chen, Jiahao Wang +3

In this paper, we tackle the problem of active robotic 3D reconstruction of an object. In particular, we study how a mobile robot with an arm-held camera can select a favorable num…

cs.CV2023

How To Not Train Your Dragon: Training-free Embodied Object Goal Navigation with Semantic Frontiers

Junting Chen, Guohao Li, Suryansh Kumar +2

Object goal navigation is an important problem in Embodied AI that involves guiding the agent to navigate to an instance of the object category in an unknown environment -- typical…

cs.CV2017

Dilated Residual Networks

Fisher Yu, Vladlen Koltun, Thomas Funkhouser

Convolutional networks for image classification progressively reduce resolution until the image is represented by tiny feature maps in which the spatial structure of the scene is n…

cs.CV2023

iDisc: Internal Discretization for Monocular Depth Estimation

Luigi Piccinelli, Christos Sakaridis, Fisher Yu

Monocular depth estimation is fundamental for 3D scene understanding and downstream applications. However, even under the supervised setup, it is still challenging and ill-posed du…

cs.CV2024

UniDepth: Universal Monocular Metric Depth Estimation

Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis +4

Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling. However, the remarkable accuracy of recent MMDE methods is c…

cs.CV2019

Few-shot Object Detection via Feature Reweighting

Bingyi Kang, Zhuang Liu, Xin Wang +3

Conventional training of a deep CNN based object detector demands a large number of bounding box annotations, which may be unavailable for rare categories. In this work we develop…

cs.CV2024

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

Yiming Li, Sihang Li, Xinhao Liu +11

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D sce…

cs.CV2023

Maskomaly:Zero-Shot Mask Anomaly Segmentation

Jan Ackermann, Christos Sakaridis, Fisher Yu

We present a simple and practical framework for anomaly segmentation called Maskomaly. It builds upon mask-based standard semantic segmentation networks by adding a simple inferenc…

cs.CV2023

Segment Anything Meets Point Tracking

Frano Rajič, Lei Ke, Yu-Wing Tai +3

The Segment Anything Model (SAM) has established itself as a powerful zero-shot image segmentation model, enabled by efficient point-centric annotation and prompt-based models. Whi…

cs.CV2016

Semantic Scene Completion from a Single Depth Image

Shuran Song, Fisher Yu, Andy Zeng +3

This paper focuses on semantic scene completion, a task for producing a complete 3D voxel representation of volumetric occupancy and semantic labels for a scene from a single-view…

cs.CV2019

Hierarchical Discrete Distribution Decomposition for Match Density Estimation

Zhichao Yin, Trevor Darrell, Fisher Yu

Explicit representations of the global match distributions of pixel-wise correspondences between pairs of images are desirable for uncertainty estimation and downstream application…

cs.CR2023

Interchain Timestamping for Mesh Security

Ertem Nusret Tas, Runchao Han, David Tse +2

Fourteen years after the invention of Bitcoin, there has been a proliferation of many permissionless blockchains. Each such chain provides a public ledger that can be written to an…

eess.IV2021

Deep Reparametrization of Multi-Frame Super-Resolution and Denoising

Goutam Bhat, Martin Danelljan, Fisher Yu +2

We propose a deep reparametrization of the maximum a posteriori formulation commonly employed in multi-frame image restoration tasks. Our approach is derived by introducing a learn…

cs.CV2023

Probabilistic Warp Consistency for Weakly-Supervised Semantic Correspondences

Prune Truong, Martin Danelljan, Fisher Yu +1

We propose Probabilistic Warp Consistency, a weakly-supervised learning objective for semantic matching. Our approach directly supervises the dense matching scores predicted by the…

cs.CV2019

Joint Monocular 3D Vehicle Detection and Tracking

Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang +5

Vehicle 3D extents and trajectories are critical cues for predicting the future location of vehicles and planning future agent ego-motion based on those predictions. In this paper,…

cs.CR2021

Accountability and Forensics in Blockchains: XDC Consensus Engine DPoS 2.0

Gerui Wang, Jerome Wang, Liam Lai +1

This document introduces XinFin DPoS 2.0, the proposed next generation decentralized consensus engine for the XinFin XDC Network. Built upon the most advanced BFT consensus protoco…

cs.RO2023

A Multiplicative Value Function for Safe and Efficient Reinforcement Learning

Nick Bührer, Zhejun Zhang, Alexander Liniger +2

An emerging field of sequential decision problems is safe Reinforcement Learning (RL), where the objective is to maximize the reward while obeying safety constraints. Being able to…

cs.CV2021

Exploring Cross-Image Pixel Contrast for Semantic Segmentation

Wenguan Wang, Tianfei Zhou, Fisher Yu +3

Current semantic segmentation methods focus only on mining "local" context, i.e., dependencies between pixels within individual images, by context-aggregation modules (e.g., dilate…

cs.CV2021

Quasi-Dense Similarity Learning for Multiple Object Tracking

Jiangmiao Pang, Linlu Qiu, Xia Li +4

Similarity learning has been recognized as a crucial step for object tracking. However, existing multiple object tracking methods only use sparse ground truth matching as the train…

cs.CL2025

Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs

Yichun Yin, Wenyong Huang, Kaikai Song +49

We present Pangu Ultra, a Large Language Model (LLM) with 135 billion parameters and dense Transformer modules trained on Ascend Neural Processing Units (NPUs). Although the field…

cs.CV2024

HiT-SR: Hierarchical Transformer for Efficient Image Super-Resolution

Xiang Zhang, Yulun Zhang, Fisher Yu

Transformers have exhibited promising performance in computer vision tasks including image super-resolution (SR). However, popular transformer-based SR methods often employ window…

cs.CV2023

MolGrapher: Graph-based Visual Recognition of Chemical Structures

Lucas Morin, Martin Danelljan, Maria Isabel Agea +5

The automatic analysis of chemical literature has immense potential to accelerate the discovery of new materials and drugs. Much of the critical information in patent documents and…

cs.CV2021

Mask Transfiner for High-Quality Instance Segmentation

Lei Ke, Martin Danelljan, Xia Li +3

Two-stage and query-based instance segmentation methods have achieved remarkable results. However, their segmented masks are still very coarse. In this paper, we present Mask Trans…

cs.CL2025

Pangu Light: Weight Re-Initialization for Pruning and Accelerating LLMs

Hanting Chen, Jiarui Qin, Jialong Guo +15

Large Language Models (LLMs) deliver state-of-the-art capabilities across numerous tasks, but their immense size and inference costs pose significant computational challenges for p…

cs.CV2024

Gaussian Grouping: Segment and Edit Anything in 3D Scenes

Mingqiao Ye, Martin Danelljan, Fisher Yu +1

The recent Gaussian Splatting achieves high-quality and real-time novel-view synthesis of the 3D scenes. However, it is solely concentrated on the appearance and geometry modeling,…

cs.CV2021

Normalizing Flow as a Flexible Fidelity Objective for Photo-Realistic Super-resolution

Andreas Lugmayr, Martin Danelljan, Fisher Yu +2

Super-resolution is an ill-posed problem, where a ground-truth high-resolution image represents only one possibility in the space of plausible solutions. Yet, the dominant paradigm…

cs.CV2023

R3D3: Dense 3D Reconstruction of Dynamic Scenes from Multiple Cameras

Aron Schmied, Tobias Fischer, Martin Danelljan +2

Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera s…

cs.CV2022

LiDAR Snowfall Simulation for Robust 3D Object Detection

Martin Hahner, Christos Sakaridis, Mario Bijelic +4

3D object detection is a central task for applications such as autonomous driving, in which the system needs to localize and classify surrounding traffic agents, even in the presen…

cs.CR2018

Characterizing Adversarial Examples Based on Spatial Consistency Information for Semantic Segmentation

Chaowei Xiao, Ruizhi Deng, Bo Li +3

Deep Neural Networks (DNNs) have been widely applied in various recognition tasks. However, recently DNNs have been shown to be vulnerable against adversarial examples, which can m…

cs.RO2024

ICGNet: A Unified Approach for Instance-Centric Grasping

René Zurbrügg, Yifan Liu, Francis Engelmann +4

Accurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of…

cs.CV2022

Tracking Every Thing in the Wild

Siyuan Li, Martin Danelljan, Henghui Ding +2

Current multi-category Multiple Object Tracking (MOT) metrics use class labels to group tracking results for per-class evaluation. Similarly, MOT methods typically only associate o…

cs.CV2026

DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving

Yung-Hsu Yang, Luigi Piccinelli, Siyuan Li +8

DVPSFormer is an online architecture that jointly estimates metric depth, semantic segmentation, and instance trajectories for autonomous driving by using explicit scene discretiza…

#depth-aware segmentation#video panoptic segmentation#online perception#autonomous driving
cs.CV2019

TAFE-Net: Task-Aware Feature Embeddings for Low Shot Learning

Xin Wang, Fisher Yu, Ruth Wang +2

Learning good feature embeddings for images often requires substantial training data. As a consequence, in settings where training data is limited (e.g., few-shot and zero-shot lea…

cs.CV2021

Monocular Quasi-Dense 3D Object Tracking

Hou-Ning Hu, Yung-Hsu Yang, Tobias Fischer +3

A reliable and accurate 3D tracking framework is essential for predicting future locations of surrounding objects and planning the observer's actions in numerous applications such…

cs.CV2018

Interactive 3D Modeling with a Generative Adversarial Network

Jerry Liu, Fisher Yu, Thomas Funkhouser

This paper proposes the idea of using a generative adversarial network (GAN) to assist a novice user in designing real-world shapes with a simple interface. The user edits a voxel…

cs.CV2017

End-to-end Learning of Driving Models from Large-scale Video Datasets

Huazhe Xu, Yang Gao, Fisher Yu +1

Robust perception-action models should be learned from training data with diverse visual appearances and realistic behaviors, yet current approaches to deep visuomotor policy learn…

cs.CV2022

Normalization Perturbation: A Simple Domain Generalization Method for Real-World Domain Shifts

Qi Fan, Mattia Segu, Yu-Wing Tai +4

Improving model's generalizability against domain shifts is crucial, especially for safety-critical applications such as autonomous driving. Real-world domain styles can vary subst…

cs.CV2019

Deep Layer Aggregation

Fisher Yu, Dequan Wang, Evan Shelhamer +1

Visual recognition requires rich representations that span levels from low to high, scales from small to large, and resolutions from fine to coarse. Even with the depth of features…

cs.CV2023

Video OWL-ViT: Temporally-consistent open-world localization in video

Georg Heigold, Matthias Minderer, Alexey Gritsenko +5

We present an architecture and a training recipe that adapts pre-trained open-world image models to localization in videos. Understanding the open visual world (without being const…

cs.CV2022

On the Practicality of Deterministic Epistemic Uncertainty

Janis Postels, Mattia Segu, Tao Sun +4

A set of novel approaches for estimating epistemic uncertainty in deep neural networks with a single forward pass has recently emerged as a valid alternative to Bayesian Neural Net…

cs.RO2022

Learning Deep Sensorimotor Policies for Vision-based Autonomous Drone Racing

Jiawei Fu, Yunlong Song, Yan Wu +2

Autonomous drones can operate in remote and unstructured environments, enabling various real-world applications. However, the lack of effective vision-based algorithms has been a s…

cs.CV2022

RePaint: Inpainting using Denoising Diffusion Probabilistic Models

Andreas Lugmayr, Martin Danelljan, Andres Romero +3

Free-form inpainting is the task of adding new content to an image in the regions specified by an arbitrary binary mask. Most existing approaches train for a certain distribution o…

cs.CV2024

MuRF: Multi-Baseline Radiance Fields

Haofei Xu, Anpei Chen, Yuedong Chen +5

We present Multi-Baseline Radiance Fields (MuRF), a general feed-forward approach to solving sparse view synthesis under multiple different baseline settings (small and large basel…