papers

Publications (106)

cs.CL2023

Decomposing Complex Queries for Tip-of-the-tongue Retrieval

Kevin Lin, Kyle Lo, Joseph E. Gonzalez +1

When re-finding items, users who forget or are uncertain about identifying details often rely on creative strategies for expressing their information needs -- complex queries that…

cs.CY2020

A Berkeley View of Teaching CS at Scale

Kevin Lin

Over the past decade, undergraduate Computer Science (CS) programs across the nation have experienced an explosive growth in enrollment as computational skills have proven increasi…

eess.IV2024

Diffusion and Multi-Domain Adaptation Methods for Eosinophil Segmentation

Kevin Lin, Donald Brown, Sana Syed +1

Eosinophilic Esophagitis (EoE) represents a challenging condition for medical providers today. The cause is currently unknown, the impact on a patient's daily life is significant,…

cs.AI2024

MemGPT: Towards LLMs as Operating Systems

Charles Packer, Sarah Wooders, Kevin Lin +4

Large language models (LLMs) have revolutionized AI, but are constrained by limited context windows, hindering their utility in tasks like extended conversations and document analy…

cs.RO2025

robosuite: A Modular Simulation Framework and Benchmark for Robot Learning

Yuke Zhu, Josiah Wong, Ajay Mandlekar +6

robosuite is a simulation framework for robot learning powered by the MuJoCo physics engine. It offers a modular design for creating robotic tasks as well as a suite of benchmark e…

cs.CV2023

Adaptive Human Matting for Dynamic Videos

Chung-Ching Lin, Jiang Wang, Kun Luo +4

The most recent efforts in video matting have focused on eliminating trimap dependency since trimap annotations are expensive and trimap-based methods are less adaptable for real-t…

math.AG2024

Proof of the geometric Langlands conjecture III: compatibility with parabolic induction

Justin Campbell, Lin Chen, Joakim Faergeman +4

We establish the compatibility of the Langlands functor with the operations of Eisenstein series constant term, and deduce that the Langlands functor induces an equivalence on Eise…

cs.CV2025

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Xiyao Wang, Zhengyuan Yang, Linjie Li +6

Despite significant advancements in vision-language models (VLMs), there lacks effective approaches to enhance response quality by scaling inference-time computation. This capabili…

cs.CV2024

GenXD: Generating Any 3D and 4D Scenes

Yuyang Zhao, Chung-Ching Lin, Kevin Lin +6

Recent developments in 2D visual generation have been remarkably successful. However, 3D and 4D generation remain challenging in real-world applications due to the lack of large-sc…

cs.CL2021

Constructing Taxonomies from Pretrained Language Models

Catherine Chen, Kevin Lin, Dan Klein

We present a method for constructing taxonomic trees (e.g., WordNet) using pretrained language models. Our approach is composed of two modules, one that predicts parenthood relatio…

cs.CV2024

DisCo: Disentangled Control for Realistic Human Dance Generation

Tan Wang, Linjie Li, Kevin Lin +6

Generative AI has made significant strides in computer vision, particularly in text-driven image/video synthesis (T2I/T2V). Despite the notable advancements, it remains challenging…

cs.CV2023

GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

An Yan, Zhengyuan Yang, Wanrong Zhu +9

We present MM-Navigator, a GPT-4V-based agent for the smartphone graphical user interface (GUI) navigation task. MM-Navigator can interact with a smartphone screen as human users,…

cs.CV2023

Neural Voting Field for Camera-Space 3D Hand Pose Estimation

Lin Huang, Chung-Ching Lin, Kevin Lin +4

We present a unified framework for camera-space 3D hand pose estimation from a single RGB image based on 3D implicit representation. As opposed to recent works, most of which first…

cs.CL2020

Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Zhuohan Li, Eric Wallace, Sheng Shen +4

Since hardware resources are limited, the objective of training deep learning models is typically to maximize accuracy subject to the time and memory constraints of training and in…

cs.CL2019

Reasoning Over Paragraph Effects in Situations

Kevin Lin, Oyvind Tafjord, Peter Clark +1

A key component of successfully reading a passage of text is the ability to apply knowledge gained from the passage to a new situation. In order to facilitate progress on this kind…

cs.CV2024

Partial-View Object View Synthesis via Filtered Inversion

Fan-Yun Sun, Jonathan Tremblay, Valts Blukis +11

We propose Filtering Inversion (FINV), a learning framework and optimization process that predicts a renderable 3D object representation from one or few partial views. FINV address…

cs.CL2024

Meta-DiffuB: A Contextualized Sequence-to-Sequence Text Diffusion Model with Meta-Exploration

Yun-Yen Chuang, Hung-Min Hsu, Kevin Lin +4

The diffusion model, a new generative modeling paradigm, has achieved significant success in generating images, audio, video, and text. It has been adapted for sequence-to-sequence…

cs.CV2021

End-to-End Human Pose and Mesh Reconstruction with Transformers

Kevin Lin, Lijuan Wang, Zicheng Liu

We present a new method, called MEsh TRansfOrmer (METRO), to reconstruct 3D human pose and mesh vertices from a single image. Our method uses a transformer encoder to jointly model…

cs.CV2023

Spatial-Frequency U-Net for Denoising Diffusion Probabilistic Models

Xin Yuan, Linjie Li, Jianfeng Wang +4

In this paper, we study the denoising diffusion probabilistic model (DDPM) in wavelet space, instead of pixel space, for visual synthesis. Considering the wavelet transform represe…

cs.AI2024

MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Weihao Yu, Zhengyuan Yang, Linjie Li +5

We propose MM-Vet, an evaluation benchmark that examines large multimodal models (LMMs) on complicated multimodal tasks. Recent LMMs have shown various intriguing abilities, such a…

cs.CV2024

Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation

Yuanhao Zhai, Kevin Lin, Zhengyuan Yang +6

Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, applying these techniques directly to video diffusion often results in unsatis…

cs.CV2024

Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation

Zhengyuan Yang, Jianfeng Wang, Linjie Li +4

We introduce ``Idea to Image,'' a system that enables multimodal iterative self-refinement with GPT-4V(ision) for automatic image design and generation. Humans can quickly identify…

cs.CL2025

Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Mrinank Sharma, Meg Tong, Jesse Mu +40

Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…

cs.CV2022

SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

Kevin Lin, Linjie Li, Chung-Ching Lin +5

The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on vid…

math.ST2021

Statistically and Computationally Efficient Change Point Localization in Regression Settings

Daren Wang, Zifeng Zhao, Kevin Lin +1

Detecting when the underlying distribution changes for the observed time series is a fundamental problem arising in a broad spectrum of applications. In this paper, we study multip…

cs.CL2018

Adversarial Ranking for Language Generation

Kevin Lin, Dianqi Li, Xiaodong He +2

Generative adversarial networks (GANs) have great successes on synthesizing data. However, the existing GANs restrict the discriminator to be a binary classifier, and thus limit th…

math.NA2016

Data-based stochastic model reduction for the Kuramoto--Sivashinsky equation

Fei Lu, Kevin Lin, Alexandre J. Chorin

The problem of constructing data-based, predictive, reduced models for the Kuramoto-Sivashinsky equation is considered, under circumstances where one has observation data only for…

cs.AI2025

Sleep-time Compute: Beyond Inference Scaling at Test-time

Kevin Lin, Charlie Snell, Yu Wang +4

Scaling test-time compute has emerged as a key ingredient for enabling large language models (LLMs) to solve difficult problems, but comes with high latency and inference cost. We…

cs.CY2021

Do Abstractions Have Politics? Toward a More Critical Algorithm Analysis

Kevin Lin

The expansion of computer science (CS) education in K--12 and higher-education in the United States has prompted deeper engagement with equity that moves beyond inclusion toward a…

cs.CV2022

GIT: A Generative Image-to-text Transformer for Vision and Language

Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu +6

In this paper, we design and train a Generative Image-to-text Transformer, GIT, to unify vision-language tasks such as image/video captioning and question answering. While generati…

cs.CY2020

Nifty Web Apps: Build a Web App for Any Text-Based Programming Assignment

Kevin Lin, Sumant Guha, Joe Spaniac +1

While many students now interact with web apps across a variety of smart devices, the vast majority of our Nifty Assignments still present traditional user interfaces such as conso…

cs.CV2025

BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation

Yuyang Peng, Shishi Xiao, Keming Wu +6

Recently, state-of-the-art text-to-image generation models, such as Flux and Ideogram 2.0, have made significant progress in sentence-level visual text rendering. In this paper, we…

cs.RO2025

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…

cs.CV2022

VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Tsu-Jui Fu, Linjie Li, Zhe Gan +4

A great challenge in video-language (VidL) modeling lies in the disconnection between fixed video representations extracted from image/video understanding models and downstream Vid…

cs.CV2021

OVIS: Open-Vocabulary Visual Instance Search via Visual-Semantic Aligned Representation Learning

Sheng Liu, Kevin Lin, Lijuan Wang +2

We introduce the task of open-vocabulary visual instance search (OVIS). Given an arbitrary textual search query, Open-vocabulary Visual Instance Search (OVIS) aims to return a rank…

cs.CL2020

Neural Module Networks for Reasoning over Text

Nitish Gupta, Kevin Lin, Dan Roth +2

Answering compositional questions that require multiple steps of reasoning against text is challenging, especially when they involve discrete, symbolic operations. Neural module ne…

cs.CL2019

QuaRTz: An Open-Domain Dataset of Qualitative Relationship Questions

Oyvind Tafjord, Matt Gardner, Kevin Lin +1

We introduce the first open-domain dataset, called QuaRTz, for reasoning about textual qualitative relationships. QuaRTz contains general qualitative statements, e.g., "A sunscreen…

cs.CV2025

Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample Optimization

Zichen Miao, Zhengyuan Yang, Kevin Lin +4

Recent advancements in timestep-distilled diffusion models have enabled high-quality image generation that rivals non-distilled multi-step models, but with significantly fewer infe…

cs.CV2021

VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning

Xiaowei Hu, Xi Yin, Kevin Lin +4

It is highly desirable yet challenging to generate image captions that can describe novel objects which are unseen in caption-labeled training data, a capability that is evaluated…

cs.DB2019

DeepBase: Deep Inspection of Neural Networks

Thibault Sellam, Kevin Lin, Ian Yiran Huang +4

Although deep learning models perform remarkably well across a range of tasks such as language translation and object recognition, it remains unclear what high-level logic, if any,…

cs.CV2025

SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Xiyao Wang, Zhengyuan Yang, Chao Feng +6

We introduce ThinkLite-VL, a family of visual reasoning models that achieve state-of-the-art (SoTA) performance using an order of magnitude fewer training samples, relying purely o…

cs.CV2017

Supervised Learning of Semantics-Preserving Hash via Deep Convolutional Neural Networks

Huei-Fang Yang, Kevin Lin, Chu-Song Chen

This paper presents a simple yet effective supervised deep hash approach that constructs binary hash codes from labeled data for large-scale image search. We assume that the semant…

cs.CV2025

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

Minheng Ni, Zhengyuan Yang, Linjie Li +4

Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, exte…

math.PR2019

Can the Stochastic Wave Equation with Strong Drift Hit Zero?

Kevin Lin, Carl Mueller

We study the stochastic wave equation with multiplicative noise and singular drift: \[ \partial_tu(t,x)=Δu(t,x)+u^{-α}(t,x)+g(u(t,x))\dot{W}(t,x) \] where lies in the circle…

cs.CV2021

Mesh Graphormer

Kevin Lin, Lijuan Wang, Zicheng Liu

We present a graph-convolution-reinforced transformer, named Mesh Graphormer, for 3D human pose and mesh reconstruction from a single image. Recently both transformers and graph co…

cs.CV2018

Adversarial Learning for Fine-grained Image Search

Kevin Lin, Fan Yang, Qiaosong Wang +1

Fine-grained image search is still a challenging problem due to the difficulty in capturing subtle differences regardless of pose variations of objects from fine-grained categories…

cs.RO2025

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Alexander Khazatsky, Karl Pertsch, Suraj Nair +98

The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. Ho…

cs.CV2023

MPT: Mesh Pre-Training with Transformers for Human Pose and Mesh Reconstruction

Kevin Lin, Chung-Ching Lin, Lin Liang +2

Traditional methods of reconstructing 3D human pose and mesh from single images rely on paired image-mesh datasets, which can be difficult and expensive to obtain. Due to this limi…

eess.AS2025

Audio-Aware Large Language Models as Judges for Speaking Styles

Cheng-Han Chiang, Xiaofei Wang, Chung-Ching Lin +8

Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to…

cs.CV2023

DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design

Kevin Lin, Zhengyuan Yang, Linjie Li +2

We introduce DEsignBench, a text-to-image (T2I) generation benchmark tailored for visual design scenarios. Recent T2I models like DALL-E 3 and others, have demonstrated remarkable…

cs.CV2026

EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing

Tianyu Chen, Yasi Zhang, Zhi Zhang +13

Instruction-based image editing has advanced rapidly, yet reliable and interpretable evaluation remains a bottleneck. Current protocols either (i) depend on paired reference images…

cs.CV2024

MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Xuehai He, Weixi Feng, Kaizhi Zheng +11

Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these ab…

cs.CV2023

MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning

Chaoyi Zhang, Kevin Lin, Zhengyuan Yang +5

We present MM-Narrator, a novel system leveraging GPT-4 with multimodal in-context learning for the generation of audio descriptions (AD). Unlike previous methods that primarily fo…

cs.CV2023

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Zhengyuan Yang, Linjie Li, Kevin Lin +4

Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper,…

cs.CY2019

Subgoals, Problem Solving Phases, and Sources of Knowledge: A Complex Mangle

Kevin Lin, David DeLiema

Educational researchers have increasingly drawn attention to how students develop computational thinking (CT) skills, including in science, math, and literacy contexts. A key compo…

cs.CV2023

An Empirical Study of Multimodal Model Merging

Yi-Lin Sung, Linjie Li, Kevin Lin +3

Model merging (e.g., via interpolation or task arithmetic) fuses multiple models trained on different tasks to generate a multi-task solution. The technique has been proven success…

cs.CY2024

"It Can Relate to Real Lives": Attitudes and Expectations in Justice-Centered Data Structures & Algorithms for Non-Majors

Anna Batra, Iris Zhou, Suh Young Choi +4

Prior work has argued for a more justice-centered approach to postsecondary computing education by emphasizing ethics, identity, and political vision. In this experience report, we…

cs.CL2023

Few-Shot Adaptation for Parsing Contextual Utterances with LLMs

Kevin Lin, Patrick Xia, Hao Fang

We evaluate the ability of semantic parsers based on large language models (LLMs) to handle contextual utterances. In real-world settings, there typically exists only a limited num…

math.AG2024

Coherent sheaves, sheared D-modules, and Hochschild cochains

Dario Beraldo, Kevin Lin, Wyatt Reeves

We show that the category of ind-coherent sheaves on a quasi-smooth scheme is naturally tensored over the category of sheared D-modules on its shifted cotangent bundle, commuting w…

cs.CL2025

SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models

Cheng-Han Chiang, Xiaofei Wang, Linjie Li +7

Current large language models (LLMs) and spoken language models (SLMs) begin thinking and taking actions only after the user has finished their turn. This prevents the model from i…

cs.CY2021

CS Education for the Socially-Just Worlds We Need: The Case for Justice-Centered Approaches to CS in Higher Education

Kevin Lin

Justice-centered approaches to equitable computer science (CS) education frame CS learning as a means for advancing peace, antiracism, and social justice rather than war, empire, a…

cs.CL2020

Evaluating Models' Local Decision Boundaries via Contrast Sets

Matt Gardner, Yoav Artzi, Victoria Basmova +23

Standard test sets for supervised learning evaluate in-distribution generalization. Unfortunately, when a dataset has systematic gaps (e.g., annotation artifacts), these evaluation…

cs.RO2025

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

NVIDIA, :, Johan Bjorck +40

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist…

cs.CV2025

ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning

Jiaqi Liao, Zhengyuan Yang, Linjie Li +4

In this work, we study the problem of Text-to-Image In-Context Learning (T2I-ICL). While Unified Multimodal LLMs (MLLMs) have advanced rapidly in recent years, they struggle with c…

cs.RO2025

DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning

Zhenyu Jiang, Yuqi Xie, Kevin Lin +5

Imitation learning from human demonstrations is an effective means to teach robots manipulation skills. But data acquisition is a major bottleneck in applying this paradigm more br…

cs.IR2024

CLARINET: Augmenting Language Models to Ask Clarification Questions for Retrieval

Yizhou Chi, Jessy Lin, Kevin Lin +1

Users often make ambiguous requests that require clarification. We study the problem of asking clarification questions in an information retrieval setting, where systems often face…

cs.CV2023

Equivariant Similarity for Vision-Language Foundation Models

Tan Wang, Kevin Lin, Linjie Li +5

This study explores the concept of equivariance in vision-language foundation models (VLMs), focusing specifically on the multimodal similarity function that is not only the major…

cs.CL2026

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

Cheng-Han Chiang, Xiaofei Wang, Linjie Li +7

Spoken Language Models (SLMs) are designed to take speech inputs and produce spoken responses. However, current SLMs lack the ability to perform an internal, unspoken thinking proc…

cs.CV2020

Learning Nonparametric Human Mesh Reconstruction from a Single Image without Ground Truth Meshes

Kevin Lin, Lijuan Wang, Ying Jin +2

Nonparametric approaches have shown promising results on reconstructing 3D human mesh from a single monocular image. Unlike previous approaches that use a parametric human model li…

cs.CV2024

MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Weihao Yu, Zhengyuan Yang, Lingfeng Ren +7

MM-Vet, with open-ended vision-language questions targeting at evaluating integrated capabilities, has become one of the most popular benchmarks for large multimodal model evaluati…

cs.CY2024

Code Interviews: Design and Evaluation of a More Authentic Assessment for Introductory Programming Assignments

Suhas Kannam, Yuri Yang, Aarya Dharm +1

Generative artificial intelligence poses new challenges around assessment, increasingly driving introductory programming educators to employ invigilated exams. But exams do not aff…

cs.CL2025

Measurement of LLM's Philosophies of Human Nature

Minheng Ni, Ennan Wu, Zidong Gong +6

The widespread application of artificial intelligence (AI) in various tasks, along with frequent reports of conflicts or violations involving AI, has sparked societal concerns abou…

cs.RO2025

Constraint-Preserving Data Generation for Visuomotor Policy Learning

Kevin Lin, Varun Ragunath, Andrew McAlinden +4

Large-scale demonstration data has powered key breakthroughs in robot manipulation, but collecting that data remains costly and time-consuming. We present Constraint-Preserving Dat…

cs.CV2023

MM-VID: Advancing Video Understanding with GPT-4V(ision)

Kevin Lin, Faisal Ahmed, Linjie Li +9

We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video unders…

stat.ML2013

Optimization for Compressed Sensing: the Simplex Method and Kronecker Sparsification

Robert Vanderbei, Han Liu, Lie Wang +1

In this paper we present two new approaches to efficiently solve large-scale compressed sensing problems. These two ideas are independent of each other and can therefore be used ei…

stat.ME2018

Post-Selection Inference for Changepoint Detection Algorithms with Application to Copy Number Variation Data

Sangwon Hyun, Kevin Lin, Max G'Sell +1

Changepoint detection methods are used in many areas of science and engineering, e.g., in the analysis of copy number variation data, to detect abnormalities in copy numbers along…

cs.CV2022

ReCo: Region-Controlled Text-to-Image Generation

Zhengyuan Yang, Jianfeng Wang, Zhe Gan +8

Recently, large-scale text-to-image (T2I) models have shown impressive performance in generating high-fidelity images, but with limited controllability, e.g., precisely specifying…

cs.GR2025

EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing

Kaizhi Zheng, Xiaotong Chen, Xuehai He +7

Given the steep learning curve of professional 3D software and the time-consuming process of managing large 3D assets, language-guided 3D scene editing has significant potential in…

cs.CV2025

List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs

An Yan, Zhengyuan Yang, Junda Wu +8

Set-of-Mark (SoM) Prompting unleashes the visual grounding capability of GPT-4V, by enabling the model to associate visual objects with tags inserted on the image. These tags, mark…

cs.CV2022

LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling

Linjie Li, Zhe Gan, Kevin Lin +4

Unified vision-language frameworks have greatly advanced in recent years, most of which adopt an encoder-decoder architecture to unify image-text tasks as sequence-to-sequence gene…

cs.CV2020

Cross-Domain Complementary Learning Using Pose for Multi-Person Part Segmentation

Kevin Lin, Lijuan Wang, Kun Luo +3

Supervised deep learning with pixel-wise training labels has great successes on multi-person part segmentation. However, data labeling at pixel-level is very expensive. To solve th…

cs.CV2026

TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation

Minheng Ni, Zhengyuan Yang, Yaowen Zhang +9

We study technical image generation, where a model must synthesize information-dense, scientifically precise illustrations from detailed descriptions rather than merely produce vis…

cs.CV2024

IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation

Yuanhao Zhai, Kevin Lin, Linjie Li +7

Significant advances have been made in human-centric video generation, yet the joint video-depth generation problem remains underexplored. Most existing monocular depth estimation…

cs.CV2023

An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling

Tsu-Jui Fu, Linjie Li, Zhe Gan +4

Masked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have…

q-bio.NC2025

Latency correction in sparse neuronal spike trains with overlapping global events

Arturo Mariani, Federico Senocrate, Jason Mikiel-Hunter +5

Background: In Kreuz et al., J Neurosci Methods 381, 109703 (2022) two methods were proposed that perform latency correction, i.e., optimize the spike time alignment of sparse neur…

cs.CV2024

LiVOS: Light Video Object Segmentation with Gated Linear Matching

Qin Liu, Jianfeng Wang, Zhengyuan Yang +4

Semi-supervised video object segmentation (VOS) has been largely driven by space-time memory (STM) networks, which store past frame features in a spatiotemporal memory to segment t…

math.RT2022

Poincare series and miraculous duality

Kevin Lin

In the setting of global geometric Langlands, we show that miraculous duality on the stack of principal bundles on a curve intertwines the functor of Poincare series with the dual…

stat.ME2016

Approximate Recovery in Changepoint Problems, from Estimation Error Rates

Kevin Lin, James Sharpnack, Alessandro Rinaldo +1

In the 1-dimensional multiple changepoint detection problem, we prove that any procedure with a fast enough error rate, in terms of its estimation of the underlying piecew…

cs.CL2023

Lost in the Middle: How Language Models Use Long Contexts

Nelson F. Liu, Kevin Lin, John Hewitt +4

While recent language models have the ability to take long contexts as input, relatively little is known about how well they use longer context. We analyze the performance of langu…

math.AG2026

On the excursion algebra

Dennis Gaitsgory, Kevin Lin, Wyatt Reeves

The excursion algebra associated to a scheme X over a finite field and a reductive group G is the algebra of global functions on the stack of arithmetic G-local systems on X. When…

eess.IV2023

Uncertainty Quantification for Eosinophil Segmentation

Kevin Lin, Donald Brown, Sana Syed +1

Eosinophilic Esophagitis (EoE) is an allergic condition increasing in prevalence. To diagnose EoE, pathologists must find 15 or more eosinophils within a single high-power field (4…

cs.RO2022

Combining optimal control and learning for autonomous aerial navigation in novel indoor environments

Kevin Lin, Brian Huo, Megan Hu

This report proposes a combined optimal control and perception framework for Micro Aerial Vehicle (MAV) autonomous navigation in novel indoor enclosed environments, relying exclusi…

cs.CV2023

OpenLEAF: Open-Domain Interleaved Image-Text Generation and Evaluation

Jie An, Zhengyuan Yang, Linjie Li +5

This work investigates a challenging task named open-domain interleaved image-text generation, which generates interleaved texts and images following an input query. We propose a n…

cs.CL2025

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding

Ming Li, Zhengyuan Yang, Xiyao Wang +4

Large reasoning models (LRMs) achieve strong reasoning performance by emitting long chains of thought. Yet, these verbose traces slow down inference and often drift into unnecessar…

cs.RO2024

Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation

Aaditya Prasad, Kevin Lin, Jimmy Wu +2

Many robotic systems, such as mobile manipulators or quadrotors, cannot be equipped with high-end GPUs due to space, weight, and power constraints. These constraints prevent these…

cs.CV2024

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Fuxiao Liu, Kevin Lin, Linjie Li +3

Despite the promising progress in multi-modal tasks, current large multi-modal models (LMMs) are prone to hallucinating inconsistent descriptions with respect to the associated ima…

cs.RO2023

Text2Motion: From Natural Language Instructions to Feasible Plans

Kevin Lin, Christopher Agia, Toki Migimatsu +2

We propose Text2Motion, a language-based planning framework enabling robots to solve sequential manipulation tasks that require long-horizon reasoning. Given a natural language ins…

cs.CV2024

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin +5

In the evolution of Vision-Language Pre-training, shifting from short-text comprehension to encompassing extended textual contexts is pivotal. Recent autoregressive vision-language…

cs.CL2020

Learning to Generate Multiple Style Transfer Outputs for an Input Sentence

Kevin Lin, Ming-Yu Liu, Ming-Ting Sun +1

Text style transfer refers to the task of rephrasing a given text in a different style. While various methods have been proposed to advance the state of the art, they often assume…

cs.CL2019

Grammar-based Neural Text-to-SQL Generation

Kevin Lin, Ben Bogin, Mark Neumann +2

The sequence-to-sequence paradigm employed by neural text-to-SQL models typically performs token-level decoding and does not consider generating SQL hierarchically from a grammar.…