papers

Publications (37)

cs.CV2026

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

Xingrui Wang, Jiang Liu, Chao Huang +7

Omni-modal large language models (OLLMs) aim to unify audio, vision, and text understanding within a single framework. While existing benchmarks primarily evaluate general cross-mo…

cs.CV2020

AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning

Ximeng Sun, Rameswar Panda, Rogerio Feris +1

Multi-task learning is an open and challenging problem in computer vision. The typical way of conducting multi-task learning with deep neural networks is either through handcrafted…

cs.LG2025

APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation

Yuzhen Zhou, Jiajun Li, Yusheng Su +15

Reinforcement learning (RL) has become a cornerstone in advancing large-scale pre-trained language models (LLMs). Successive generations, including GPT-o series, DeepSeek-R1, Kimi-…

cs.CV2021

AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition

Rameswar Panda, Chun-Fu Chen, Quanfu Fan +4

Multi-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learni…

cs.HC2025

Agent Laboratory: Using LLM Agents as Research Assistants

Samuel Schmidgall, Yusheng Su, Ze Wang +7

Historically, scientific discovery has been a lengthy and costly process, demanding substantial time and resources from initial conception to final results. To accelerate scientifi…

cs.CV2018

A Temporally-Aware Interpolation Network for Video Frame Inpainting

Ximeng Sun, Ryan Szeto, Jason J. Corso

We propose the first deep learning solution to video frame inpainting, a challenging instance of the general video inpainting problem with applications in video editing, manipulati…

cs.CL2026

Reliable Use of Lemmas via Eligibility Reasoning and SectionAware Reinforcement Learning

Zhikun Xu, Xiaodong Yu, Ben Zhou +6

Recent large language models (LLMs) perform strongly on mathematical benchmarks yet often misapply lemmas, importing conclusions without validating assumptions. We formalize lemma$…

cs.CV2021

Dynamic Network Quantization for Efficient Video Inference

Ximeng Sun, Rameswar Panda, Chun-Fu Chen +3

Deep convolutional networks have recently achieved great success in video recognition, yet their practical realization remains a challenge due to the large amount of computational…

cs.CV2026

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Shenghui Chen, Po-han Li, Ximeng Sun +5

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation…

cs.CV2023

DIME-FM: DIstilling Multimodal and Efficient Foundation Models

Ximeng Sun, Pengchuan Zhang, Peizhao Zhang +3

Large Vision-Language Foundation Models (VLFM), such as CLIP, ALIGN and Florence, are trained on large-scale datasets of image-caption pairs and achieve superior transferability an…

cs.CV2019

Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition

Ping Hu, Ximeng Sun, Kate Saenko +1

Learning from a few examples is a challenging task for machine learning. While recent progress has been made for this problem, most of the existing methods ignore the compositional…

cs.CV2025

Unleashing Hour-Scale Video Training for Long Video-Language Understanding

Jingyang Lin, Jialian Wu, Ximeng Sun +8

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has…

cs.CV2026

DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning

Chao Huang, Zeliang Zhang, Jiang Liu +7

Multimodal large language models (MLLMs) have made rapid progress, yet their reasoning ability often lags behind strong text-only LLMs. Bridging this gap typically requires large-s…

cs.CV2023

DualCoOp++: Fast and Effective Adaptation to Multi-Label Recognition with Limited Annotations

Ping Hu, Ximeng Sun, Stan Sclaroff +1

Multi-label image recognition in the low-label regime is a task of great challenge and practical significance. Previous works have focused on learning the alignment between textual…

cs.AI2026

TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents

Kaijie Zhu, Yuzhou Nie, Yijiang Li +10

Executing complex terminal tasks remains a significant challenge for open-weight LLMs, constrained by two fundamental limitations. First, high-fidelity, executable training environ…

cs.CV2024

Koala: Key frame-conditioned long video-LLM

Reuben Tan, Ximeng Sun, Ping Hu +5

Long video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Lar…

cs.CV2025

KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation

Xingrui Wang, Jiang Liu, Ze Wang +7

Generating video from various conditions, such as text, image, and audio, enables both spatial and temporal control, leading to high-quality generation results. Videos with dramati…

cs.CV2025

Latent Visual Reasoning

Bangzheng Li, Ximeng Sun, Jiang Liu +7

Multimodal Large Language Models (MLLMs) have achieved notable gains in various tasks by incorporating Chain-of-Thought (CoT) reasoning in language spaces. Recent work extends this…

cs.CV2021

Improved Techniques for Quantizing Deep Networks with Adaptive Bit-Widths

Ximeng Sun, Rameswar Panda, Chun-Fu Chen +6

Quantizing deep networks with adaptive bit-widths is a promising technique for efficient inference across many devices and resource constraints. In contrast to static methods that…

cs.CV2025

ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning

Yuxiang Guo, Jiang Liu, Ze Wang +7

The rapid advancement of text-to-image (T2I) models has increased the need for reliable human preference modeling, a demand further amplified by recent progress in reinforcement le…

cs.CV2025

MOVi: Training-free Text-conditioned Multi-Object Video Generation

Aimon Rahman, Jiang Liu, Ze Wang +7

Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objec…

cs.CL2025

Instella: Fully Open Language Models with Stellar Performance

Jiang Liu, Jialian Wu, Xiaodong Yu +10

Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks, yet the majority of high-performing models remain closed-source or partially ope…

cs.CV2025

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer

Hao Chen, Ze Wang, Xiang Li +7

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models. We present SoftVQ-VAE, a continuous image tokenizer that leve…

cs.CV2019

Domain Agnostic Learning with Disentangled Representations

Xingchao Peng, Zijun Huang, Ximeng Sun +1

Unsupervised model transfer has the potential to greatly improve the generalizability of deep models to novel domains. Yet the current literature assumes that the separation of tar…

cs.CV2025

Learning from Online Videos at Inference Time for Computer-Use Agents

Yujian Liu, Ze Wang, Hao Chen +7

Computer-use agents can operate computers and automate laborious tasks, but despite recent rapid progress, they still lag behind human users, especially when tasks require domain-s…

cs.CV2022

DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations

Ximeng Sun, Ping Hu, Kate Saenko

Solving multi-label recognition (MLR) for images in the low-label regime is a challenging task with many real-world applications. Recent work learns an alignment between textual an…

cs.CV2024

CLAMP: Contrastive LAnguage Model Prompt-tuning

Piotr Teterwak, Ximeng Sun, Bryan A. Plummer +2

Large language models (LLMs) have emerged as powerful general-purpose interfaces for many machine learning problems. Recent work has adapted LLMs to generative visual tasks like im…

cs.CV2020

Revisiting Few-shot Activity Detection with Class Similarity Control

Huijuan Xu, Ximeng Sun, Eric Tzeng +3

Many interesting events in the real world are rare making preannotated machine learning ready videos a rarity in consequence. Thus, temporal activity detection models that are able…

cs.CL2025

Self-Taught Agentic Long Context Understanding

Yufan Zhuang, Xiaodong Yu, Jialian Wu +7

Answering complex, long-context questions remains a major challenge for large language models (LLMs) as it requires effective question clarifications and context retrieval. We prop…

cs.LG2023

Label Budget Allocation in Multi-Task Learning

Ximeng Sun, Kihyuk Sohn, Kate Saenko +2

The cost of labeling data often limits the performance of machine learning systems. In multi-task learning, related tasks provide information to each other and improve overall perf…

cs.CV2018

Similarity R-C3D for Few-shot Temporal Activity Detection

Huijuan Xu, Bingyi Kang, Ximeng Sun +3

Many activities of interest are rare events, with only a few labeled examples available. Therefore models for temporal activity detection which are able to learn from a few example…

cs.CV2020

TwoStreamVAN: Improving Motion Modeling in Video Generation

Ximeng Sun, Huijuan Xu, Kate Saenko

Video generation is an inherently challenging task, as it requires modeling realistic temporal dynamics as well as spatial content. Existing methods entangle the two intrinsically…

cs.CL2026

CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models

Yihao Liang, Ze Wang, Hao Chen +7

Autoregressive large language models achieve strong results on many benchmarks, but decoding remains fundamentally latency-limited by sequential dependence on previously generated…

cs.CL2026

Stabilizing Efficient Reasoning with Step-Level Advantage Selection

Han Wang, Xiaodong Yu, Jialian Wu +4

Large language models (LLMs) achieve strong reasoning performance by allocating substantial computation at inference time, often generating long and verbose reasoning traces. While…

cs.CV2026

CaptionQA: Is Your Caption as Useful as the Image Itself?

Shijia Yang, Yunong Liu, Bohan Zhai +5

Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic inference pipelines. Yet current eva…

cs.CV2025

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation

Ze Wang, Hao Chen, Benran Hu +7

Image tokenization plays a critical role in reducing the computational demands of modeling high-resolution images, significantly improving the efficiency of image and multimodal un…

cs.CV2026

VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking

Jingyang Lin, Jialian Wu, Jiang Liu +6

Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, result…