papers

Publications (85)

cs.CL2024

Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts

Junmo Kang, Leonid Karlinsky, Hongyin Luo +7

cs.CV2025

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

Granite Vision Team, Leonid Karlinsky, Assaf Arbelle +60

cs.CV2021

CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

Chun-Fu Chen, Quanfu Fan, Rameswar Panda

cs.LG2020

A Maximal Correlation Approach to Imposing Fairness in Machine Learning

Joshua Lee, Yuheng Bu, Prasanna Sattigeri +4

cs.CV2020

Exploiting Global Camera Network Constraints for Unsupervised Video Person Re-identification

Xueping Wang, Rameswar Panda, Min Liu +2

cs.CV2023

Learning Human Action Recognition Representations Without Real Humans

Howard Zhong, Samarth Mishra, Donghyun Kim +7

cs.CV2021

AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition

Rameswar Panda, Chun-Fu Chen, Quanfu Fan +4

cs.CL2026

CodeAlchemy: Synthetic Code Rewriting at Scale

Ankit Gupta, Aditya Prasad, Rameswar Panda

cs.CV2016

Video Summarization in a Multi-View Camera Network

Rameswar Panda, Abir Das, Amit K. Roy-Chowdhury

cs.CV2017

Collaborative Summarization of Topic-Related Videos

Rameswar Panda, Amit K. Roy-Chowdhury

cs.CV2021

Dynamic Distillation Network for Cross-Domain Few-Shot Recognition with Unlabeled Data

Ashraful Islam, Chun-Fu Chen, Rameswar Panda +3

cs.CV2020

Mitigating Dataset Imbalance via Joint Generation and Classification

Aadarsh Sahoo, Ankit Singh, Rameswar Panda +2

eess.IV2020

Non-Adversarial Video Synthesis with Learned Priors

Abhishek Aich, Akash Gupta, Rameswar Panda +3

cs.CV2021

Dynamic Network Quantization for Efficient Video Inference

Ximeng Sun, Rameswar Panda, Chun-Fu Chen +3

cs.CV2022

Semi-Supervised Domain Adaptation with Auto-Encoder via Simultaneous Learning

Md Mahmudur Rahman, Rameswar Panda, Mohammad Arif Ul Alam

cs.AI2026

Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models

Xinyi Wang, Shawn Tan, Shenbo Xu +4

cs.CV2021

AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition

Yue Meng, Rameswar Panda, Chung-Ching Lin +5

cs.CV2020

Large Scale Neural Architecture Search with Polyharmonic Splines

Ulrich Finkler, Michele Merler, Rameswar Panda +8

cs.CV2020

Adversarial Knowledge Transfer from Unlabeled Data

Akash Gupta, Rameswar Panda, Sujoy Paul +2

cs.CL2023

Neural Architecture Search for Effective Teacher-Student Knowledge Transfer in Language Models

Aashka Trivedi, Takuma Udagawa, Michele Merler +3

cs.CV2024

LangNav: Language as a Perceptual Representation for Navigation

Bowen Pan, Rameswar Panda, SouYoung Jin +4

cs.LG2024

Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Bowen Pan, Yikang Shen, Haokun Liu +5

cs.CV2023

CODA-Prompt: COntinual Decomposed Attention-based Prompting for Rehearsal-Free Continual Learning

James Seale Smith, Leonid Karlinsky, Vyshnavi Gutta +6

cs.CV2021

Contrast and Mix: Temporal Contrastive Video Domain Adaptation with Background Mixing

Aadarsh Sahoo, Rutav Shah, Rameswar Panda +2

cs.DC2025

The infrastructure powering IBM's Gen AI model development

Talia Gershon, Seetharami Seelam, Brian Belgodere +143

cs.CV2017

Unsupervised Adaptive Re-identification in Open World Dynamic Camera Networks

Rameswar Panda, Amran Bhuiyan, Vittorio Murino +1

cs.LG2024

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

William Brandon, Mayank Mishra, Aniruddha Nrusimha +2

cs.CV2022

Task2Sim : Towards Effective Pre-training and Transfer from Synthetic Data

Samarth Mishra, Rameswar Panda, Cheng Perng Phoo +5

cs.CV2022

RegionViT: Regional-to-Local Attention for Vision Transformers

Chun-Fu Chen, Rameswar Panda, Quanfu Fan

cs.CV2020

Measurement-driven Security Analysis of Imperceptible Impersonation Attacks

Shasha Li, Karim Khalil, Rameswar Panda +4

cs.CV2016

Continuous Adaptation of Multi-Camera Person Identification Models through Sparse Non-redundant Representative Selection

Abir Das, Rameswar Panda, Amit K. Roy-Chowdhury

cs.LG2025

TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments

Zhangchen Xu, Adriana Meza Soria, Shawn Tan +4

cs.LG2023

Learning to Grow Pretrained Models for Efficient Transformer Training

Peihao Wang, Rameswar Panda, Lucas Torroba Hennigen +6

cs.CV2021

Semi-Supervised Action Recognition with Temporal Contrastive Learning

Ankit Singh, Omprakash Chakraborty, Ashutosh Varshney +4

cs.LG2026

Dynamic Short Convolutions Improve Transformers

Oliver Sieberling, Bharat Runwal, Rameswar Panda +1

cs.CV2018

Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias

Rameswar Panda, Jianming Zhang, Haoxiang Li +3

cs.CV2021

VA-RED: Video Adaptive Redundancy Reduction

Bowen Pan, Rameswar Panda, Camilo Fosco +6

cs.CV2018

FFNet: Video Fast-Forwarding via Reinforcement Learning

Shuyue Lan, Rameswar Panda, Qi Zhu +1

cs.LG2025

FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference

Aniruddha Nrusimha, William Brandon, Mayank Mishra +4

cs.CV2021

IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers

Bowen Pan, Rameswar Panda, Yifan Jiang +3

cs.CL2025

Calibrating Expressions of Certainty

Peiqi Wang, Barbara D. Lam, Yingcheng Liu +5

cs.CV2020

Camera On-boarding for Person Re-identification using Hypothesis Transfer Learning

Sk Miraj Ahmed, Aske R Lejbølle, Rameswar Panda +1

cs.CV2022

Can An Image Classifier Suffice For Action Recognition?

Quanfu Fan, Chun-Fu, Chen +1

cs.AI2024

Scaling Granite Code Models to 128K Context

Matt Stallone, Vaibhav Saxena, Leonid Karlinsky +19

cs.LG2026

PRISM: Demystifying Retention and Interaction in Mid-Training

Bharat Runwal, Ashish Agrawal, Anurag Roy +1

cs.CL2025

Distilling to Hybrid Attention Models via KL-Guided Layer Selection

Yanhong Li, Songlin Yang, Shawn Tan +4

cs.CV2020

AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning

Ximeng Sun, Rameswar Panda, Rogerio Feris +1

cs.CL2024

Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler

Yikang Shen, Matthew Stallone, Mayank Mishra +6

cs.CL2023

Synthetic Pre-Training Tasks for Neural Machine Translation

Zexue He, Graeme Blackwood, Rameswar Panda +2

cs.CV2021

A Broad Study on the Transferability of Visual Representations with Contrastive Learning

Ashraful Islam, Chun-Fu Chen, Rameswar Panda +3

cs.CV2022

FETA: Towards Specializing Foundation Models for Expert Task Applications

Amit Alfassy, Assaf Arbelle, Oshri Halimi +10

cs.CV2021

Deep Analysis of CNN-based Spatio-temporal Representations for Action Recognition

Chun-Fu Chen, Rameswar Panda, Kandan Ramakrishnan +4

cs.CL2026

PaTH Attention: Position Encoding via Accumulating Householder Transformations

Songlin Yang, Yikang Shen, Kaiyue Wen +5

cs.LG2023

Energy Transformer

Benjamin Hoover, Yuchen Liang, Bao Pham +5

cs.CV2023

Teaching Structured Vision&Language Concepts to Vision&Language Models

Sivan Doveh, Assaf Arbelle, Sivan Harary +8

cs.CV2023

Select, Label, and Mix: Learning Discriminative Invariant Feature Representations for Partial Domain Adaptation

Aadarsh Sahoo, Rameswar Panda, Rogerio Feris +2

cs.LG2023

ConStruct-VL: Data-Free Continual Structured VL Concepts Learning

James Seale Smith, Paola Cascante-Bonilla, Assaf Arbelle +7

cs.LG2024

Scattered Mixture-of-Experts Implementation

Shawn Tan, Yikang Shen, Rameswar Panda +1

cs.LG2024

Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks

Ibrahim Abdelaziz, Kinjal Basu, Mayank Agarwal +23

cs.LG2025

Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study

Shawn Tan, Songlin Yang, Aaron Courville +2

cs.CV2020

AR-Net: Adaptive Frame Resolution for Efficient Action Recognition

Yue Meng, Chung-Ching Lin, Rameswar Panda +5

cs.CV2019

Estimating Skin Tone and Effects on Classification Performance in Dermatology Datasets

Newton M. Kinyanjui, Timothy Odonga, Celia Cintas +4

cs.CL2025

API Pack: A Massive Multi-Programming Language Dataset for API Call Generation

Zhen Guo, Adriana Meza Soria, Wei Sun +2

cs.CV2023

Going Beyond Nouns With Vision & Language Models Using Synthetic Data

Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith +8

cs.CV2021

Detector-Free Weakly Supervised Grounding by Separation

Assaf Arbelle, Sivan Doveh, Amit Alfassy +14

cs.CV2021

NASTransfer: Analyzing Architecture Transferability in Large Scale Neural Architecture Search

Rameswar Panda, Michele Merler, Mayoore Jaiswal +8

cs.LG2022

Selective Regression Under Fairness Criteria

Abhin Shah, Yuheng Bu, Joshua Ka-Wing Lee +4

cs.CV2021

Improved Techniques for Quantizing Deep Networks with Adaptive Bit-Widths

Ximeng Sun, Rameswar Panda, Chun-Fu Chen +6

cs.LG2024

Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization

Aniruddha Nrusimha, Mayank Mishra, Naigang Wang +3

cs.CL2023

Multitask Prompt Tuning Enables Parameter-Efficient Transfer Learning

Zhen Wang, Rameswar Panda, Leonid Karlinsky +3

cs.CV2023

Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models

Sivan Doveh, Assaf Arbelle, Sivan Harary +9

cs.CV2017

Diversity-aware Multi-Video Summarization

Rameswar Panda, Niluthpol Chowdhury Mithun, Amit K. Roy-Chowdhury

cs.AI2024

Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Mayank Mishra, Matt Stallone, Gaoyuan Zhang +43

cs.LG2024

Gated Linear Attention Transformers with Hardware-Efficient Training

Songlin Yang, Bailin Wang, Yikang Shen +2

cs.LG2024

Diversity Measurement and Subset Selection for Instruction Tuning Datasets

Peiqi Wang, Yikang Shen, Zhen Guo +4

cs.CV2021

AVLnet: Learning Audio-Visual Language Representations from Instructional Videos

Andrew Rouditchenko, Angie Boggust, David Harwath +11

cs.CL2024

Data Engineering for Scaling Language Models to 128K Context

Yao Fu, Rameswar Panda, Xinyao Niu +4

cs.CV2022

VALHALLA: Visual Hallucination for Machine Translation

Yi Li, Rameswar Panda, Yoon Kim +4

cs.CL2021

Cascaded Multilingual Audio-Visual Learning from Videos

Andrew Rouditchenko, Angie Boggust, David Harwath +8

cs.CL2026

Variable-Width Transformers

Zhaofeng Wu, Oliver Sieberling, Shawn Tan +3

cs.MM2018

Webly Supervised Joint Embedding for Cross-Modal Image-Text Retrieval

Niluthpol Chowdhury Mithun, Rameswar Panda, Evangelos E. Papalexakis +1

cs.CV2024

SITAR: Semi-supervised Image Transformer for Action Recognition

Owais Iqbal, Omprakash Chakraborty, Aftab Hussain +2

cs.CV2021

Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Brian Chen, Andrew Rouditchenko, Kevin Duarte +10

cs.CV2023

MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge

Wei Lin, Leonid Karlinsky, Nina Shvetsova +6

cs.CV2017

Multi-View Surveillance Video Summarization via Joint Embedding and Sparse Optimization

Rameswar Panda, Amit K. Roy-Chowdhury