#data augmentation

try —

15 papers match

cs.SD2026

Teffic-Audio: Tell Fact from Fiction

Wan Lin, Li Wang, Jindong Wang +2

The paper presents Teffic-Audio, a speech deepfake detection system that uses a Conformer-based encoder with attentive pooling and a training recipe focused on multi-source data an…

#speech deepfake detection#audio forensics#conformer encoder#data augmentation
cs.CV2026

Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification

Alexander Kozachok, Ilya Latyshev, Evgeny Karpulevich +3

The paper evaluates how different data augmentation strategies, especially photometric transformations and a mixed augmentation policy, can improve the out-of-domain robustness of…

#skin cancer classification#data augmentation#out-of-domain generalization#dermoscopic imaging
cs.AI2026

Multi-Sensor Alignment for Weather Simulations

Samsad Alam, Devyani Lambhate, Aditya Mohan +2

The paper introduces methods to align weather simulations across multiple sensors for autonomous vehicle perception, including ReDAM for fog intensity and Unified-weather-edit for…

#weather simulation#sensor alignment#autonomous vehicles#3d object detection
cs.CV2026

Simulating Automotive Radar with Lidar and Camera Inputs

Peili Song, Dezhen Song, Yifan Yang +2

The paper introduces a method that uses camera images, lidar point clouds, and ego-velocity to generate realistic 4D automotive radar signals via two neural networks (DIS‑Net and R…

#radar simulation#sensor fusion#autonomous driving#neural networks
cs.AI2026

QDA-SQL: Questions Enhanced Dialogue Augmentation for Multi-Turn Text-to-SQL

Yinggang Sun, Ziming Guo, Haining Yu +5

The paper introduces QDA-SQL, a data augmentation technique that uses large language models to generate and validate multi‑turn question‑answer pairs, improving fine‑tuned models'…

#text-to-sql#multi-turn dialogue#data augmentation#large language models
cs.RO2026

Pipette: An Embodied Simulation Platform, Benchmark, and Data-Efficient Augmentation Framework for Wet-Lab Robotics

Zhe Liu, Huanbo Jin, Zhaohui Du +10

Pipette is an embodied simulation platform that provides open-source wet‑lab assets, a benchmark of 12 robotic tasks, and a data‑efficient augmentation pipeline to turn a few human…

#wet-lab robotics#simulation platform#data augmentation#benchmark
cs.CV2026

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method

Jackie Alex, Justin Petter

The paper proposes a framework that uses knowledge embedding and a hypernetwork‑guided diffusion model to generate realistic defect images of substation meters from very few annota…

#few-shot image generation#industrial defect detection#stable diffusion#hypernetwork control
cs.LG2026

Overcoming the Modality Gap in Context-Aided Forecasting

Vincent Zhihao Zheng, Étienne Marcotte, Arjun Ashok +4

The paper introduces a semi‑synthetic data augmentation technique to create high‑quality contextual information for time‑series forecasting, producing a 7 million‑sample dataset (C…

#time series forecasting#multimodal learning#data augmentation#contextual modeling
cs.RO2026

Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments

Mingyu Liu, Zeju Li, Jiuhe Shu +4

The paper shows that smooth robot demonstrations can miss critical alignment moments, and proposes slowing down and resampling key motion segments, plus a spatio‑temporal feature c…

#imitation learning#manipulation#alignment#data augmentation
cs.CR2026

When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training

Na Li, Boyu Kuang, Hongsheng Hu +4

The paper shows that mixing text‑to‑image synthetic data with real data during training can increase privacy leakage of the real samples, and introduces a method (RSMixLeak) to mea…

#privacy#synthetic data#text-to-image generation#membership inference attack
cs.LG2026

Fast Rates for Semi-Supervised Learning via Data-Augmentation Graph Regularization

Adam M. Oberman

The paper provides a theoretical analysis showing that self‑supervised learning with data augmentation can achieve a fast O(1/n_L) error rate in semi‑supervised settings, linking t…

#semi-supervised learning#self-supervised learning#data augmentation#graph regularization
cs.CL2026

Constraint-Aware Counterfactual Editing for Aspect-Based Sentiment Analysis

S M Rafiuddin, Vamsi Krishna Pavuluri, Atriya Sen

The paper introduces CAVE-ABSA, a framework that generates and validates aspect-level counterfactual sentences for aspect‑based sentiment analysis, ensuring the target aspect’s sen…

#aspect-based sentiment analysis#counterfactual generation#constraint-aware editing#validation
cs.CV2026

Steering Diffusion Models via Class-Contrastive Influence for Few-Shot Medical Classification

Jeeyung Kim, Erfan Esmaeili, Qiang Qiu

The paper introduces Class-Contrastive Influence (C2I) to evaluate how useful diffusion‑generated images are for few‑shot medical classification, and uses reinforcement learning to…

#few-shot learning#diffusion models#medical image classification#data augmentation
cs.CV2026

MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentation

Teerath Kumar, Raja Vavekanand, Muhammad Turab

The paper introduces MedDiffuseMix, a saliency‑aware diffusion‑based augmentation method that mixes low‑importance regions of medical images while preserving diagnostically importa…

#data augmentation#diffusion models#saliency maps#medical image classification
cs.LG2026

Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs

Max Weltevrede, Caroline Horsch, Matthijs T. J. Spaan +1

The paper shows that training reinforcement‑learning agents on additional, irrelevant states acts like data augmentation and can improve zero‑shot generalization in contextual MDPs…

#contextual mdp#zero-shot policy transfer#generalization#exploration

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.