#self-supervised learning

35 results
cs.CL2026

Training Skills Like Parameters via Self-Supervised Semantic Diffusion

Mo Li, Zixin Yin, Ting Cao +1

The paper introduces a self‑supervised framework that lets a language model acquire and store textual skills in an external library using diffusion‑style reconstruction loss, witho…

#self-supervised learning#diffusion models#skill extraction#language model adaptation
cs.CV2026

PhiZero: A World Model Built Around Physical Language

Shuyao Shang, Yuqi Wang, Ruopeng Gao +4

PhiZero is a physical world model that learns a compact discrete "physical language" from videos to predict future world states as language sequences before rendering them into rea…

#physical world modeling#language-based representation#video prediction#self-supervised learning
cs.SD2026

Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances

Dmitrii Gavrilev, Ilya Borovik, Vladimir Viro

The paper proposes using contextual embeddings from self‑supervised symbolic music models to evaluate expressive MIDI piano performances, showing that these embeddings align with h…

#expressive performance evaluation#symbolic music#contextual embeddings#self-supervised learning
cs.CV2026

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

Jiwen Liu, Shujuan Li, Xiaohan Li +5

The paper introduces TARS, a 3D‑free video re‑shooting framework that uses text‑driven semantic viewpoint specifications and self‑supervised training to control camera motion and p…

#video re-shooting#camera control#text-driven viewpoint#self-supervised learning
cs.CV2026

Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures

Sweta Banerjee, Alireza Teimoury, Nils Porsche +11

The paper evaluates whether pathology foundation models can serve as effective backbones for dense detection of mitotic figures, comparing several self‑supervised models to a ResNe…

#pathology foundation models#mitotic figure detection#dense object detection#self-supervised learning
cs.LG2026

Building a User Foundation Model for the Open Web

Solal Vernier, Ivan Can Arisoy, Merwan Barlier +1

The paper introduces a self‑supervised transformer model trained on fragmented web browsing histories to create user representations that improve click prediction and bidding perfo…

#user modeling#self-supervised learning#transformer encoder#real-time bidding
cs.CV2026

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

Ionuţ Grigore, Călin-Adrian Popa

The paper introduces JEPADepth, a self‑supervised monocular depth estimation method that adds a masked predictive loss using a pretrained DINOv3 Vision Transformer encoder to the u…

#monocular depth estimation#self-supervised learning#masked predictive modeling#vision transformers
cs.LG2026

Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision

Zhiyuan Ma, Zeyuan Li, Zhiyi Lu +7

The paper introduces BridgeMIL, a two-stage method that first learns EEG instance representations without using inherited labels and then applies subject-level supervision via a mu…

#electroencephalography#disease diagnosis#multiple instance learning#self-supervised learning
cs.NI2026

WALoMA: A Multitask Wireless Foundation Model via Adaptive Low-Rank Masked Autoencoders

Madi Makin, Asmaa Abdallah, Abdulkadir Celik +1

The paper introduces WALoMA, a multitask wireless foundation model that uses adaptive low-rank masked autoencoders to learn self-supervised representations of channel state informa…

#foundation models#self-supervised learning#masked autoencoders#low-rank adaptation
cs.CV2026

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

Quoc-Huy Trinh, Minh-Van Nguyen, Ulas Bagci

The paper presents Rad-JEPA 3D, a self‑supervised joint‑embedding model that learns 3D CT representations by predicting latent features of a full scan from a masked view, using a h…

#self-supervised learning#3d medical imaging#ct scan analysis#joint embedding
cs.SD2026

Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3

The paper investigates how the language used to train neural audio codecs and self‑supervised speech models affects performance, finding that codec training language has little imp…

#self-supervised learning#neural audio codecs#speech representation#language sensitivity
cs.SD2026

MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

Scott H. Hawley

The paper introduces MIDI-RAE-JEPA, a self‑supervised model that learns hierarchical, equivariant representations of symbolic music from piano‑roll images using a Swin Transformer…

#symbolic music#self-supervised learning#equivariance#transformer encoder
cs.CV2026

Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality

Kunal Pratap Singh, Ali Garjani, Rishubh Singh +6

The paper introduces Test-Space Training, a self‑supervised approach that collects multimodal sensor data directly in a target test environment and uses cross‑modal learning to pre…

#self-supervised learning#multimodal learning#cross-modal prediction#test-time specialization
cs.RO2026

Image-to-Point Cloud Registration Made Easy with Rectified Flow-based LiDAR Upsampling

Reon Tabata, Kenji Koide, Shuji Oishi +4

The paper presents a method that converts a sparse LiDAR scan into a dense intensity image using conditional rectified flow, matches it to a camera image, and estimates the 6‑DoF p…

#image-to-point cloud registration#lidar upsampling#sensor fusion#pose estimation
cs.CV2026

The TIME Machine: On The Power of Motion for Efficient Perception

Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara

The paper introduces TIME, a video representation learned from motion point-tracks using a masked autoencoder, enabling self‑supervised, language‑free training that requires far le…

#video representation#motion modeling#self-supervised learning#masked autoencoder
cs.CV2026

Learning from Complementary Ultrasound Representations for Liver Disease Classification

Sabahattin Mert Daloglu, Gokce Bekar, Ceren Coskun +5

The paper studies whether adding physics-guided and local phase ultrasound representations to conventional B-mode images improves classification of NASH versus NAFLD, using self-su…

#ultrasound imaging#liver disease classification#self-supervised learning#graph convolutional networks
cs.RO2026

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots

Ishneet Sukhvinder Singh, Dhanoosh Pooranakumaran, Alex Nguyen +1

The paper introduces Kepler-Encoder-v0.1, a self‑supervised multimodal encoder that fuses vision, proprioception, and force/torque data into a shared latent space, enabling a visio…

#multimodal embedding#robot perception#self-supervised learning#cross-modal attention
cs.CL2026

Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

Stephen McIntosh, Reuben Smit, Daisuke Saito +2

The paper explores using dynamic time warping on self‑supervised WavLM speech representations to automatically score phonetic accuracy, rhythm, and intonation of L2 English and Jap…

#l2 speech assessment#self-supervised learning#wavlm representations#dtw alignment
cs.CV2026

EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent

Junlong Li, Junxi Li, Yuxiang Yang +3

The paper introduces EgoProceVQA, a new egocentric video question answering benchmark focused on procedural reasoning, and proposes EgoProceAgent, a self-exploring agent that learn…

#egocentric video#procedural understanding#video question answering#self-supervised learning
cs.CV2026

BenthiCat: An opti-acoustic dataset for advancing benthic classification and habitat mapping

Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles +3

The paper presents BenthiCat, a large multi‑modal dataset of side‑scan sonar tiles, bathymetric maps, and co‑registered optical images for training and benchmarking machine learnin…

#benthic classification#side-scan sonar#multimodal dataset#underwater habitat mapping
cs.LG2026

Fast Rates for Semi-Supervised Learning via Data-Augmentation Graph Regularization

Adam M. Oberman

The paper provides a theoretical analysis showing that self‑supervised learning with data augmentation can achieve a fast O(1/n_L) error rate in semi‑supervised settings, linking t…

#semi-supervised learning#self-supervised learning#data augmentation#graph regularization
cs.LG2026

Leveraging unlabelled data for generalizable neural population decoding

Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo +2

The paper presents MOJO, a framework that combines masked autoencoding self‑supervised learning with supervised training for spike‑tokenizing neural decoders, yielding better decod…

#neural decoding#self-supervised learning#masked autoencoding#spike tokenization
cs.LG2026

LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration

Jagan Mohan Reddy Dwarampudi, Veena Kochat, Suresh Satpati +3

LATTICE is a graph-based self‑supervised framework that learns spot‑level embeddings by integrating multimodal spatial omics data (RNA, ATAC, CUT&Tag) using a TransformerConv encod…

#spatial omics#multimodal integration#graph neural networks#self-supervised learning
cs.LG2026

Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis

Siwon Kim

The paper introduces ER-JEPA, a lightweight self‑supervised learning framework that builds hierarchical representations for multivariate time‑series data, demonstrated on 12‑lead E…

#self-supervised learning#time series#electrocardiogram#hierarchical models
← Prev1 / 2Next →