#self-supervised learning
35 resultsTraining Skills Like Parameters via Self-Supervised Semantic Diffusion
Mo Li, Zixin Yin, Ting Cao +1
The paper introduces a self‑supervised framework that lets a language model acquire and store textual skills in an external library using diffusion‑style reconstruction loss, witho…
PhiZero: A World Model Built Around Physical Language
Shuyao Shang, Yuqi Wang, Ruopeng Gao +4
PhiZero is a physical world model that learns a compact discrete "physical language" from videos to predict future world states as language sequences before rendering them into rea…
Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances
Dmitrii Gavrilev, Ilya Borovik, Vladimir Viro
The paper proposes using contextual embeddings from self‑supervised symbolic music models to evaluate expressive MIDI piano performances, showing that these embeddings align with h…
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting
Jiwen Liu, Shujuan Li, Xiaohan Li +5
The paper introduces TARS, a 3D‑free video re‑shooting framework that uses text‑driven semantic viewpoint specifications and self‑supervised training to control camera motion and p…
Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
Sweta Banerjee, Alireza Teimoury, Nils Porsche +11
The paper evaluates whether pathology foundation models can serve as effective backbones for dense detection of mitotic figures, comparing several self‑supervised models to a ResNe…
Building a User Foundation Model for the Open Web
Solal Vernier, Ivan Can Arisoy, Merwan Barlier +1
The paper introduces a self‑supervised transformer model trained on fragmented web browsing histories to create user representations that improve click prediction and bidding perfo…
JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation
IonuÅ£ Grigore, CÄlin-Adrian Popa
The paper introduces JEPADepth, a self‑supervised monocular depth estimation method that adds a masked predictive loss using a pretrained DINOv3 Vision Transformer encoder to the u…
Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision
Zhiyuan Ma, Zeyuan Li, Zhiyi Lu +7
The paper introduces BridgeMIL, a two-stage method that first learns EEG instance representations without using inherited labels and then applies subject-level supervision via a mu…
WALoMA: A Multitask Wireless Foundation Model via Adaptive Low-Rank Masked Autoencoders
Madi Makin, Asmaa Abdallah, Abdulkadir Celik +1
The paper introduces WALoMA, a multitask wireless foundation model that uses adaptive low-rank masked autoencoders to learn self-supervised representations of channel state informa…
Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
Quoc-Huy Trinh, Minh-Van Nguyen, Ulas Bagci
The paper presents Rad-JEPA 3D, a self‑supervised joint‑embedding model that learns 3D CT representations by predicting latent features of a full scan from a masked view, using a h…
Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3
The paper investigates how the language used to train neural audio codecs and self‑supervised speech models affects performance, finding that codec training language has little imp…
MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
Scott H. Hawley
The paper introduces MIDI-RAE-JEPA, a self‑supervised model that learns hierarchical, equivariant representations of symbolic music from piano‑roll images using a Swin Transformer…
Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality
Kunal Pratap Singh, Ali Garjani, Rishubh Singh +6
The paper introduces Test-Space Training, a self‑supervised approach that collects multimodal sensor data directly in a target test environment and uses cross‑modal learning to pre…
Image-to-Point Cloud Registration Made Easy with Rectified Flow-based LiDAR Upsampling
Reon Tabata, Kenji Koide, Shuji Oishi +4
The paper presents a method that converts a sparse LiDAR scan into a dense intensity image using conditional rectified flow, matches it to a camera image, and estimates the 6‑DoF p…
The TIME Machine: On The Power of Motion for Efficient Perception
Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
The paper introduces TIME, a video representation learned from motion point-tracks using a masked autoencoder, enabling self‑supervised, language‑free training that requires far le…
Learning from Complementary Ultrasound Representations for Liver Disease Classification
Sabahattin Mert Daloglu, Gokce Bekar, Ceren Coskun +5
The paper studies whether adding physics-guided and local phase ultrasound representations to conventional B-mode images improves classification of NASH versus NAFLD, using self-su…
Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots
Ishneet Sukhvinder Singh, Dhanoosh Pooranakumaran, Alex Nguyen +1
The paper introduces Kepler-Encoder-v0.1, a self‑supervised multimodal encoder that fuses vision, proprioception, and force/torque data into a shared latent space, enabling a visio…
Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring
Stephen McIntosh, Reuben Smit, Daisuke Saito +2
The paper explores using dynamic time warping on self‑supervised WavLM speech representations to automatically score phonetic accuracy, rhythm, and intonation of L2 English and Jap…
EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent
Junlong Li, Junxi Li, Yuxiang Yang +3
The paper introduces EgoProceVQA, a new egocentric video question answering benchmark focused on procedural reasoning, and proposes EgoProceAgent, a self-exploring agent that learn…
BenthiCat: An opti-acoustic dataset for advancing benthic classification and habitat mapping
Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles +3
The paper presents BenthiCat, a large multi‑modal dataset of side‑scan sonar tiles, bathymetric maps, and co‑registered optical images for training and benchmarking machine learnin…
Fast Rates for Semi-Supervised Learning via Data-Augmentation Graph Regularization
Adam M. Oberman
The paper provides a theoretical analysis showing that self‑supervised learning with data augmentation can achieve a fast O(1/n_L) error rate in semi‑supervised settings, linking t…
Leveraging unlabelled data for generalizable neural population decoding
Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo +2
The paper presents MOJO, a framework that combines masked autoencoding self‑supervised learning with supervised training for spike‑tokenizing neural decoders, yielding better decod…
LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration
Jagan Mohan Reddy Dwarampudi, Veena Kochat, Suresh Satpati +3
LATTICE is a graph-based self‑supervised framework that learns spot‑level embeddings by integrating multimodal spatial omics data (RNA, ATAC, CUT&Tag) using a TransformerConv encod…
Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis
Siwon Kim
The paper introduces ER-JEPA, a lightweight self‑supervised learning framework that builds hierarchical representations for multivariate time‑series data, demonstrated on 12‑lead E…