#representation learning
32 papers match
Explaining Image Similarity with Automatically Extracted Concept Activation Vectors
Isaac Roberts, Petra Bevandic, Alexander Schulz +1
The paper proposes a model‑agnostic method that uses automatically discovered concept activation vectors to explain why two images are considered similar, by perturbing embeddings…
Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories
Mengfei Ran, Yifeng Shen, Ruijie Guan
The paper introduces a method called Doubly Robust Functional Representation Learning (DR-FRL) that transforms irregular, time‑varying data into targeted representations for longit…
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
Yanshi Li, Xueru Bai, Shuman Liu +1
The paper introduces RepBench, a framework that aggregates thousands of benchmark datasets into a large set of probe texts to evaluate capability-aligned representations of large l…
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
Kaizhen Tan, Xin Xu, Siru Tao +4
The paper investigates which physical properties (mass, drag, stiffness) are encoded in latent world models by using controlled interventions in a simulated multimodal environment…
Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation
Jialu Xu, Mengkun Liang, Guannan Liu +2
The paper introduces CURL, a method that uses uncertainty estimates to guide a frozen large language model in creating semantic representations for covariates, improving heterogene…
Sky sphere representation in language models
Aleksandr Berdnikov, Yevgeny Liokumovich
The paper investigates whether large language models (~100B parameters) contain a decodable representation of the night sky map within their residual streams, showing that most exa…
Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision
Zhiyuan Ma, Zeyuan Li, Zhiyi Lu +7
The paper introduces BridgeMIL, a two-stage method that first learns EEG instance representations without using inherited labels and then applies subject-level supervision via a mu…
Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance
Gaspard Lambrechts, Adrien Bolland, Daniel Ebi +1
The paper introduces Reinforced Dreamer, an asymmetric model‑based reinforcement learning algorithm that uses latent guidance to improve representation learning from privileged inf…
IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment
Xinran Liu, Shengtao Li, Shouqian Shi +2
The paper introduces IRIS, a training-free method that uses frozen large language models to generate stable identity signatures for entities, enabling direct similarity-based align…
Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
Jiaxin Bai, Jiaxuan Xiong
The paper introduces Temporal-Distance JEPA, a method that learns a directed temporal cost from offline trajectories to improve latent world model predictive control, enhancing pla…
Probing Spatial Structure in Pretrained Audio Representations
Chuyang Chen, Sivan Ding, Adrian S. Roman +1
The paper introduces SARL, a benchmark for evaluating how pretrained audio models encode spatial information such as source direction and room acoustics, and analyzes the strengths…
AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization
Yiyang Yao, Shanglin Liu, Jianming Lv +4
The paper introduces AspectCLIP, a method that groups image-text pairs by shared textual aspects and applies consistency regularization within these groups to avoid forcing unrelat…
Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality
Kunal Pratap Singh, Ali Garjani, Rishubh Singh +6
The paper introduces Test-Space Training, a self‑supervised approach that collects multimodal sensor data directly in a target test environment and uses cross‑modal learning to pre…
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models
Yufeng Ji, Wenhao Tang, Haoyi Niu +3
The paper introduces Action QFormer, a query-based interface that reorganizes multimodal information into action-focused representations to improve vision-language-action models, e…
Contrastive Conformal Sets
Yahya Alkhatib, Wee Peng Tay
The paper introduces a method that combines contrastive learning with conformal prediction to create learnable geometric sets that guarantee a user‑specified coverage of positive s…
Comparison of Dimension Reduction Methods for EEG Seizure Detection Using Autonomous AI-Driven Optimization
Annika Stiehl, Vishal Kagade, Nicolas Weeger +3
The paper compares four dimension‑reduction techniques for multichannel EEG seizure detection and uses an autonomous AI framework to jointly optimize the representation and deep‑le…
Factorized Spectral Representations for Reinforcement Learning
Junyi Wu, Dan Li
The paper introduces FaStR, a method that factorizes the transition kernel of a reinforcement learning environment as a three-way tensor using CP decomposition, learning separate e…
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
Zhihao Xie, Junfeng Wu, Xinting Hu +2
VideoRAE is a representation autoencoder that leverages frozen video foundation model features to create compact, generation‑friendly video latents, supporting both continuous diff…
Prime Fourier Embeddings: A Principled Basis for Modular Arithmetic
Hyunsang Hwang, Suhyun Bae, Donghun Lee
The paper proposes Prime Fourier Embeddings, a sinusoidal encoding of integers based on prime indices that makes modular arithmetic operations explicit, and shows theoretically tha…
Not Only NTP: Extending Training Signal Coverage for Generative Recommendation
Changhao Li, Shuli Wang, Junwei Yin +6
The paper introduces NONTP, a method that augments next‑token prediction for recommendation models with temporal contrastive learning and trans‑domain learning to capture longer‑ra…
UMSS: Towards Unsupervised Multi-modal Semantic Segmentation
Haitian Zhang, Thai Duy Nguyen, Xiangyuan Wang +2
The paper introduces UniM2, an unsupervised framework for multimodal semantic segmentation that learns a shared latent space across sensors using cross‑modal correspondence and a h…
Contrastive-Augmented Flow Matching for Style-Content Disentanglement
Yusong Li, Pingchuan Ma, Ming Gui +2
The paper proposes Contrastive Augmented Flow Matching (CAtFM), a method that adds contrastive regularization to invertible flow matching to learn disentangled content and style re…
Beyond Perceptual Distance: Discrepancy Assessment on Deep Representation for Out-of-Distribution Detection with Diffusion Model
Kun Fang, Zuopeng Yang, Haibo Hu +3
The paper introduces DDR, a method that evaluates the difference between an input image and its diffusion‑model reconstruction using the classifier’s deep feature and logit represe…
Contrastive-Collapsed Loss for Flexible and Geometrically Optimal Embeddings and Faster Convergence
Blanca Cano-Camarero, Ãngela Fernández-Pascual, José R. Dorronsoro
The paper proposes CoCo, a contrastive-collapsed loss that encourages intra‑class collapse and inter‑class contrast to produce normalized, geometrically optimal embeddings with lar…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.