On the Opportunities and Risks of Foundation Models
arXiv:2108.07258
Abstract
AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks. We call these models foundation models to underscore their critically central yet incomplete character. This report provides a thorough account of the opportunities and risks of foundation models, ranging from their capabilities (e.g., language, vision, robotics, reasoning, human interaction) and technical principles(e.g., model architectures, training procedures, data, systems, security, evaluation, theory) to their applications (e.g., law, healthcare, education) and societal impact (e.g., inequity, misuse, economic and environmental impact, legal and ethical considerations). Though foundation models are based on standard deep learning and transfer learning, their scale results in new emergent capabilities,and their effectiveness across so many tasks incentivizes homogenization. Homogenization provides powerful leverage but demands caution, as the defects of the foundation model are inherited by all the adapted models downstream. Despite the impending widespread deployment of foundation models, we currently lack a clear understanding of how they work, when they fail, and what they are even capable of due to their emergent properties. To tackle these questions, we believe much of the critical research on foundation models will require deep interdisciplinary collaboration commensurate with their fundamentally sociotechnical nature.
Authored by the Center for Research on Foundation Models (CRFM) at the Stanford Institute for Human-Centered Artificial Intelligence (HAI). Report page with citation guidelines: https://crfm.stanford.edu/report.html
References in corpus (156)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Distilling the Knowledge in a Neural Network
- Learning Transferable Visual Models From Natural Language Supervision
- WaveNet: A Generative Model for Raw Audio
- Bootstrap your own latent: A new approach to self-supervised Learning
- Towards A Rigorous Science of Interpretable Machine Learning
- Language Models are Few-Shot Learners
- LoRA: Low-Rank Adaptation of Large Language Models
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Scaling Laws for Neural Language Models
- Learning agile and dynamic motor skills for legged robots
- Evaluating Large Language Models Trained on Code
- MLP-Mixer: An all-MLP Architecture for Vision
- Frustratingly Easy Domain Adaptation
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
- NICE: Non-linear Independent Components Estimation
- EfficientNetV2: Smaller Models and Faster Training
- Zero-Shot Text-to-Image Generation
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Linformer: Self-Attention with Linear Complexity
- Snorkel: Rapid Training Data Creation with Weak Supervision
- Poisoning Attacks against Support Vector Machines
- Solving Rubik's Cube with a Robot Hand
- Multilingual Denoising Pre-training for Neural Machine Translation
- Conservative Q-Learning for Offline Reinforcement Learning
- REALM: Retrieval-Augmented Language Model Pre-Training
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Generating Long Sequences with Sparse Transformers
- Decision Transformer: Reinforcement Learning via Sequence Modeling
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- Deep Learning Scaling is Predictable, Empirically
- Massively Multitask Networks for Drug Discovery
- Do ImageNet Classifiers Generalize to ImageNet?
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
- Reformer: The Efficient Transformer
- Towards a Critical Race Methodology in Algorithmic Fairness
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- One Model To Learn Them All
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- The Evolved Transformer
- True Few-Shot Learning with Language Models
- Unifying Vision-and-Language Tasks via Text Generation
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- Unsupervised Domain Adaptation through Self-Supervision
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- Predictive Inequity in Object Detection
- Learning and Evaluating General Linguistic Intelligence
- Data Shapley: Equitable Valuation of Data for Machine Learning
- Scaling Laws for Autoregressive Generative Modeling
- Time-Aware Language Models as Temporal Knowledge Bases
- Algorithmic Monoculture and Social Welfare
- Quantifying the Carbon Emissions of Machine Learning
- Carbon Emissions and Large Neural Network Training
- Artificial Intelligence in Clinical Health Care Applications: Viewpoint
- Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
- Multilingual is not enough: BERT for Finnish
- Rethinking Attention with Performers
- Carbontracker: Tracking and Predicting the Carbon Footprint of Training Deep Learning Models
- Pretrained Transformers as Universal Computation Engines
- CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review
- Unsolved Problems in ML Safety
- PatentBERT: Patent Classification with Fine-Tuning a pre-trained BERT Model
- Multimodal Few-Shot Learning with Frozen Language Models
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled Data
- Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
- An Investigation of Why Overparameterization Exacerbates Spurious Correlations
- Query2box: Reasoning over Knowledge Graphs in Vector Space using Box Embeddings
- Keeping Community in the Loop: Understanding Wikipedia Stakeholder Values for Machine Learning-Based Systems
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- From Language to Goals: Inverse Reinforcement Learning for Vision-Based Instruction Following
- Language-Agnostic Representation Learning of Source Code from Structure and Context
- Data and its (dis)contents: A survey of dataset development and use in machine learning research
- Noise or Signal: The Role of Image Backgrounds in Object Recognition
- ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
- MERLOT: Multimodal Neural Script Knowledge Models
- Understanding the Failure Modes of Out-of-Distribution Generalization
- QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering
- Graph-based, Self-Supervised Program Repair from Diagnostic Feedback
- The Role of Cooperation in Responsible AI Development
- Energy Usage Reports: Environmental awareness as part of algorithmic accountability
- Characterising Bias in Compressed Models
- Context-Aware Legal Citation Recommendation using Deep Learning
- Model Patching: Closing the Subgroup Performance Gap with Data Augmentation
- Generative Language Modeling for Automated Theorem Proving
- Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation
- Modifying Memories in Transformer Models
- Towards Ecologically Valid Research on Language User Interfaces
- Alignment of Language Agents
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss
- On Adversarial Bias and the Robustness of Fair Machine Learning
- Learning to summarize from human feedback
- Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus
- BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments
- MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale
- Understanding Self-supervised Learning with Dual Deep Networks
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt Tuning
- BERT Goes to Law School: Quantifying the Competitive Advantage of Access to Large Legal Corpora in Contract Understanding
- Pretraining Representations for Data-Efficient Reinforcement Learning
- HTLM: Hyper-Text Pre-Training and Prompting of Language Models
- End-to-End Robotic Reinforcement Learning without Reward Engineering
- Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization
- Break-It-Fix-It: Unsupervised Learning for Program Repair
- Contrastive Learning Inverts the Data Generating Process
- Combiner: Full Attention Transformer with Sparse Computation Cost
- A Benchmark for Lease Contract Review
- Facts as Experts: Adaptable and Interpretable Neural Memory over Symbolic Knowledge
- A Neural Topic-Attention Model for Medical Term Abbreviation Disambiguation
- How much is Wikipedia Lagging Behind News?
- Poisoning the Unlabeled Dataset of Semi-Supervised Learning
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning
- Top-KAST: Top-K Always Sparse Training
- Revealing Persona Biases in Dialogue Systems
- Poisoning and Backdooring Contrastive Learning
- BREEDS: Benchmarks for Subpopulation Shift
- Contrastive estimation reveals topic posterior information to linear models
- The iWildCam 2021 Competition Dataset
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- CraftAssist: A Framework for Dialogue-enabled Interactive Agents
- C5T5: Controllable Generation of Organic Molecules with Transformers
- Fast, Structured Clinical Documentation via Contextual Autocomplete
- Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling
- Affirmative Algorithms: The Legal Grounds for Fairness as Awareness
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
- OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning
- HULK: An Energy Efficiency Benchmark Platform for Responsible Natural Language Processing
- When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset
- ESR: Ethics and Society Review of Artificial Intelligence Research
- MathBERT: A Pre-trained Language Model for General NLP Tasks in Mathematics Education
- Persuasive Natural Language Generation -- A Literature Review
- Domain-Specific Pretraining for Vertical Search: Case Study on Biomedical Literature
- Interpretable Multi-Step Reasoning with Knowledge Extraction on Complex Healthcare Question Answering
- Mandoline: Model Evaluation under Distribution Shift
- Fair Machine Learning Under Partial Compliance
- The Benchmark Lottery
- ProtoTransformer: A Meta-Learning Approach to Providing Student Feedback
- A Theory of Label Propagation for Subpopulation Shift
- On the proper role of linguistically-oriented deep net analysis in linguistic theorizing
- Label Noise SGD Provably Prefers Flat Global Minimizers
- Benchmarking Differential Privacy and Federated Learning for BERT Models
- DABS: A Domain-Agnostic Benchmark for Self-Supervised Learning
- Proof Artifact Co-training for Theorem Proving with Language Models
- Causal Analysis of Agent Behavior for AI Safety
- Performance in the Courtroom: Automated Processing and Visualization of Appeal Court Decisions in France
- Impossibility results for fair representations
- Using Transformers to Provide Teachers with Personalized Feedback on their Classroom Discourse: The TalkMoves Application
- Training a First-Order Theorem Prover from Synthetic Data
- Distributed Deep Learning in Open Collaborations
- Cross-neutralising: Probing for joint encoding of linguistic information in multilingual models
- When to reply? Context Sensitive Models to Predict Instructor Interventions in MOOC Forums
- Feedback Effects in Repeat-Use Criminal Risk Assessments
- How could Neural Networks understand Programs?
- The Law of Large Documents: Understanding the Structure of Legal Contracts Using Visual Cues
Cited by in corpus (194)
- Learning to Prompt for Vision-Language Models
- Generative AI
- DINOv2: Learning Robust Visual Features without Supervision
- SpectralGPT: Spectral Remote Sensing Foundation Model
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- The Ethics of ChatGPT in Medicine and Healthcare: A Systematic Review on Large Language Models (LLMs)
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- Self-Supervised Speech Representation Learning: A Review
- Large language models in medicine: the potentials and pitfalls
- Sustainable AI: Environmental Implications, Challenges and Opportunities
- Structured Pruning for Deep Convolutional Neural Networks: A survey
- Florence: A New Foundation Model for Computer Vision
- Human heuristics for AI-generated language are flawed
- TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting
- Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
- The Programmer's Assistant: Conversational Interaction with a Large Language Model for Software Development
- The Metacognitive Demands and Opportunities of Generative AI
- Co-Writing with Opinionated Language Models Affects Users' Views
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs
- Compute and Energy Consumption Trends in Deep Learning Inference
- Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
- A Comprehensive Survey of Deep Transfer Learning for Anomaly Detection in Industrial Time Series: Methods, Applications, and Directions
- A Taxonomy of Prompt Modifiers for Text-To-Image Generation
- Geo-knowledge-guided GPT models improve the extraction of location descriptions from disaster-related social media messages
- Multimodal Data Integration for Oncology in the Era of Deep Neural Networks: A Review
- Auditing of AI: Legal, Ethical and Technical Approaches
- A Survey on Unsupervised Anomaly Detection Algorithms for Industrial Images
- Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- Harms from Increasingly Agentic Algorithmic Systems
- Machine Culture
- FAIR for AI: An interdisciplinary and international community building perspective
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- Unsolved Problems in ML Safety
- A Study on the Implementation of Generative AI Services Using an Enterprise Data-Based LLM Application Architecture
- General Purpose Artificial Intelligence Systems (GPAIS): Properties, Definition, Taxonomy, Societal Implications and Responsible Governance
- Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation
- Large Language Models Can Be Strong Differentially Private Learners
- Addressing Bias in Generative AI: Challenges and Research Opportunities in Information Management
- Physics-informed machine learning for building performance simulation-A review of a nascent field
- Finetuned Language Models Are Zero-Shot Learners
- Learning from models beyond fine-tuning
- Merak: An Efficient Distributed DNN Training Framework with Automated 3D Parallelism for Giant Foundation Models
- AstroCLIP: A Cross-Modal Foundation Model for Galaxies
- Domain Adversarial Spatial-Temporal Network: A Transferable Framework for Short-term Traffic Forecasting across Cities
- Exploring the Role of AI Assistants in Computer Science Education: Methods, Implications, and Instructor Perspectives
- Anatomy of an AI-powered malicious social botnet
- MERLOT: Multimodal Neural Script Knowledge Models
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine Learning
- Reinforcement Learning for Generative AI: State of the Art, Opportunities and Open Research Challenges
- Hyperspectral unmixing for Raman spectroscopy via physics-constrained autoencoders
- Fine-tuning and Utilization Methods of Domain-specific LLMs
- The Impact of ChatGPT and LLMs on Medical Imaging Stakeholders: Perspectives and Use Cases
- ERNIE-GeoL: A Geography-and-Language Pre-trained Model and its Applications in Baidu Maps
- Debiasing Methods for Fairer Neural Models in Vision and Language Research: A Survey
- FoodSAM: Any Food Segmentation
- Exploring Adapter-based Transfer Learning for Recommender Systems: Empirical Studies and Practical Insights
- A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models
- Multimodal Foundation Models for Material Property Prediction and Discovery
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Transfer Language Selection for Zero-Shot Cross-Lingual Abusive Language Detection
- The worst of both worlds: A comparative analysis of errors in learning from data in psychology and machine learning
- Pre-trained Language Models in Biomedical Domain: A Systematic Survey
- The Survey on Multi-Source Data Fusion in Cyber-Physical-Social Systems:Foundational Infrastructure for Industrial Metaverses and Industries 5.0
- Normative Challenges of Risk Regulation of Artificial Intelligence and Automated Decision-Making
- Zero-shot cross-lingual transfer language selection using linguistic similarity
- Self-supervised representations in speech-based depression detection
- The Clever Hans Effect in Unsupervised Learning
- On the design space between molecular mechanics and machine learning force fields
- Message Ritual: A Posthuman Account of Living with Lamp
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation Models
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- Dual Use Concerns of Generative AI and Large Language Models
- Domain-specific ChatBots for Science using Embeddings
- Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Federated Few-Shot Learning for Mobile NLP
- FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation Models
- Opportunities for Large Language Models and Discourse in Engineering Design
- PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
- On the Sample Complexity of Quantum Boltzmann Machine Learning
- Building Privacy-Preserving and Secure Geospatial Artificial Intelligence Foundation Models
- Probing Speech Emotion Recognition Transformers for Linguistic Knowledge
- Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
- RadFusion: Benchmarking Performance and Fairness for Multimodal Pulmonary Embolism Detection from CT and EHR
- Leveraging large language models for structured information extraction from pathology reports
- Embedded Visual Prompt Tuning
- Toward parallel intelligence: an interdisciplinary solution for complex systems
- Probing transfer learning with a model of synthetic correlated datasets
- Will Code Remain a Relevant User Interface for End-User Programming with Generative AI Models?
- The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models
- Welcome Your New AI Teammate: On Safety Analysis by Leashing Large Language Models
- Perceptron Theory Can Predict the Accuracy of Neural Networks
- M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining
- ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
- Mapping the Landscape of Generative AI in Network Monitoring and Management
- Mapping the individual, social, and biospheric impacts of Foundation Models
- Jack and Masters of all Trades: One-Pass Learning Sets of Model Sets From Large Pre-Trained Models
- Foundation Models and Transformers for Anomaly Detection: A Survey
- SeisT: A foundational deep learning model for earthquake monitoring tasks
- Perceived Trustworthiness of Natural Language Generators
- LLM-Mediated Domain-Specific Voice Agents: The Case of TextileBot
- Findings of the The RuATD Shared Task 2022 on Artificial Text Detection in Russian
- Low-resource finetuning of foundation models beats state-of-the-art in histopathology
- Engineering Artificial Intelligence: Framework, Challenges, and Future Direction
- How Will It Drape Like? Capturing Fabric Mechanics from Depth Images
- Rethinking model prototyping through the MedMNIST+ dataset collection
- Tucano: Advancing Neural Text Generation for Portuguese
- What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure
- Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs
- Differentiable Physics: A Position Piece
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Perturbing Attention Gives You More Bang for the Buck: Subtle Imaging Perturbations That Efficiently Fool Customized Diffusion Models
- Towards Symbolic XAI -- Explanation Through Human Understandable Logical Relationships Between Features
- Emergency Department Decision Support using Clinical Pseudo-notes
- INTERN: A New Learning Paradigm Towards General Vision
- FAIR AI Models in High Energy Physics
- Undesirable Biases in NLP: Addressing Challenges of Measurement
- Towards Foundation Models for Materials Science: The Open MatSci ML Toolkit
- Practically implementing an LLM-supported collaborative vulnerability remediation process: a team-based approach
- Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis
- Intent Tagging: Exploring Micro-Prompting Interactions for Supporting Granular Human-GenAI Co-Creation Workflows
- Towards Foundation Model for Chemical Reactor Modeling: Meta-Learning with Physics-Informed Adaptation
- Parameter-Efficient Learning for Text-to-Speech Accent Adaptation
- CoLLIE: Continual Learning of Language Grounding from Language-Image Embeddings
- ForgetMe: Evaluating Selective Forgetting in Generative Models
- Distribution-based Emotion Recognition in Conversation
- Adapting an ASR Foundation Model for Spoken Language Assessment
- Analytical Engines With Context-Rich Processing: Towards Efficient Next-Generation Analytics
- Evaluating the Fairness of Discriminative Foundation Models in Computer Vision
- Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for Conversations
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- Should attention be all we need? The epistemic and ethical implications of unification in machine learning
- Shaking the foundations: delusions in sequence models for interaction and control
- LLM-Assisted Visual Analytics: Opportunities and Challenges
- Preemptively Pruning Clever-Hans Strategies in Deep Neural Networks
- Adaptive Intelligence: leveraging insights from adaptive behavior in animals to build flexible AI systems
- How to Estimate Model Transferability of Pre-Trained Speech Models?
- TADA: Task-Agnostic Dialect Adapters for English
- Relational Programming with Foundation Models
- Tissue Concepts: supervised foundation models in computational pathology
- A Survey of Spatio-Temporal EEG data Analysis: from Models to Applications
- Solving Probability and Statistics Problems by Program Synthesis
- Saturn Platform: Foundation Model Operations and Generative AI for Financial Services
- Stuck-at Faults in ReRAM Neuromorphic Circuit Array and their Correction through Machine Learning
- Generalized Decision Transformer for Offline Hindsight Information Matching
- MaskTerial: A Foundation Model for Automated 2D Material Flake Detection
- A Scalable Model Specialization Framework for Training and Inference using Submodels and its Application to Speech Model Personalization
- Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches
- SMILE: Evaluation and Domain Adaptation for Social Media Language Understanding
- Towards Collaborative Plan Acquisition through Theory of Mind Modeling in Situated Dialogue
- 3D Object Detection and High-Resolution Traffic Parameters Extraction Using Low-Resolution LiDAR Data
- ESG Accountability Made Easy: DocQA at Your Service
- Optimizing Context-Enhanced Relational Joins
- Fine-tuning Strategies for Domain Specific Question Answering under Low Annotation Budget Constraints
- AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
- On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models
- Language Models as a Knowledge Source for Cognitive Agents
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- Insulin Resistance Prediction From Wearables and Routine Blood Biomarkers
- A Good Foundation is Worth Many Labels: Label-Efficient Panoptic Segmentation
- On the attribution of confidence to large language models
- Bridging Text and Crystal Structures: Literature-driven Contrastive Learning for Materials Science
- Utilizing Grounded SAM for self-supervised frugal camouflaged human detection
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- Design considerations for a hierarchical semantic compositional framework for medical natural language understanding
- Towards More Robust NLP System Evaluation: Handling Missing Scores in Benchmarks
- How to Compliment a Human -- Designing Affective and Well-being Promoting Conversational Things
- Risks of AI Foundation Models in Education
- One to Transfer All: A Universal Transfer Framework for Vision Foundation Model with Few Data
- Measure Twice, Cut Once: Quantifying Bias and Fairness in Deep Neural Networks
- Cross-Lingual Language Model Meta-Pretraining
- Reconstructing hadronically decaying tau leptons with a jet foundation model
- A Systematic Review of Open Datasets Used in Text-to-Image (T2I) Gen AI Model Safety
- Beyond Partisan Leaning: A Comparative Analysis of Political Bias in Large Language Models
- PØDA: Prompt-driven Zero-shot Domain Adaptation
- Toward Foundation Models for Earth Monitoring: Proposal for a Climate Change Benchmark
- 10 Security and Privacy Problems in Large Foundation Models
- An argument for the impossibility of machine intelligence
- CrossedWires: A Dataset of Syntactically Equivalent but Semantically Disparate Deep Learning Models
- ÚFAL at MultiLexNorm 2021: Improving Multilingual Lexical Normalization by Fine-tuning ByT5
- Bag-of-Vectors Autoencoders for Unsupervised Conditional Text Generation
- Accelerating Deep Learning with Dynamic Data Pruning
- Active Learning at the ImageNet Scale
- The TYC Dataset for Understanding Instance-Level Semantics and Motions of Cells in Microstructures
- Collaborating Foundation Models for Domain Generalized Semantic Segmentation
- Radio Galaxy Zoo: Morphological classification by Fanaroff-Riley designation using self-supervised pre-training
- Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- A Novel Information-Theoretic Objective to Disentangle Representations for Fair Classification
- Efficient Rotation Invariance in Deep Neural Networks through Artificial Mental Rotation
- Who Decides if AI is Fair? The Labels Problem in Algorithmic Auditing