An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design
arXiv:2510.18860 · doi:10.1038/s44387-026-00087-1
Abstract
Traditional machine learning has advanced polymer discovery, yet direct generation of chemically valid and synthesizable polymers without exhaustive enumeration remains a challenge. Here we present polyT5, an encoder-decoder chemical language model based on the T5 architecture, trained to understand and generate polymer structures. polyT5 enables both property prediction and the targeted generation of polymers conditioned on desired property values. We demonstrate its utility for dielectric polymer design, seeking candidates with dielectric constant >3, bandgap >4 eV, and glass transition temperature >400 K, alongside melt-processability and solubility requirements. From over 20,000 generated promising candidates, one was experimentally synthesized and validated, showing strong agreement with predictions. To further enhance usability, we integrated polyT5 within an agentic AI framework that couples it with a general-purpose LLM, allowing natural language interaction for property prediction and generative design. Together, these advances establish a versatile and accessible framework for accelerated polymer discovery.
References in corpus (7)
- Self-Referencing Embedded Strings (SELFIES): A 100% robust molecular string representation
- SELFIES and the future of molecular string representations
- polyBERT: A chemical language model to enable fully machine-driven ultrafast polymer informatics
- TransPolymer: a Transformer-based language model for polymer property predictions
- Artificial Intelligence in Materials Science and Engineering: Current Landscape, Key Challenges, and Future Trajectorie
- An overview of domain-specific foundation model: key technologies, applications and challenges
- Benchmarking Large Language Models for Polymer Property Predictions