Domain-Adapted Molecular Language Models Improve Virtual Library Search Efficiency
This study assesses four pretrained molecular language models as encoders for predicting structure-property relationships, evaluating their performance across six virtual molecular libraries aimed at drug discovery, organic materials, and catalysis. Findings reveal that while native embeddings exhibit inconsistent results, molecular fingerprints establish a robust baseline. The authors suggest a mismatch in domain representation and explore domain adaptation by fine-tuning encoders on library structures, resulting in enhanced performance and better sample efficiency. Several adapted encoders surpass the fingerprint baseline. The preprint titled 'Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries' is available on arXiv. The research underscores the significance of domain adaptation and the role of classical baselines in molecular discovery processes.
Key facts
- Four molecular language models were benchmarked.
- Six virtual molecular libraries were used.
- Libraries span drug discovery, organic materials, and catalysis.
- Native embeddings show substantial performance variation across libraries.
- Molecular fingerprints provide a consistently strong and robust baseline.
- Explicit domain adaptation substantially improves representation performance.
- Fine-tuning on target library structures consistently improves sample efficiency.
- Several adapted encoders emerge as top performers.
Entities
Institutions
- arXiv