BayBE one more time
Hackathon project — Bayesian optimisation with BayBE to screen small-molecule corrosion inhibitors for aluminium alloys, including transfer learning between alloys.
Code: GitHub · Video: YouTube · Event: Bayesian Optimization Hackathon for Chemistry and Materials (Acceleration Consortium × Merck KGaA, March 2024) · Team: Surface Science Syndicate
Finding a good corrosion inhibitor means testing many candidate molecules across many conditions, and each test is a real electrochemical experiment. We asked how much of that screening Bayesian optimisation can save. We used BayBE, Merck’s open-source Bayesian optimisation library, on a published database of electrochemical responses of small organic molecules on aluminium and five aluminium alloys (AA1000, AA2024, AA5000, AA6000, AA7075).
Setup
- Search space: the inhibitor molecule (as SMILES), exposure time, pH, inhibitor concentration and salt concentration, with inhibition efficiency as the target to maximise.
- Molecular encodings compared: one-hot, Mordred descriptors, RDKit descriptors and Morgan fingerprints, each against a random-sampling baseline. This tests whether telling the optimiser something about chemistry actually helps.
- Protocol: simulated campaigns against the measured data, with 50 experiments each, one experiment per round, and 10 Monte Carlo repetitions.
- Transfer learning: a campaign on AA2024 seeded with prior data from AA1000, using BayBE’s task parameters, compared with a campaign starting from scratch.
What we found
- On AA2024, every strategy reached near-maximal efficiency within roughly 20 experiments, including random sampling. That says as much about the dataset, which contains many strong inhibitors, as about the optimiser. Morgan fingerprints were the slowest to get going.
- Transfer learning paid off. Seeding the AA2024 campaign with AA1000 data reached about 96% efficiency by the 12th experiment. The fresh campaign was still at about 89% after 25.
- The main lesson: benchmark against random sampling, and pick datasets where the optimum is actually hard to find. Otherwise an easy benchmark can make any optimiser look good.
The hackathon’s outcomes, including this project, are summarised in the event paper on ChemRxiv.
Stack: BayBE · RDKit · Mordred · pandas · Jupyter
Data: T. L. P. Galvão et al., CORDATA, npj Mater. Degrad. 6, 48 (2022); C. Özkan et al., npj Mater. Degrad. 8, 21 (2024).