Skip to main navigation Skip to search Skip to main content

Machine Learning-guided directed protein evolution

Activity: Talk or presentationInvited talkscience-to-science

Description

Efficient protein engineering and optimization is essential for development of novel biotechnologies. Obtaining a protein with desired properties implies finding a solution to an optimization problem of large dimensionality. The current state-of-the-art technique for optimization of biological sequences is Directed Evolution (DE). It involves multiple rounds of selection and optionally diversification steps. This procedure imitates the process of natural selection akin to evolutionary algorithms to iteratively refine the desired properties of the sequences in the pool. Recently, machine learning methods have been proposed that intend to augment and speed up this process. These methods attempt to utilize the prior knowledge, usually coming from earlier DE rounds, to suggest the best sequences for further testing. Usually, such setups consist of a) a generative method that suggests new sequences and b) one or multiple scoring functions that assess the activity of the proteins (usually by means of predicting the performance of the sequence in prior selection rounds). We evaluated various such methods in a practical setting, e.g. regression approaches, MBE, and a discretized regression (“robust regression”). We have evaluated these approaches with respect to the AUC and ΔAUC-PR metrics during the optimization of a specific protein of interest. Additionally, we evaluate the benefits of the epistemic uncertainty estimates from the scoring functions, as it is a key element in many efficient exploration techniques. We found that robust regression performs best with noisy and irregularly distributed real world sequencing data from a DE experiment. For this kind of data, regression models succumbed to noise and outliers and were found to be unsuitable. We show benefits of alternating between in-vitro DE rounds and in-silico ML-DE rounds to speed up the optimization process and improve the resulting sequences.
Period04 Jun 2024
Event titleELLIS SYMPOSIUM on Machine Learning for Drug Discovery
Event typeWorkshop
Degree of RecognitionInternational

Fields of science

  • 101019 Stochastics
  • 102003 Image processing
  • 103029 Statistical physics
  • 101018 Statistics
  • 101017 Game theory
  • 102001 Artificial intelligence
  • 202017 Embedded systems
  • 101016 Optimisation
  • 101015 Operations research
  • 101014 Numerical mathematics
  • 101029 Mathematical statistics
  • 101028 Mathematical modelling
  • 101026 Time series analysis
  • 101024 Probability theory
  • 102032 Computational intelligence
  • 102004 Bioinformatics
  • 102013 Human-computer interaction
  • 101027 Dynamical systems
  • 305907 Medical statistics
  • 101004 Biomathematics
  • 305905 Medical informatics
  • 101031 Approximation theory
  • 102033 Data mining
  • 102 Computer Sciences
  • 305901 Computer-aided diagnosis and therapy
  • 102019 Machine learning
  • 106007 Biostatistics
  • 102018 Artificial neural networks
  • 106005 Bioinformatics
  • 202037 Signal processing
  • 202036 Sensor systems
  • 202035 Robotics

JKU Focus areas

  • Digital Transformation