Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Machine Learning-guided directed protein evolution

Aktivität: Vortrag oder PräsentationEingeladener VortragScience-to-science

Beschreibung

Efficient protein engineering and optimization is essential for development of novel biotechnologies. Obtaining a protein with desired properties implies finding a solution to an optimization problem of large dimensionality. The current state-of-the-art technique for optimization of biological sequences is Directed Evolution (DE). It involves multiple rounds of selection and optionally diversification steps. This procedure imitates the process of natural selection akin to evolutionary algorithms to iteratively refine the desired properties of the sequences in the pool. Recently, machine learning methods have been proposed that intend to augment and speed up this process. These methods attempt to utilize the prior knowledge, usually coming from earlier DE rounds, to suggest the best sequences for further testing. Usually, such setups consist of a) a generative method that suggests new sequences and b) one or multiple scoring functions that assess the activity of the proteins (usually by means of predicting the performance of the sequence in prior selection rounds). We evaluated various such methods in a practical setting, e.g. regression approaches, MBE, and a discretized regression (“robust regression”). We have evaluated these approaches with respect to the AUC and ΔAUC-PR metrics during the optimization of a specific protein of interest. Additionally, we evaluate the benefits of the epistemic uncertainty estimates from the scoring functions, as it is a key element in many efficient exploration techniques. We found that robust regression performs best with noisy and irregularly distributed real world sequencing data from a DE experiment. For this kind of data, regression models succumbed to noise and outliers and were found to be unsuitable. We show benefits of alternating between in-vitro DE rounds and in-silico ML-DE rounds to speed up the optimization process and improve the resulting sequences.
Zeitraum04 Juni 2024
EreignistitelELLIS SYMPOSIUM on Machine Learning for Drug Discovery
VeranstaltungstypWorkshop
BekanntheitsgradInternational

Wissenschaftszweige

  • 101019 Stochastik
  • 102003 Bildverarbeitung
  • 103029 Statistische Physik
  • 101018 Statistik
  • 101017 Spieltheorie
  • 102001 Artificial Intelligence
  • 202017 Embedded Systems
  • 101016 Optimierung
  • 101015 Operations Research
  • 101014 Numerische Mathematik
  • 101029 Mathematische Statistik
  • 101028 Mathematische Modellierung
  • 101026 Zeitreihenanalyse
  • 101024 Wahrscheinlichkeitstheorie
  • 102032 Computational Intelligence
  • 102004 Bioinformatik
  • 102013 Human-Computer Interaction
  • 101027 Dynamische Systeme
  • 305907 Medizinische Statistik
  • 101004 Biomathematik
  • 305905 Medizinische Informatik
  • 101031 Approximationstheorie
  • 102033 Data Mining
  • 102 Informatik
  • 305901 Computerunterstützte Diagnose und Therapie
  • 102019 Machine Learning
  • 106007 Biostatistik
  • 102018 Künstliche Neuronale Netze
  • 106005 Bioinformatik
  • 202037 Signalverarbeitung
  • 202036 Sensorik
  • 202035 Robotik

JKU-Schwerpunkte

  • Digital Transformation