Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Weighted similarity-based clustering of chemical structures and bioactivity data in early drug discovery

  • Nolen Perualila-Tan
  • , Ziv Shkedy
  • , Willem Talloen
  • , Hinrich W.H. Göhlmann
  • , QSTAR Consortium
  • , Adetayo Kasim
  • , Marijk Van Moerbeke

Publikation: Beitrag in FachzeitschriftArtikelBegutachtung

Abstract

The modern process of discovering candidate molecules in early drug discovery phase includes a wide range of approaches to extract vital information from the intersection of biology and chemistry. A typical strategy in compound selection involves compound clustering based on chemical similarity to obtain representative chemically diverse compounds (not incorporating potency information). In this paper, we propose an integrative clustering approach that makes use of both biological (compound efficacy) and chemical (structural features) data sources for the purpose of discovering a subset of compounds with aligned structural and biological properties. The datasets are integrated at the similarity level by assigning complementary weights to produce a weighted similarity matrix, serving as a generic input in any clustering algorithm. This new analysis work flow is semi-supervised method since, after the determination of clusters, a secondary analysis is performed wherein it finds differentially expressed genes associated to the derived integrated cluster(s) to further explain the compound-induced biological effects inside the cell. In this paper, datasets from two drug development oncology projects are used to illustrate the usefulness of the weighted similarity-based clustering approach to integrate multi-source high-dimensional information to aid drug discovery. Compounds that are structurally and biologically similar to the reference compounds are discovered using this proposed integrative approach.
OriginalspracheEnglisch
Aufsatznummer1650018
Seitenumfang22
FachzeitschriftJournal of Bioinformatics and Computational Biology
Volume14
Ausgabenummer4
DOIs
PublikationsstatusVeröffentlicht - 01 Aug. 2016

Wissenschaftszweige

  • 303 Gesundheitswissenschaften
  • 304 Medizinische Biotechnologie
  • 304003 Gentechnik
  • 305 Andere Humanmedizin, Gesundheitswissenschaften
  • 101004 Biomathematik
  • 101018 Statistik
  • 102 Informatik
  • 102001 Artificial Intelligence
  • 102004 Bioinformatik
  • 102010 Datenbanksysteme
  • 102015 Informationssysteme
  • 102019 Machine Learning
  • 106023 Molekularbiologie
  • 106002 Biochemie
  • 106005 Bioinformatik
  • 106007 Biostatistik
  • 106041 Strukturbiologie
  • 301 Medizinisch-theoretische Wissenschaften, Pharmazie
  • 302 Klinische Medizin

JKU-Schwerpunkte

  • Computation in Informatics and Mathematics
  • Nano-, Bio- and Polymer-Systems: From Structure to Function
  • MED Allgemein
  • Versorgungsforschung
  • Klinische Altersforschung

Dieses zitieren