How Cheminformatics Has Been Changed by Machine Learning? – Recent Progress
How Cheminformatics Has Been Changed by Machine Learning? – Recent Progress
Cheminformatics refers to computational management and analysis of chemical data, including visualization of molecular structures and analyzed datasets. It covers a broad range of disciplines and is developing with the aid of advancements in computation and information technology. Now is the age of machine leaning (ML) and it’s enhancing the capability of cheminformatics and expanding its perspective.
Cheminformatics is a crackpot concept for the application of accumulated data in not only pharmaceutical sciences but also in any field related to chemistry. In drug discovery and development, cheminformatics has been utilized for molecular design against the protein of interest through quantitative structure-activity relationship (QSAR). Since the proposal of the concept of cheminformatics in 1998,1) it became a formidable area of innovation.
Recent cheminformatics have recruited ML, with keeping the well-established knowledge on statistics, physics, chemistry, biology and mathematics.2) There is a possibility of the advent of paradigm shift in computer-aided drug design (CADD).3) So-called “Digital Transfomartion (DX)” in drug discovery field could be brought about by ML-driven cheminformatics.
The review2) summarized the current state and the future perspective of ML-driven cheminformatics from the very basis, from how you deal with the chemical dataset like the choice of descriptor. Here we would like to focus on the application side in drug discovery.5)
The authors focus on the possibility of expanding the possibility of cheminformatics by the appropriate merge with ML-based QSAR. Since big pharmaceutical companies has long been accumulated their own data, large-scale analysis of the treasure enhanced by ML would provide their own outcomes and gain their unique strength in CADD.
Basically, QSAR aims for the prediction of biological outcome by physicochemical characteristics of chemical entities.4) DL are playing its significant role in QSAR modelling and quantitative structure-property relationship (QSPR) is also possible now.6) Cheminformatics builds up an infrastructure for the sophisticated ML-based approach for QSAR and QSPR.
Machine-learning-based QSAR comprises of four steps: modelling, molecular encoding, feature selection and model training. Modeling and training are featuring and differentiating factors of QSAR performance and quality, and so many algorisms have been developed so far. Supervised learning and unsupervised learning are the main categories of ML models and both approaches have advantages on QSAR.
Supervised learning trains the model with labeled data on known input-output correlation.
In supervised learning, the model obeys the rules mostly deduced by statistical analysis. But the speed of multiple analyses against the data is higher and precise to predict biological activities, ADMET profiles and so on.
Unsupervised learning uses unlabeled data to discover underlying relationships without known correlation or guidance. The model tries to uncover the hidden pattens of data through the basic algorism. Hence unsupervised learning-based model has a potential for serendipitous output, even though precisely designed model and dataset with uniformity is necessary for phenomenal discovery.
The crucial challenge on ML-based QSAR is to overcome the interpretability and explainability. The model often provides you an output that is not understandable by a researcher’s mind. ML is a kind of black box and some techniques like LIME and SHAP have been developed.7),8) Solving this issue allows us to learn from ML-based QSAR and mutual collaboration comes true between machines and humans.
As you realized it, ML-based QSAR is changing or adding the role of cheminformatics. It is no need to mention that management and quality control of the precious and meaningful data is a must for ML. Data selection for training changes the result of ML-based QSAR. Researchers in cheminformatics who are willing to expand their strength according to the age of computational science would get a ticket for jumping on a different and exciting stage.
We are eagerly working on expanding the capability of cheminformatics, or we would say, informatics and computational sciences. Please join us if you want to make an innovation in drug discovery field.
1)https://doi.org/10.1016/s0065-7743(08)61100-8
2)https://doi.org/10.3390/ijms241411488
3)https://doi.org/10.3390/ph17010022
4)https://doi.org/10.1016/j.drudis.2018.05.010
5)https://doi.org/10.1080/17460441.2021.1909567
6)https://doi.org/10.1016/j.ymeth.2014.09.009
7)https://c3.ai/glossary/data-science/lime-local-interpretable-model-agnostic-explanations/
8)https://christophm.github.io/interpretable-ml-book/shap.html


