A Machine Learning and Explainable AI Framework Tailored for Unbalanced Experimental Catalyst Discovery

The successful application of machine learning (ML) in catalyst design relies on high-quality and diverse data to ensure effective generalization to novel compositions, thereby aiding in catalyst discovery. However, due to complex interactions, catalyst design has long relied on trial-and-error, a c...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Hauptverfasser:	Semnani, Parastoo, Bogojeski, Mihail, Bley, Florian, Zhang, Zizheng, Wu, Qiong, Kneib, Thomas, Herrmann, Jan, Weisser, Christoph, Patcas, Florina, Müller, Klaus-Robert
Format:	Artikel
Sprache:	eng
Schlagworte:	Computer Science - Learning Physics - Chemical Physics
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

container_end_page
container_issue
container_start_page
container_title
container_volume
creator	Semnani, Parastoo Bogojeski, Mihail Bley, Florian Zhang, Zizheng Wu, Qiong Kneib, Thomas Herrmann, Jan Weisser, Christoph Patcas, Florina Müller, Klaus-Robert
description	The successful application of machine learning (ML) in catalyst design relies on high-quality and diverse data to ensure effective generalization to novel compositions, thereby aiding in catalyst discovery. However, due to complex interactions, catalyst design has long relied on trial-and-error, a costly and labor-intensive process leading to scarce data that is heavily biased towards undesired, low-yield catalysts. Despite the rise of ML in this field, most efforts have not focused on dealing with the challenges presented by such experimental data. To address these challenges, we introduce a robust machine learning and explainable AI (XAI) framework to accurately classify the catalytic yield of various compositions and identify the contributions of individual components. This framework combines a series of ML practices designed to handle the scarcity and imbalance of catalyst data. We apply the framework to classify the yield of various catalyst compositions in oxidative methane coupling, and use it to evaluate the performance of a range of ML models: tree-based models, logistic regression, support vector machines, and neural networks. These experiments demonstrate that the methods used in our framework lead to a significant improvement in the performance of all but one of the evaluated models. Additionally, the decision-making process of each ML model is analyzed by identifying the most important features for predicting catalyst performance using XAI methods. Our analysis found that XAI methods, providing class-aware explanations, such as Layer-wise Relevance Propagation, identified key components that contribute specifically to high-yield catalysts. These findings align with chemical intuition and existing literature, reinforcing their validity. We believe that such insights can assist chemists in the development and identification of novel catalysts with superior performance.
doi_str_mv	10.48550/arxiv.2407.18935
format	Article
fullrecord	<record><control><sourceid>arxiv_GOX</sourceid><recordid>TN_cdi_arxiv_primary_2407_18935</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2407_18935</sourcerecordid><originalsourceid>FETCH-arxiv_primary_2407_189353</originalsourceid><addsrcrecordid>eNqFjjEOgkAQRbexMOoBrJwLiCAQsSQI0UQ7rckAA25cFjIQhNuLxN7mv-a_5AmxtkzD8VzX3CH3sjP2jnkwLO9ou3NR-HDD9Ck1wZWQtdQFoM4g7GuFUmOiCPwLRIwlvSt-wR2lqpgyyCuGh05QoU5pEohlSbpFBQGOOzQtnGSTVh3xsBSzHFVDqx8XYhOF9-C8nZLielSRh_ibFk9p9v_HBxY3RNk</addsrcrecordid><sourcetype>Open Access Repository</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype></control><display><type>article</type><title>A Machine Learning and Explainable AI Framework Tailored for Unbalanced Experimental Catalyst Discovery</title><source>arXiv.org</source><creator>Semnani, Parastoo ; Bogojeski, Mihail ; Bley, Florian ; Zhang, Zizheng ; Wu, Qiong ; Kneib, Thomas ; Herrmann, Jan ; Weisser, Christoph ; Patcas, Florina ; Müller, Klaus-Robert</creator><creatorcontrib>Semnani, Parastoo ; Bogojeski, Mihail ; Bley, Florian ; Zhang, Zizheng ; Wu, Qiong ; Kneib, Thomas ; Herrmann, Jan ; Weisser, Christoph ; Patcas, Florina ; Müller, Klaus-Robert</creatorcontrib><description>The successful application of machine learning (ML) in catalyst design relies on high-quality and diverse data to ensure effective generalization to novel compositions, thereby aiding in catalyst discovery. However, due to complex interactions, catalyst design has long relied on trial-and-error, a costly and labor-intensive process leading to scarce data that is heavily biased towards undesired, low-yield catalysts. Despite the rise of ML in this field, most efforts have not focused on dealing with the challenges presented by such experimental data. To address these challenges, we introduce a robust machine learning and explainable AI (XAI) framework to accurately classify the catalytic yield of various compositions and identify the contributions of individual components. This framework combines a series of ML practices designed to handle the scarcity and imbalance of catalyst data. We apply the framework to classify the yield of various catalyst compositions in oxidative methane coupling, and use it to evaluate the performance of a range of ML models: tree-based models, logistic regression, support vector machines, and neural networks. These experiments demonstrate that the methods used in our framework lead to a significant improvement in the performance of all but one of the evaluated models. Additionally, the decision-making process of each ML model is analyzed by identifying the most important features for predicting catalyst performance using XAI methods. Our analysis found that XAI methods, providing class-aware explanations, such as Layer-wise Relevance Propagation, identified key components that contribute specifically to high-yield catalysts. These findings align with chemical intuition and existing literature, reinforcing their validity. We believe that such insights can assist chemists in the development and identification of novel catalysts with superior performance.</description><identifier>DOI: 10.48550/arxiv.2407.18935</identifier><language>eng</language><subject>Computer Science - Learning ; Physics - Chemical Physics</subject><creationdate>2024-07</creationdate><rights>http://creativecommons.org/licenses/by-nc-nd/4.0</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>228,230,780,885</link.rule.ids><linktorsrc>$$Uhttps://arxiv.org/abs/2407.18935$$EView_record_in_Cornell_University$$FView_record_in_$$GCornell_University$$Hfree_for_read</linktorsrc><backlink>$$Uhttps://doi.org/10.48550/arXiv.2407.18935$$DView paper in arXiv$$Hfree_for_read</backlink></links><search><creatorcontrib>Semnani, Parastoo</creatorcontrib><creatorcontrib>Bogojeski, Mihail</creatorcontrib><creatorcontrib>Bley, Florian</creatorcontrib><creatorcontrib>Zhang, Zizheng</creatorcontrib><creatorcontrib>Wu, Qiong</creatorcontrib><creatorcontrib>Kneib, Thomas</creatorcontrib><creatorcontrib>Herrmann, Jan</creatorcontrib><creatorcontrib>Weisser, Christoph</creatorcontrib><creatorcontrib>Patcas, Florina</creatorcontrib><creatorcontrib>Müller, Klaus-Robert</creatorcontrib><title>A Machine Learning and Explainable AI Framework Tailored for Unbalanced Experimental Catalyst Discovery</title><description>The successful application of machine learning (ML) in catalyst design relies on high-quality and diverse data to ensure effective generalization to novel compositions, thereby aiding in catalyst discovery. However, due to complex interactions, catalyst design has long relied on trial-and-error, a costly and labor-intensive process leading to scarce data that is heavily biased towards undesired, low-yield catalysts. Despite the rise of ML in this field, most efforts have not focused on dealing with the challenges presented by such experimental data. To address these challenges, we introduce a robust machine learning and explainable AI (XAI) framework to accurately classify the catalytic yield of various compositions and identify the contributions of individual components. This framework combines a series of ML practices designed to handle the scarcity and imbalance of catalyst data. We apply the framework to classify the yield of various catalyst compositions in oxidative methane coupling, and use it to evaluate the performance of a range of ML models: tree-based models, logistic regression, support vector machines, and neural networks. These experiments demonstrate that the methods used in our framework lead to a significant improvement in the performance of all but one of the evaluated models. Additionally, the decision-making process of each ML model is analyzed by identifying the most important features for predicting catalyst performance using XAI methods. Our analysis found that XAI methods, providing class-aware explanations, such as Layer-wise Relevance Propagation, identified key components that contribute specifically to high-yield catalysts. These findings align with chemical intuition and existing literature, reinforcing their validity. We believe that such insights can assist chemists in the development and identification of novel catalysts with superior performance.</description><subject>Computer Science - Learning</subject><subject>Physics - Chemical Physics</subject><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2024</creationdate><recordtype>article</recordtype><sourceid>GOX</sourceid><recordid>eNqFjjEOgkAQRbexMOoBrJwLiCAQsSQI0UQ7rckAA25cFjIQhNuLxN7mv-a_5AmxtkzD8VzX3CH3sjP2jnkwLO9ou3NR-HDD9Ck1wZWQtdQFoM4g7GuFUmOiCPwLRIwlvSt-wR2lqpgyyCuGh05QoU5pEohlSbpFBQGOOzQtnGSTVh3xsBSzHFVDqx8XYhOF9-C8nZLielSRh_ibFk9p9v_HBxY3RNk</recordid><startdate>20240710</startdate><enddate>20240710</enddate><creator>Semnani, Parastoo</creator><creator>Bogojeski, Mihail</creator><creator>Bley, Florian</creator><creator>Zhang, Zizheng</creator><creator>Wu, Qiong</creator><creator>Kneib, Thomas</creator><creator>Herrmann, Jan</creator><creator>Weisser, Christoph</creator><creator>Patcas, Florina</creator><creator>Müller, Klaus-Robert</creator><scope>AKY</scope><scope>GOX</scope></search><sort><creationdate>20240710</creationdate><title>A Machine Learning and Explainable AI Framework Tailored for Unbalanced Experimental Catalyst Discovery</title><author>Semnani, Parastoo ; Bogojeski, Mihail ; Bley, Florian ; Zhang, Zizheng ; Wu, Qiong ; Kneib, Thomas ; Herrmann, Jan ; Weisser, Christoph ; Patcas, Florina ; Müller, Klaus-Robert</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-arxiv_primary_2407_189353</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2024</creationdate><topic>Computer Science - Learning</topic><topic>Physics - Chemical Physics</topic><toplevel>online_resources</toplevel><creatorcontrib>Semnani, Parastoo</creatorcontrib><creatorcontrib>Bogojeski, Mihail</creatorcontrib><creatorcontrib>Bley, Florian</creatorcontrib><creatorcontrib>Zhang, Zizheng</creatorcontrib><creatorcontrib>Wu, Qiong</creatorcontrib><creatorcontrib>Kneib, Thomas</creatorcontrib><creatorcontrib>Herrmann, Jan</creatorcontrib><creatorcontrib>Weisser, Christoph</creatorcontrib><creatorcontrib>Patcas, Florina</creatorcontrib><creatorcontrib>Müller, Klaus-Robert</creatorcontrib><collection>arXiv Computer Science</collection><collection>arXiv.org</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext_linktorsrc</fulltext></delivery><addata><au>Semnani, Parastoo</au><au>Bogojeski, Mihail</au><au>Bley, Florian</au><au>Zhang, Zizheng</au><au>Wu, Qiong</au><au>Kneib, Thomas</au><au>Herrmann, Jan</au><au>Weisser, Christoph</au><au>Patcas, Florina</au><au>Müller, Klaus-Robert</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>A Machine Learning and Explainable AI Framework Tailored for Unbalanced Experimental Catalyst Discovery</atitle><date>2024-07-10</date><risdate>2024</risdate><abstract>The successful application of machine learning (ML) in catalyst design relies on high-quality and diverse data to ensure effective generalization to novel compositions, thereby aiding in catalyst discovery. However, due to complex interactions, catalyst design has long relied on trial-and-error, a costly and labor-intensive process leading to scarce data that is heavily biased towards undesired, low-yield catalysts. Despite the rise of ML in this field, most efforts have not focused on dealing with the challenges presented by such experimental data. To address these challenges, we introduce a robust machine learning and explainable AI (XAI) framework to accurately classify the catalytic yield of various compositions and identify the contributions of individual components. This framework combines a series of ML practices designed to handle the scarcity and imbalance of catalyst data. We apply the framework to classify the yield of various catalyst compositions in oxidative methane coupling, and use it to evaluate the performance of a range of ML models: tree-based models, logistic regression, support vector machines, and neural networks. These experiments demonstrate that the methods used in our framework lead to a significant improvement in the performance of all but one of the evaluated models. Additionally, the decision-making process of each ML model is analyzed by identifying the most important features for predicting catalyst performance using XAI methods. Our analysis found that XAI methods, providing class-aware explanations, such as Layer-wise Relevance Propagation, identified key components that contribute specifically to high-yield catalysts. These findings align with chemical intuition and existing literature, reinforcing their validity. We believe that such insights can assist chemists in the development and identification of novel catalysts with superior performance.</abstract><doi>10.48550/arxiv.2407.18935</doi><oa>free_for_read</oa></addata></record>
fulltext	fulltext_linktorsrc
identifier	DOI: 10.48550/arxiv.2407.18935
ispartof
issn
language	eng
recordid	cdi_arxiv_primary_2407_18935
source	arXiv.org
subjects	Computer Science - Learning Physics - Chemical Physics
title	A Machine Learning and Explainable AI Framework Tailored for Unbalanced Experimental Catalyst Discovery
url	https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-12T19%3A39%3A02IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-arxiv_GOX&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=A%20Machine%20Learning%20and%20Explainable%20AI%20Framework%20Tailored%20for%20Unbalanced%20Experimental%20Catalyst%20Discovery&rft.au=Semnani,%20Parastoo&rft.date=2024-07-10&rft_id=info:doi/10.48550/arxiv.2407.18935&rft_dat=%3Carxiv_GOX%3E2407_18935%3C/arxiv_GOX%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_id=info:pmid/&rfr_iscdi=true