An Interpretable Transformer-Based Autoencoder Framework for Identifying Driver Mutation in Colorectal Cancer
Öz
Identification of driving genetic mutations in colorectal cancer is one of the main challenges given the multi-dimensional nature and the difficulty of genomic mutation data. In this study, a deep learning model that is based on the transformer-based Autoencoder was proposed to identify influential genetic mutations linked with colorectal cancer, using data from the cosmic database.
The proposed model is dependent on a self-attention mechanism to deal with complex interfaces between genes and long-range associations between mutations. To assess the its efficacy, the model was compared with a classical Autoencoder that is dependent on dense layers as a reference framework. Additionally, a new indicator identified as the Mutation Importance Score (MIS), that integrates attention weights and reconstruction error, with an attempt to rank mutations based on their possible potential biological implication.
Experimentation results showed that the Transformer-based model attained a significant reduction in reconstruction error in comparison with the reference model, as the test loss value reduced from 0.154 to 0.107, to reflect an enhancement around 30%. The proposed model indicated a clear enhancement in the ability to find biologically significant genes, with an advanced overlap with the Cancer Gene Census database of (87% vs. 63%), as well as an enhancement in the Precision@20 scale of (0.85 vs. 0.65). These results indicate the model's ability to give exact priority to possible motor mutations.
These results indicate that deep learning models based on the self- attention mechanism are strong tools to analyze large-scale genomic mutation data, and may contribute to the growth of computational methods in the field of cancer genomics and accuracy medicine.
Anahtar Kelimeler
Destekleyen Kurum
The authors received no specific funding or institutional support for this research.
Etik Beyan
The authors confirm that this research was conducted in accordance with ethical standards. The manuscript is original, has not been published previously, and is not under consideration elsewhere. All authors have approved the final version of the manuscript and declare that there are no conflicts of interest related to this work.
Teşekkür
The authors would like to express their sincere gratitude to Al-Nahrain University for providing support and facilities that contributed to the completion of this research.
An Interpretable Transformer-Based Autoencoder Framework for Identifying Driver Mutation in Colorectal Cancer
Öz
Identification of driving genetic mutations in colorectal cancer is one of the main challenges given the multi-dimensional nature and the difficulty of genomic mutation data. In this study, a deep learning model that is based on the transformer-based Autoencoder was proposed to identify influential genetic mutations linked with colorectal cancer, using data from the cosmic database.
The proposed model is dependent on a self-attention mechanism to deal with complex interfaces between genes and long-range associations between mutations. To assess the its efficacy, the model was compared with a classical Autoencoder that is dependent on dense layers as a reference framework. Additionally, a new indicator identified as the Mutation Importance Score (MIS), that integrates attention weights and reconstruction error, with an attempt to rank mutations based on their possible potential biological implication.
Experimentation results showed that the Transformer-based model attained a significant reduction in reconstruction error in comparison with the reference model, as the test loss value reduced from 0.154 to 0.107, to reflect an enhancement around 30%. The proposed model indicated a clear enhancement in the ability to find biologically significant genes, with an advanced overlap with the Cancer Gene Census database of (87% vs. 63%), as well as an enhancement in the Precision@20 scale of (0.85 vs. 0.65). These results indicate the model's ability to give exact priority to possible motor mutations.
These results indicate that deep learning models based on the self- attention mechanism are strong tools to analyze large-scale genomic mutation data, and may contribute to the growth of computational methods in the field of cancer genomics and accuracy medicine.
Anahtar Kelimeler
Destekleyen Kurum
Yazarlar bu araştırma için herhangi bir özel fon veya kurumsal destek almamıştır.
Etik Beyan
Yazarlar, bu araştırmanın etik standartlara uygun olarak yürütüldüğünü teyit ederler. Makale özgündür, daha önce yayınlanmamıştır ve başka bir yerde değerlendirme aşamasında değildir. Tüm yazarlar makalenin son halini onaylamış ve bu çalışmayla ilgili herhangi bir çıkar çatışması olmadığını beyan etmişlerdir.
Teşekkür
Yazarlar, bu araştırmanın tamamlanmasına katkıda bulunan destek ve olanakları sağladığı için El-Nahrain Üniversitesi'ne içten teşekkürlerini sunmak isterler.