Identifying Individual-Cancer-Related Genes by Rebalancing the Training Samples

Bolin Chen; Xuequn Shang; Min Li; Jianxin Wang; Fang Xiang Wu

doi:10.1109/TNB.2016.2553119

Identifying Individual-Cancer-Related Genes by Rebalancing the Training Samples

Bolin Chen, Xuequn Shang, Min Li, Jianxin Wang, Fang Xiang Wu

计算机学院

科研成果: 期刊稿件 › 文章 › 同行评审

18 引用（Scopus）

摘要

The identification of individual-cancer-related genes typically is an imbalanced classification issue. The number of known cancer-related genes is far less than the number of all unknown genes, which makes it very hard to detect novel predictions from such imbalanced training samples. A regular machine learning method can either only detect genes related to all cancers or add clinical knowledge to circumvent this issue. In this study, we introduce a training sample rebalancing strategy to overcome this issue by using a two-step logistic regression and a random resampling method. The two-step logistic regression is to select a set of genes that related to all cancers. While the random resampling method is performed to further classify those genes associated with individual cancers. The issue of imbalanced classification is circumvented by randomly adding positive instances related to other cancers at first, and then excluding those unrelated predictions according to the overall performance at the following step. Numerical experiments show that the proposed resampling method is able to identify cancer-related genes even when the number of known genes related to it is small. The final predictions for all individual cancers achieve AUC values around 0.93 by using the leave-one-out cross validation method, which is very promising, compared with existing methods.

源语言	英语
文章编号	7451278
页（从-至）	309-315
页数	7
期刊	IEEE Transactions on Nanobioscience
卷	15
期	4
DOI	https://doi.org/10.1109/TNB.2016.2553119
出版状态	已出版 - 6月 2016

联合国可持续发展目标

此成果有助于实现下列可持续发展目标：

访问文件

10.1109/TNB.2016.2553119

其它文件与链接

链接到 Scopus 的出版物

引用此

@article{7d8833ca06f9411d9a7f8c80364c158a,

title = "Identifying Individual-Cancer-Related Genes by Rebalancing the Training Samples",

abstract = "The identification of individual-cancer-related genes typically is an imbalanced classification issue. The number of known cancer-related genes is far less than the number of all unknown genes, which makes it very hard to detect novel predictions from such imbalanced training samples. A regular machine learning method can either only detect genes related to all cancers or add clinical knowledge to circumvent this issue. In this study, we introduce a training sample rebalancing strategy to overcome this issue by using a two-step logistic regression and a random resampling method. The two-step logistic regression is to select a set of genes that related to all cancers. While the random resampling method is performed to further classify those genes associated with individual cancers. The issue of imbalanced classification is circumvented by randomly adding positive instances related to other cancers at first, and then excluding those unrelated predictions according to the overall performance at the following step. Numerical experiments show that the proposed resampling method is able to identify cancer-related genes even when the number of known genes related to it is small. The final predictions for all individual cancers achieve AUC values around 0.93 by using the leave-one-out cross validation method, which is very promising, compared with existing methods.",

keywords = "Cancer-related gene, imbalanced classification, logistic regression, resampling method",

author = "Bolin Chen and Xuequn Shang and Min Li and Jianxin Wang and Wu, {Fang Xiang}",

note = "Publisher Copyright: {\textcopyright} 2016 IEEE.",

year = "2016",

month = jun,

doi = "10.1109/TNB.2016.2553119",

language = "英语",

volume = "15",

pages = "309--315",

journal = "IEEE Transactions on Nanobioscience",

issn = "1536-1241",

publisher = "Institute of Electrical and Electronics Engineers Inc.",

number = "4",

}

TY - JOUR

T1 - Identifying Individual-Cancer-Related Genes by Rebalancing the Training Samples

AU - Chen, Bolin

AU - Shang, Xuequn

AU - Li, Min

AU - Wang, Jianxin

AU - Wu, Fang Xiang

PY - 2016/6

Y1 - 2016/6

N2 - The identification of individual-cancer-related genes typically is an imbalanced classification issue. The number of known cancer-related genes is far less than the number of all unknown genes, which makes it very hard to detect novel predictions from such imbalanced training samples. A regular machine learning method can either only detect genes related to all cancers or add clinical knowledge to circumvent this issue. In this study, we introduce a training sample rebalancing strategy to overcome this issue by using a two-step logistic regression and a random resampling method. The two-step logistic regression is to select a set of genes that related to all cancers. While the random resampling method is performed to further classify those genes associated with individual cancers. The issue of imbalanced classification is circumvented by randomly adding positive instances related to other cancers at first, and then excluding those unrelated predictions according to the overall performance at the following step. Numerical experiments show that the proposed resampling method is able to identify cancer-related genes even when the number of known genes related to it is small. The final predictions for all individual cancers achieve AUC values around 0.93 by using the leave-one-out cross validation method, which is very promising, compared with existing methods.

AB - The identification of individual-cancer-related genes typically is an imbalanced classification issue. The number of known cancer-related genes is far less than the number of all unknown genes, which makes it very hard to detect novel predictions from such imbalanced training samples. A regular machine learning method can either only detect genes related to all cancers or add clinical knowledge to circumvent this issue. In this study, we introduce a training sample rebalancing strategy to overcome this issue by using a two-step logistic regression and a random resampling method. The two-step logistic regression is to select a set of genes that related to all cancers. While the random resampling method is performed to further classify those genes associated with individual cancers. The issue of imbalanced classification is circumvented by randomly adding positive instances related to other cancers at first, and then excluding those unrelated predictions according to the overall performance at the following step. Numerical experiments show that the proposed resampling method is able to identify cancer-related genes even when the number of known genes related to it is small. The final predictions for all individual cancers achieve AUC values around 0.93 by using the leave-one-out cross validation method, which is very promising, compared with existing methods.

KW - Cancer-related gene

KW - imbalanced classification

KW - logistic regression

KW - resampling method

UR - http://www.scopus.com/inward/record.url?scp=84982298202&partnerID=8YFLogxK

U2 - 10.1109/TNB.2016.2553119

DO - 10.1109/TNB.2016.2553119

M3 - 文章

C2 - 27093705

AN - SCOPUS:84982298202

SN - 1536-1241

VL - 15

SP - 309

EP - 315

JO - IEEE Transactions on Nanobioscience

JF - IEEE Transactions on Nanobioscience

IS - 4

M1 - 7451278

ER -

Identifying Individual-Cancer-Related Genes by Rebalancing the Training Samples

摘要

联合国可持续发展目标

访问文件

其它文件与链接

指纹

引用此