Extracting gene expression patterns and identifying co-expressed genes from microarray data reveals biologically responsive processes

被引:42
|
作者
Chou, Jeff W. [1 ]
Zhou, Tong [2 ,3 ]
Kaufmann, William K. [2 ,3 ]
Paules, Richard S. [1 ]
Bushel, Pierre R. [1 ]
机构
[1] Natl Inst Environm Hlth Sci, Microarray Grp, Res Triangle Pk, NC 27709 USA
[2] Univ N Carolina, Ctr Environm Hlth & Susceptibil, Dept Pathol & Lab Med, Chapel Hill, NC 27599 USA
[3] Univ N Carolina, Lineberger Comprehens Canc Ctr, Chapel Hill, NC 27599 USA
关键词
D O I
10.1186/1471-2105-8-427
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background: A common observation in the analysis of gene expression data is that many genes display similarity in their expression patterns and therefore appear to be co-regulated. However, the variation associated with microarray data and the complexity of the experimental designs make the acquisition of co-expressed genes a challenge. We developed a novel method for Extracting microarray gene expression Patterns and Identifying co-expressed Genes, designated as EPIG. The approach utilizes the underlying structure of gene expression data to extract patterns and identify co-expressed genes that are responsive to experimental conditions. Results: Through evaluation of the correlations among profiles, the magnitude of variation in gene expression profiles, and profile signal-to-noise ratio's, EPIG extracts a set of patterns representing co-expressed genes. The method is shown to work well with a simulated data set and microarray data obtained from time-series studies of dauer recovery and LI starvation in C. elegans and after ultraviolet (UV) or ionizing radiation (IR)-induced DNA damage in diploid human fibroblasts. With the simulated data set, EPIG extracted the appropriate number of patterns which were more stable and homogeneous than the set of patterns that were determined using the CLICK or CAST clustering algorithms. However, CLICK performed better than EPIG and CAST with respect to the average correlation between clusters/patterns of the simulated data. With real biological data, EPIG extracted more dauer-specific patterns than CLICK. Furthermore, analysis of the IR/UV data revealed 18 unique patterns and 2661 genes out of approximately 17,000 that were identified as significantly expressed and categorized to the patterns by EPIG. The time-dependent patterns displayed similar and dissimilar responses between IR and UV treatments. Gene Ontology analysis applied to each pattern-related subset of co-expressed genes revealed underlying biological processes affected by IR-and/or UV-induced DNA damage. Conclusion: EPIG competed with CLICK and performed better than CAST in extracting patterns from simulated data. EPIG extracted more biological informative patterns and co-expressed genes from both C. elegans and IR/UV-treated human fibroblasts. Using Gene Ontology analysis of the genes in the patterns extracted by EPIG, several key biological categories related to p53-dependent cell cycle control were revealed from the IR/UV data. Among them were mitotic cell cycle, DNA replication, DNA repair, cell cycle checkpoint, and G(0)-like status transition. EPIG can be applied to data sets from a variety of experimental designs.
引用
收藏
页数:16
相关论文
共 50 条
  • [1] Extracting gene expression patterns and identifying co-expressed genes from microarray data reveals biologically responsive processes
    Jeff W Chou
    Tong Zhou
    William K Kaufmann
    Richard S Paules
    Pierre R Bushel
    [J]. BMC Bioinformatics, 8
  • [2] Covariance thresholding to detect differentially co-expressed genes from microarray gene expression data
    Oh, Mingyu
    Kim, Kipoong
    Sun, Hokeun
    [J]. JOURNAL OF BIOINFORMATICS AND COMPUTATIONAL BIOLOGY, 2020, 18 (01)
  • [3] EPIG-Seq: extracting patterns and identifying co-expressed genes from RNA-Seq data
    Li, Jianying
    Bushel, Pierre R.
    [J]. BMC GENOMICS, 2016, 17
  • [4] EPIG-Seq: extracting patterns and identifying co-expressed genes from RNA-Seq data
    Jianying Li
    Pierre R. Bushel
    [J]. BMC Genomics, 17
  • [5] Genomic positions of co-expressed genes: Echoes of chromosome organisation in gene expression data
    Szczepińska T.
    Pawłowski K.
    [J]. BMC Research Notes, 6 (1)
  • [6] Analysis of differentially co-expressed genes based on microarray data of hepatocellular carcinoma
    Wang, Y.
    Jiang, T.
    Li, Z.
    Lu, L.
    Zhang, R.
    Zhang, D.
    Wang, X.
    Tan, J.
    [J]. NEOPLASMA, 2017, 64 (02) : 216 - +
  • [7] Clust: automatic extraction of optimal co-expressed gene clusters from gene expression data
    Basel Abu-Jamous
    Steven Kelly
    [J]. Genome Biology, 19
  • [8] Clust: automatic extraction of optimal co-expressed gene clusters from gene expression data
    Abu-Jamous, Basel
    Kelly, Steven
    [J]. GENOME BIOLOGY, 2018, 19
  • [9] Extracting biologically significant patterns from short time series gene expression data
    Alain B Tchagang
    Kevin V Bui
    Thomas McGinnis
    Panayiotis V Benos
    [J]. BMC Bioinformatics, 10
  • [10] Novel implementation of conditional co-regulation by graph theory to derive co-expressed genes from microarray data
    Arun Rawat
    Georg J Seifert
    Youping Deng
    [J]. BMC Bioinformatics, 9