Algorithms for matching partially labelled sequence graphs

被引：0

作者：

Taylor, William R. ^{[1
]}

机构：

[1] Francis Crick Inst, 1 Midland Rd, London NW1 1AT, England

来源：

ALGORITHMS FOR MOLECULAR BIOLOGY | 2017年 / 12卷

基金：

英国惠康基金;

关键词：

Phylogenetic tree matching; Correlated substitution analysis; Bipartite graph matching; PROTEIN; CONTACTS; COEVOLUTION; PREDICTION;

D O I：

10.1186/s13015-017-0115-y

中图分类号：

Q5 [生物化学];

学科分类号：

071010 ; 081704 ;

摘要：

Background: In order to find correlated pairs of positions between proteins, which are useful in predicting interactions, it is necessary to concatenate two large multiple sequence alignments such that the sequences that are joined together belong to those that interact in their species of origin. When each protein is unique then the species name is sufficient to guide this match, however, when there are multiple related sequences (paralogs) in each species then the pairing is more difficult. In bacteria a good guide can be gained from genome co-location as interacting proteins tend to be in a common operon but in eukaryotes this simple principle is not sufficient. Results: The methods developed in this paper take sets of paralogs for different proteins found in the same species and make a pairing based on their evolutionary distance relative to a set of other proteins that are unique and so have a known relationship (singletons). The former constitute a set of unlabelled nodes in a graph while the latter are labelled. Two variants were tested, one based on a phylogenetic tree of the sequences (the topology-based method) and a simpler, faster variant based only on the inter-sequence distances (the distance-based method). Over a set of test proteins, both gave good results, with the topology method performing slightly better. Conclusions: The methods develop here still need refinement and augmentation from constraints other than the sequence data alone, such as known interactions from annotation and databases, or non-trivial relationships in genome location. With the ever growing numbers of eukaryotic genomes, it is hoped that the methods described here will open a route to the use of these data equal to the current success attained with bacterial sequences.

引用

页数：22

共 50 条

[1] Algorithms for matching partially labelled sequence graphs
William R. Taylor
Algorithms for Molecular Biology, 12
[2] Adaptive Matching Based Kernels for Labelled Graphs
Woznica, Adam
Kalousis, Alexandros
Hilario, Melanie
ADVANCES IN KNOWLEDGE DISCOVERY AND DATA MINING, PT II, PROCEEDINGS, 2010, 6119 : 374 - 385
[3] Multivariate matching polynomials of cyclically labelled graphs
McSorley, John P.
Feinsilver, Philip
DISCRETE MATHEMATICS, 2009, 309 (10) : 3205 - 3218
[4] Matching algorithms are fast in sparse random graphs
Bast, H
Mehlhorn, K
Schäfer, G
Tamaki, H
THEORY OF COMPUTING SYSTEMS, 2006, 39 (01) : 3 - 14
[5] EFFICIENT ALGORITHMS FOR FINDING MAXIMUM MATCHING IN GRAPHS
GALIL, Z
COMPUTING SURVEYS, 1986, 18 (01) : 23 - 38
[6] Scaling Algorithms for Weighted Matching in General Graphs
Duan, Ran
Pettie, Seth
Su, Hsin-Hao
ACM TRANSACTIONS ON ALGORITHMS, 2018, 14 (01)
[7] Multithreaded Algorithms for Maximum Matching in Bipartite Graphs
Azad, Ariful
Halappanavar, Mahantesh
Rajamanickam, Sivasankaran
Boman, Erik G.
Khan, Arif
Pothen, Alex
2012 IEEE 26TH INTERNATIONAL PARALLEL AND DISTRIBUTED PROCESSING SYMPOSIUM (IPDPS), 2012, : 860 - 872
[8] Matching algorithms are fast in sparse random graphs
Bast, H
Mehlhorn, K
Schäfer, G
Tamaki, H
STACS 2004, PROCEEDINGS, 2004, 2996 : 81 - 92
[9] Matching Algorithms Are Fast in Sparse Random Graphs
Holger Bast
Kurt Mehlhorn
Guido Schafer
Hisao Tamaki
Theory of Computing Systems, 2006, 39 : 3 - 14
[10] EFFICIENT ALGORITHMS FOR FINDING MAXIMAL MATCHING IN GRAPHS
GALIL, Z
LECTURE NOTES IN COMPUTER SCIENCE, 1983, 159 : 90 - 113

← 1 2 3 4 5 →