A polynomial time algorithm for calculating the probability of a ranked gene tree given a species tree
© Stadler and Degnan; licensee BioMed Central Ltd. 2012
Received: 30 September 2011
Accepted: 2 April 2012
Published: 30 April 2012
The ancestries of genes form gene trees which do not necessarily have the same topology as the species tree due to incomplete lineage sorting. Available algorithms determining the probability of a gene tree given a species tree require exponential computational runtime.
In this paper, we provide a polynomial time algorithm to calculate the probability of a ranked gene tree topology for a given species tree, where a ranked tree topology is a tree topology with the internal vertices being ordered. The probability of a gene tree topology can thus be calculated in polynomial time if the number of orderings of the internal vertices is a polynomial number. However, the complexity of calculating the probability of a gene tree topology with an exponential number of rankings for a given species tree remains unknown.
Polynomial algorithms for calculating ranked gene tree probabilities may become useful in developing methodology to infer species trees based on a collection of gene trees, leading to a more accurate reconstruction of ancestral species relationships.
KeywordsIncomplete lineage sorting Coalescent history Anomalous gene tree Dynamic programming
Given a fixed species tree, and assuming the gene tree evolved under the multi-species coalescent , the most probable gene tree topology can have a different topology from that of the species tree. Such a gene tree topology is called an anomalous gene tree. In fact, for every species tree topology with at least 5 leaves, we can choose edge lengths in the species tree topology such that anomalous gene trees exist . This implies that the gene tree topology appearing most often when considering different genes might not agree with the species tree topology, thus we cannot use a simple majority-heuristic to infer the species tree from a collection of gene trees. Instead we need statistical tools rather than majority rule heuristics for inferring the species tree based on gene trees.
Current methods for inferring species trees from gene trees in this setting can be divided into topology-based and genealogy-based methods, in which the input for a reconstruction algorithm accepts either gene tree topologies or genealogies, i.e., gene trees with branch lengths (coalescence times). Topology-based methods include Minimize Deep Coalescence (MDC) [3, 4], STAR , STELLS , rooted triple consensus  and other consensus and supertree methods [8, 9]. Genealogy-based methods include Bayesian and likelihood methods such as BEST, *BEAST, and STEM [10–12] and clustering and distance-based methods [5, 13–15]. Possible pros and cons of the two approaches are that topology-based methods can be computationally faster and less sensitive to errors in estimating gene trees (and gene tree branch lengths) from sequence data , while methods that use coalescence times, particularly using Bayesian modelling, can be the most accurate when model assumptions are correct .
Another possibility that has been so far unexplored in methods for inferring species trees from gene trees is to use ranked gene trees, in which the temporal order of the nodes of the gene tree (the coalescence times) is used, but not the continuous-valued branch lengths. This approach might therefore be intermediate between purely topology-based methods and genealogy-based methods. By preserving more of the temporal information in the gene tree nodes, the hope is to develop methods that are more powerful than purely topology-based methods and that are still computationally efficient and robust to errors in estimating gene trees and gene tree branch lengths from sequence data.
In , a first step toward developing methods that use ranked gene trees for inferring species trees was taken by providing formulae to calculate the probability of a ranked gene tree given a species tree. The previous work, however, was based on an exponential enumeration of what were called ranked coalescent histories and did not provide an algorithm for computing some of the key terms in the probability of individual ranked histories. In this paper, we improve this previous (computationally inefficient) approach, by providing a method for computing probabilities of ranked gene trees given species trees which is polynomial in the number of leaves using a dynamic programming approach.
Methods for computing probabilities of ranked gene trees efficiently may also be of interest in the context of computing probabilities of unranked gene trees, particularly because no polynomial time algorithm has been found for calculating the probability of a gene tree topology given a species tree under the multispecies coalescent [6, 19–21]. The probability of an unranked gene tree topology can be obtained by summing over all ranked gene tree topologies with the same topology. Thus, for unranked gene trees with particular shapes where the number of rankings increases in polynomial time, using ranked gene trees can potentially increase the speed of computing probabilities of unranked gene trees as well. We note that a completely unbalanced gene tree has only one ranking, while the number of rankings can be exponential in the number of leaves when gene trees become more balanced. Thus, our approach for calculating unranked gene tree probabilities will be most useful for less balanced ranked gene trees.
The bulk of the paper consists of the derivation of the polynomial time method for computing ranked gene tree probabilities. The algorithm is summarized in section ‘An algorithm’. This is followed by a discussion of applications to computing probabilities of unranked gene tree topologies and to inferring ranked species trees under maximum likelihood and a modification to the MDC criterion.
Calculating the probability of a ranked gene tree topology
In the following, we will derive the probability of a ranked gene tree topology given a species tree, . Equations (1, 2, 3, 4, 8, 10) allow the calculation of in time O(n 5). The model giving rise to the gene tree is the multi-species coalescent with constant population sizes . Each species consists of a population of constant size where lineages merge according to the coalescent. Thus, lineages from two different species may coalesce any time previous to the split of the two species.
Notation used in the paper
Species tree with real-valued divergence times
Ranked gene tree (real-valued coalescence times not specified)
The number of leaves of and
Speciation times, with s 1 >⋯> s n−1, let s 0 = ∞
Intervals between speciation times, τ i = [s i ,s i−1)
The number of gene tree lineages at time s i
The number of coalescence events in interval τ i
The ranked gene tree observed from time 0 to time s i
The minimum number of gene tree lineages at time s i
Population z in interval τ i in beaded tree
Internal node (coalescence) with rank i in the gene tree, u 1 is most ancient, u n−1 is the most recent
The number of lineages available for coalescence in population y i,z just after the jth coalescence (considered forward in time) in interval τ i ; k i,0,z is the number of lineages “exiting” at time s i−1
The set of leaves descended from a node of the species tree or gene tree, respectively
For a node u of the gene tree, the node y of the species tree with largest rank such that δ(u) ⊂ δ(y)
For a node y with rank i on the species tree, we denote τ(y) = τ i (the interval immediately above y)
The overall coalescence rate in interval τ i immediately preceding (backwards in time) the jth coalescence
Number of sequences of coalescences above the root of the species tree starting with k lineages
The joint density of coalescence times in interval τ i
Let be a ranked gene tree topology. It is convenient to use the same labels for the leaves of and of . This is a slight abuse of notation, as leaf A of refers to a population (or species), and A of refers to a gene sampled from population A. We denote the nodes of (which are coalescence events) by u 1,…,u n−1, where node u j has rank j, and where higher rank indicates a more recent coalescence. A ranked tree topology can be notated similarly to Newick notation, putting the rank as a subscript for each node, see also Figure 1.
Let be part of a ranked gene tree evolving on a species tree between time s i and time 0 (i.e. the present). consists of ℓ i gene tree lineages at speciation time s i and the coalescent history of in time interval (0,s i ) is consistent with the ranked gene tree . Let g i be the minimum number of lineages required in the ranked gene tree at time s i such that can be embedded into the species tree . Note that n ≥ ℓ i ≥ g i > i. Next we provide a dynamic programming approach for calculating the probability of a ranked gene tree given a species tree. An efficient way to determine the required quantities g 1,…,g n−1 is provided in Section ‘Calculation of g i and k i,j,z ’.
Essentially, in our approach, we traverse the intervals between speciation events going back in time, τ n−1,…,τ 2 (formalized in Theorem 2), and calculate the probability of the appropriate coalescent events occurring in interval τ i based on how many coalescent events happened in the later intervals τ i+1,…,τ n−1 (Theorem 3). Finally with Theorem 1, we account for the most ancestral time interval τ 1.
is the probability for the coalescences above the root appearing in the right order.
For precalculated (ℓ 1 = 2,…,n) the complexity of calculating is thus O(n). Next, we will provide a recursive way to calculate for ℓ 1 = 2,…,n in polynomial time, thus can be calculated in polynomial time.
The complexity of calculating for ℓ 1 = 2,…,n is O(n 3), given we know for all i,ℓ i ,ℓ i+1.
Suppose is known. Given we calculated the probability for ℓ i+1= i + 2,…,n, then calculating for ℓ i = i + 1,…,n requires calculations. Summing up over i = 1,…,n − 1 yields a complexity of . □
It remains to determine . Note that during the interval τ i , we have i branches in the species tree. Let m i be the number of coalescent events in τ i , so m i = ℓ i − ℓ i−1. Let the number of lineages on branch z just after the jth coalescent event (going forward in time) in τ i be k i,j,z . Calculation of k i,j,z can be done efficiently as shown in Section ‘Calculation of g i and k i,j,z ’.
where and .
The density for the coalescence events in interval τ i can be obtained by considering the waiting time to the “next” coalescent event (going backwards in time) as being due to competing exponentials in the different branches, where the coalescence rate within branch z is . Thus, the waiting time until the next coalescent event has rate .
We denote the time between the jth and (j + 1)st coalescent event as v j , where v 0 is the time between s i−1 and the first (least recent) coalescent event in τ i and with being the time between s i and coalescent event m i .
It remains to integrate over v, for which we distinguish between case (i) λ i,0= 0, and case (ii) λ i,0> 0.
where the second line follows because −λ i,j = λ i,0− λ i,j .
which is the same expression as for the λ i,0= 0 case (6). Note that for case (i) we made use of the cumulative distribution function of the hypoexponential distribution, while for case (ii) we made use of the density function of the hypoexponential distribution. Both cases yield the same final expression for , which establishes the proof. □
The probabilities for all possible i, m i and ℓ i (recall that m i = ℓ i − ℓ i−1) are calculated in O(n 5), given all λ i,j .
The quantities λ i,j can be calculated for all possible i, m i , ℓ i and j in O(n 5), given all k i,j,z .
showing that the ratio grows faster than polynomial in n.
Calculation of g i and k i,j,z
Calculation of g i
where τ j < τ i iff j < i, and where I(·) is an indicator function taking the value 1 if the condition holds and otherwise 0. Assuming each lca() operation is O(1) [29, 30], preprocessing allows all lca terms to be computed in O(n) time. Thus, calculating g 1,…,g n−1 can be done, based on Equation 8, in O(n 3).
Calculation of k i,j,z
The value of k i,j,z depends on the number of lineages entering branch i, ℓ i , as well as the number of lineages exiting the branch, and not just on the number of coalescence events in the interval. For example, in Figure 1c, k 2,0,1 = 1 and k 2,1,1 = 2, while in Figure 1d, k 2,0,1 = 2 and k 2,1,1 = 3, although the two gene trees have the same ranked topology and m 2 = 1 for both cases.
To determine the terms k i,j,z we note that the number of coalescences that have occurred more recently than interval τ i is n − ℓ i . In a given interval τ i , we let z (1) and z (2) be the left and right children, respectively, of population z of outdegree 2, and let z (1) = z (2) be the only child of a node z of outdegree 1.
Note that taking the sum over all z is not necessary, as in all but one branch the k i,j,z equals the k i,j+1,z .
In this paper, we provide a polynomial-time algorithm (O(n 5) where n is the number of species) to calculate the probability of a ranked gene tree topology given a species tree, summarized in Section ‘An algorithm’. We now discuss applying these results to computing probabilities of unranked gene tree topologies and to inferring ranked species trees.
Computing probabilities of unranked gene tree topologies
Previous work on computing probabilities of unranked gene tree topologies used the concept of coalescent histories, which specify the branches in the species tree in which each node of the gene tree occurs. An unranked gene tree probability can then be computed by enumerating all coalescent histories and computing the probability of each. The number of coalescent histories grows at least exponentially when the (unranked) gene tree matches the species tree, making this approach computationally intensive. Coalescent histories can be enumerated either recursively (e.g., in PHYLONET  or ) or nonrecursively (COAL ).
A much faster approach using dynamic programming similar to that used in this paper is implemented in STELLS , which conditions on the ancestral configuration in each branch rather than the number of lineages. Here an ancestral configuration keeps track not only of the number of lineages in a branch in the species tree, but also the particular nodes of the gene tree. Different ancestral configurations can potentially have the same number of lineages within a population. Enumerating ancestral configurations turns out to have exponential running time for arbitrarily shaped trees, but the number of ancestral configurations is still much smaller than the number of coalescent histories. When computing probabilities of ranked gene tree topologies, however, the ranking specifies the sequence of coalescence events, leading to a unique ancestral configuration given the number of lineages in a time interval. This fortuitously enables probabilities of ranked gene tree topologies to be computed in polynomial time.
We note that although the number of rankings for a gene tree is not polynomial in the number of leaves in general, the number of rankings can be small for certain tree shapes. For example, if the gene tree has a caterpillar shape, in which each internal node has a leaf as a descendant, then there is only one ranking, and thus computing the ranked and unranked gene tree are equivalent. For a pseudo-caterpillar, a tree made by replacing the subtree with four leaves of a caterpillar with a balanced tree on four leaves , there are only two rankings possible, and for a bicaterpillar, for which the left subtree is a caterpillar with n L leaves and the right subtree is a caterpillar with n − n L leaves, there are rankings. Thus computing unranked gene tree probabilities by summing ranked gene tree probabilities can be done in polynomial time for some tree shapes. We note that for the approach used by STELLS, some tree shapes can also be computed in polynomial time, including the cases we mentioned with a polynomial number of rankings (caterpillar and pseudo-caterpillar). An open question is whether there are any classes of unranked gene trees which have a polynomial number of rankings but an exponential number of ancestral configurations, or vice versa.
Inferring species trees from ranked gene trees
is a multinomial likelihood. Here can be determined with our polynomial-time algorithm, we let denote the ith ranked topology, and n i is the number of times ranked topology i is observed, with . Note in particular that the ranked topology of might differ from the most frequent ranked gene tree topology .
Minimizing (12) as a criterion for the ranked species tree will tend to penalize long edges of the species tree which have multiple lineages persisting through multiple species divergence events. As an example, in Figure 1b, the gene tree has a MDC cost of 1 since there are two lineages exiting the population immediately ancestral to A and B; however the cost according (12) is 2 because there are two edges on the beaded version of the species tree (Figure 2) that each have an extra lineage. In Figure 1c, the gene tree has a MDC cost of 0 for the species tree since it has the matching unranked topology; however, the number of extra lineages from equation (12) is 1. We note that in Figure 1c, interval τ 3, incomplete lineage sorting (and deep coalescence) have not occurred as these concepts are normally used. To capture the idea that coalescence has nevertheless occurred in a more ancient time interval than allowed, we might refer to the coalescence of A and B in Figure 1c as an “ancient lineage sorting” event (rather than incomplete lineage sorting event) or an ancient coalescence rather than a deep coalescence. We could therefore refer to minimizing equation (12) as the Minimize Ancient Coalescence (MAC) criterion, which would provide an interesting comparison to the usual topology-based MDC criterion.
In practice, a method of inferring a species tree from ranked gene trees would require estimating the ranked gene trees. This would require clock-like gene trees, or trees with times estimated for nodes, which can also be inferred under relaxed clock models in BEAST . To account for the uncertainty in the gene trees, the counts for different ranked gene trees could be weighted by their posterior probabilities obtained from Bayesian estimation of the gene trees . Thus, in equation (11), we would let n i k be the posterior probability of ranked topology i at locus k, and use as the estimated number of times that ranked topology i was observed. Similarly, for equation (12), the coalescence cost at a locus could be distributed over multiple topologies weighted by their posterior probabilities.
We thank David Bryant for suggesting the dynamic programming approach to this problem and two anonymous referees for valuable comments, particularly on calculating g i and k i,j,z . JHD was funded by the New Zealand Marsden fund and by a Sabbatical Fellowship at the National Institute for Mathematical and Biological Synthesis, an Institute sponsored by the National Science Foundation, the U.S. Department of Homeland Security, and the U.S. Department of Agriculture through NSF Award #EF-0832858, with additional support from The University of Tennessee, Knoxville. TS was funded by the Swiss National Science Foundation.
- Degnan JH, Rosenberg NA: Gene tree discordance, phylogenetic inference, and the multispecies coalescent. Trends Ecol Evol 2009, 24:332–340.PubMedView Article
- Degnan JH, Rosenberg NA: Discordance of species trees with their most likely gene trees. PLoS Genet 2006, 2:762–768.View Article
- Maddison WP, Knowles LL: Inferring phylogeny despite incomplete lineage sorting. Syst Biol 2006, 55:21–30.PubMedView Article
- Than C, Nakhleh L: Species tree inference by minimizing deep coalescences. PLoS Comput Biol 2009, 5:e1000501.PubMedView Article
- Liu L, Yu L, Pearl DK, Edwards SV: Estimating species phylogenies using coalescence times among sequences. Syst Biol 2009, 58:468–477.PubMedView Article
- Wu Y: Coalescent-based species tree inference from gene tree topologies under incomplete lineage sorting by maximum likelihood. Evolution 2011,.
- Ewing GB, Ebersberger I, Schmidt HA, von Haeseler A: Rooted triple consensus and anomalous gene trees. BMC Evol Biol 2008, 8:118.PubMedView Article
- Degnan JH, DeGiorgio M, Bryant D, Rosenberg NA: Properties of consensus methods for inferring species trees from gene trees. Syst Biol 2009, 58:35–54.PubMedView Article
- Wang Y, Degnan JH: Performance of matrix representation with parsimony for inferring species from gene trees. Stat Appl Genet Mol Biol 2011, 10:21.
- Heled J, Drummond AJ: Bayesian inference of species trees from multilocus data. Mol Biol Evol 2010, 27:570–580.PubMedView Article
- Kubatko LS, Carstens BC, Knowles LL: STEM: Species tree estimation using maximum likelihood for gene trees under coalescence. Bioinformatics 2009, 25:971–973.PubMedView Article
- Liu L, Pearl DK: Species trees from gene trees: Reconstructing bayesian posterior distributions of a species phylogeny using estimated gene tree distributions. Syst Biol 2007, 56:504–514.PubMedView Article
- Liu L, Yu L: Estimating species trees from unrooted gene trees. Syst Biol 2011, 60:661–667.PubMedView Article
- Liu L, Yu L, Pearl DK: Maximum tree: a consistent estimator of the species tree. J Math Biol 2010, 60:95–106.PubMedView Article
- Mossel E, Roch S: Incomplete lineage sorting: consistent phylogeny estimation from multiple loci. IEEE/ACM Trans Comp Biol Bioinf 2010, 7:166–171.View Article
- Huang H, He Q, Kubatko LS, Knowles LL: Sources of error for species-tree estimation: Impact of mutational and coalescent effects on accuracy and implications for choosing among different methods. Syst Biol 2009, 59:573–583.View Article
- Liu L, Yu L, Kubatko LS, Pearl DK, Edwards SV: Coalescent methods for estimating phylogenetic trees. Mol Phylogenet Evol 2009, 53:320–328.PubMedView Article
- Degnan JH, Rosenberg N, Stadler T: The probability distribution of ranked gene trees on a species tree. Math Biosci 2012, 235:45–55.PubMedView Article
- Degnan JH, Salter LA: Gene tree distributions under the coalescent process. Evolution 2005, 59:24–37.PubMed
- Rosenberg NA: Counting coalescent histories. J Comput Biol 2007, 14:360–377.PubMedView Article
- Than C, Ruths D, Innan H, Nakhleh L: Confounding factors in HGT detection: statistical error, coalescent effects, and multiple solutions. J Comput Biol 2007, 14:517–535.PubMedView Article
- Edwards AWF: Estimation of the branch points of a branching diffusion process. J R Stat Soc Ser B 1970, 32:155–174.
- Ross S: Introduction to Probability Models. San Diego: Academic Press; 2007.
- Tavaré S: Line-of-descent and genealogical processes, and their applications in population genetics models. Theor Popul Biol 1984, 26:119–164.PubMedView Article
- Wakeley J: Coalescent Theory. Greenwood Village: Roberts & Company; 2008.
- Pamilo P, Nei M: Relationships between gene trees and species trees. Mol Biol Evol 1988, 5:568–583.PubMed
- Rosenberg NA: The probability of topological concordance of gene trees and species trees. Theor Pop Biol 2002, 61:225–247.View Article
- Semple C, Steel M: Phylogenetics, vol. 24 of Oxford Lecture Series in Mathematics and its Applications. Oxford: Oxford University Press; 2003.
- Harel D, Tarjan RE: Fast algorithms for finding nearest common ancestors. SIAM J Comput 1984, 13:338–355.View Article
- Schiever B, Vishkin U: On finding lowest common ancestors: simplification and parallelization. SIAM J Comput 1988, 17:1253–1262.View Article
- Than C, Ruths D, Nakhleh L: Phylonet: A software package for analyzing and reconstructing reticulate evolutionary relationships. BMC Bioinformatics 2008, 9:322.PubMedView Article
- Drummond AJ, Rambaut A: Beast: Bayesian evolutionary analysis by sampling trees. BMC Evolut Biol 2007, 7:214.View Article
- Allman ES, Degnan JH, Rhodes JA: Identifying the rooted species tree from the distribution of unrooted gene trees under the coalescent. J Math Biol 2011, 62:833–862.PubMedView Article
This article is published under license to BioMed Central Ltd. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.