Dated Ancestral Trees from Binary Trait Data and Their Application to the Diversification of Languages
- 10 April 2008
- journal article
- Published by Oxford University Press (OUP) in Journal of the Royal Statistical Society Series B: Statistical Methodology
- Vol. 70 (3) , 545-566
- https://doi.org/10.1111/j.1467-9868.2007.00648.x
Abstract
Summary: Binary trait data record the presence or absence of distinguishing traits in individuals. We treat the problem of estimating ancestral trees with time depth from binary trait data. Simple analysis of such data is problematic. Each homology class of traits has a unique birth event on the tree, and the birth event of a trait that is visible at the leaves is biased towards the leaves. We propose a model-based analysis of such data and present a Markov chain Monte Carlo algorithm that can sample from the resulting posterior distribution. Our model is based on using a birth–death process for the evolution of the elements of sets of traits. Our analysis correctly accounts for the removal of singleton traits, which are commonly discarded in real data sets. We illustrate Bayesian inference for two binary trait data sets which arise in historical linguistics. The Bayesian approach allows for the incorporation of information from ancestral languages. The marginal prior distribution of the root time is uniform. We present a thorough analysis of the robustness of our results to model misspecification, through analysis of predictive distributions for external data, and fitting data that are simulated under alternative observation models. The reconstructed ages of tree nodes are relatively robust, whereas posterior probabilities for topology are not reliable.All Related Versions
This publication has 18 references indexed in Scilit:
- Temporal phylogenetic networks and logic programmingTheory and Practice of Logic Programming, 2006
- From words to dates: water into wine, mathemagic or phylogenetic inference?Transactions of the Philological Society, 2005
- Perfect Phylogenetic Networks: A New Methodology for Reconstructing the Evolutionary History of Natural LanguagesLanguage, 2005
- Phylogenetic trees based on gene contentBioinformatics, 2004
- Language-tree divergence times support the Anatolian theory of Indo-European originNature, 2003
- Indo‐European and Computational CladisticsTransactions of the Philological Society, 2002
- Phylogenies from Restriction Sites: A Maximum-Likelihood ApproachEvolution, 1992
- An Indoeuropean Classification: A Lexicostatistical ExperimentTransactions of the American Philosophical Society, 1992
- Evolutionary trees from DNA sequences: A maximum likelihood approachJournal of Molecular Evolution, 1981
- The Basis of GlottochronologyLanguage, 1953