Peer-reviewed Article
How should gaps be treated in parsimony? A comparison of approaches using simulation
Abstract
Simulation with indels was used to produce alignments where true site homologies in DNA sequences were known; the gaps from these datasets were removed and the sequences were then aligned to produce hypothesized alignments. Both alignments were then analyzed under three widely used methods of treating gaps during tree reconstruction under the maximum parsimony principle. With the true alignments, for many cases (82%), there was no difference in topological accuracy for the different methods of gap coding. However, in cases where a difference was present, coding gaps as a fifth state character or as separate presence/absence characters outperformed treating gaps as unknown/missing data nearly 90% of the time. For the hypothesized alignments, on average, all gap treatment approaches performed equally well. Data sets with higher sequence divergence and more pectinate tree shapes with variable branch lengths are more affected by gap coding than datasets associated with shallower non-pectinate tree shapes.
Full Citation
Ogden, T.H., and M.S. Rosenberg (2007) How should gaps be treated in parsimony? A comparison of approaches using simulation. Molecular Phylogenetics and Evolution 42(3):817–826. [Erratum: 2008, 46(2):807-808.]
DOI
Associated Software
IndelCoderAssociated Data
Direct Download
Download PDFPubMed Record
PMID: 17011794Google Scholar Data
104 citations as of 2024-011-19
Altmetrics