Showing posts with label E3b. Show all posts
Showing posts with label E3b. Show all posts

Friday, February 21, 2014

YDNA E-M123; A closer look

E-M123 (as well as E-M34) was first discovered by Underhill(2000) and is found with a low to medium frequency distribution in East Africa and the Middle East, while it has a low frequency distribution in North Africa and Europe.

Phylogeny:
Figure 1 - Current and previous E-M215 phylogenetic structure 

Figure 1 shows a comparison of the basic phylogeny of E-M215/M35 as was known before 2011 (a) and after (b), with a 'who and when' key for the Discovery of the UEPs. Notice the impact the rearrangement has on the phylogenetic placement of E-M123, specifically the fact that E-M123 is shown to have a more recent common ancestor with the East and Southern African variants of E-M35, i.e. E-V42 and E-M293, before it does with any of the other variants of E-M35.

Previous publications:

While it is unfortunate that all of the research that has previously been published on E-M123 was done under the consideration of the older (and rather out of date) configuration of the basic structure of E-M35, it is still worth while to look at articles that have tried to untangle the origins and history of this lineage, of these, 3 come to mind:

Wednesday, May 8, 2013

Another Extensive thesis on East African DNA


It was brought to my attention last week, thanks to a comment on this blog made by the user 'Umi', that another thesis on East African DNA variation was publicly available online:

Complex Genetic History of East African Human Populations

This is also an extensive thesis with a wealth of information akin to Plaster's thesis, the primary differences being that this one was more focused on parts of East Africa that are found further to the South of Ethiopia, and in addition to uni-parental analysis, it also included some Autosomal model-based inference, albeit of quite low resolution in today's standards; 848 microsattelites and 479 indels (refer to Tishkoff et al. 2009 for marker details).

Due to the extensive nature of the report I haven't had a chance to cover its entire scope, instead, for starters, I have first focused on the YDNA data by creating a relative frequency chart from the results reported in Fig. 3.3.2. 

Several things to initially point out here,

  • The report outlines the discovery of 4 new SNPs, TL1-4. The first two were found in Haplogroup B and downstream from B-M150 and B-M112 respectively. The last two, TL3 and TL4, were found in haplogroup E and downstream from E-U174 and E-V32 respectively. Incidentally, the fourth SNP that is under E-V32, TL4, could potentially be the same as Z808/Z809 as identified recently by the geneological community, however, as the report does not give the Y-Chromosome location of the SNP in a NCBI Build 36/37 format, this can not be verified, at least by me, at the moment.
  • A couple of the frequency results in Fig. 3.3.2 do not add up, in particular, the frequency results for the Boni and the Baggara, but also to a lesser extent for the Kanuri and Teita.  I have labeled the missing frequency results with a “?” in the relative charts for those specific populations.
  • The Burji and Konso are labeled as being only from Kenya throughout the report, however most Burji are from Ethiopia, and the Konso are exclusively found in Ethiopia, I have reflected this in the charts.
  • STR data is not readily available to perform TMRCA estimates on, however, some TMRCA results are reported using Zhivotovsky's rates in Table 3.3.1, nevertheless, these are estimates only for different lineages found in the dataset for all the samples and not necessarily comparing TMRCAs in the different populations under study.
  • J-M62, while a subclade of J-M267, is not the main subclade of J-M267 found in East Africa, that would be J-P58, therefore, the results for J-12f2.1 (x M62, M172) reported, may after all be, or largely include, J-P58 lineages, off-course those results could also include variants of J-M267 other than J-P58 and J-M62 as well since the SNP was not directly tested. 
  • E-P2* lineages are abundantly found (> 30%) in the Konso, Burji and Mbugwe, however on closer examination and correlation with current data, these could be E-M329, E-V38* or even E-M215*, as none of these SNPs were directly tested. Genuine E-P2* lineages would be positive for E-P2 and negative for V38 and M215 (See Trombetta et al. 2011)
  • Similarly, the E-M35* lineages reported could be members of relatively newly discovered lineages of E-Z830*( See this post for details), or some of the untested variantes of E-M35, i.e.  E-V42, V92 and maybe even E-V68 (x M78)

Thursday, February 21, 2013

The Zhivotovsky Multiplier


It is reported that Zhivotovsky's effective mutation rate [1] has the effect of increasing the TMRCA of a lineage, as computed by the use of Microsattelite Genetic Distances[2], by a factor of 3-4 fold relative to TMRCAs computed via mutation rates observed in pedigree and family studies [3].

By utilizing my TMRCA calculating program, I want to explore,
  1. What effect does different marker combinations have on this multiplier ?
  2. What effect does marker size have on this multiplier ?
  3. Is there a variation in this multiplier for different data-sets?

First, to ensure that my program correctly calculates the TMRCA when the Zhivotovsky mutation rate of 0.00069 is applied to all the markers in my database consistently (versus only the marker specific Pedigree mutation rates I have thus far been utilizing), I attempted to replicate the TMRCA computations of the following publication;


Friday, February 8, 2013

Sudan YDNA

This is from a relatively old study, but it seems that it is the most comprehensive YDNA breakdown we have of North and South Sudan to date.

Y-chromosome variation among Sudanese: restricted gene flow, concordance with language, geography, and history. Hassan (2008)

Here is a map of the populations tested from Fig.1 of the Study
Populations Studied

Here below is the phylogeny (as known back in 2008) of the SNPs tested, note that those in bold; E-M75, E-P2, G-M201 and T-M70 were NOT tested in the study.

SNPs tested (except those in bold)
The E-M78+ cases from above were also tested for Cruciani's V-Series SNPs as well for further resolution,


Cruciani's V-Series SNPs (2007)

Some notes:


  • The high level (38%) of E-M215 (x M78) in the Borgu is quite intriguing, I wonder what variant/s of E-M215 it is?
  • Almost all the J-12f2(x M172) should be J-M267.
  • B-M60 is found in Southern Nilo-Saharan speakers and not the North Western ones, while A-M13 is found in both.
  • The F-M89(x M52,M170,I2f2, M9) found in the north is also interesting, although it could possibly be G-M201, at least part of it.
  • E-V22 has a relatively high presence in these samples, even when compared to the Egyptian samples from Cruciani '07, and most certainly higher than its presence in Ethiopia.
  • The High presence of E-V12 (x V32) is also concordant with its putative area of origin, all the E-M78 found in the Nuer and the Copts is of this variety.
  • The presence of E-M78* in the Masalit and the Nuba is notable.
  • Off course the strangest result is the 54% R-M173 (x P25) in the Fulani, this could be some R1b*(R-M343), or some type of R1a, the latter would be very out of place for the region, while the former could be reconciled with the presence of more downstream R1b variants in Africa. 


Monday, February 4, 2013

A speculative superimposition of E-M35 variants onto Afroasiatic.

Here is a speculative superimposition of the variants of YDNA E-M215/M35 (E1b1b/1) onto an Afroasiatic internal classification, Lionel Bender's (1997) classification. 


The red question marks represent a less unsure fit.

Saturday, January 5, 2013

TMRCA calculations from Plaster NRY data : Correcting an Error


Previously, I had computed TMRCAs for the YDNA STR data from the additional material that was provided along with Dr.Chris Plaster's thesis. However, after a brief communication with the author, I found out that the marker order of the STRs in the excel file was reported wrongly, the correct order for the markers are thus as follows:

DYS19 DYS388 DYS389I DYS389II DYS390 DYS391 DYS392 DYS393 DYS437 DYS438 DYS439 DYS448 DYS456 DYS635 Y GATA H4

This changes my TMRCA calculations because I am not computing the coalescent using a generic mutation rate that is equivalent for all the markers, but rather each marker has its own mutation rate attributed to it.

When I rerun my program using the newly corrected order above I get the following:


As can be seen, using the new order of markers generally reduces the number of generations to coalescent for the Plaster data-set. The previous observation of a relatively lower TMRCA for the haplozone data of E-M123 versus that of the E-M34 Plaster data-set largely disappears. 

To check if the fact that the high number of samples (129) present in the E-M123 haplozone data-set was skewing the results, I took 23 random samples (which equals the same number of samples available in the Plaster E-M34 data-set) from the larger E-M123 Haplozone dataset and re-run the TMRCA calculations on just those samples, I repeated this process 300 times, only 28% of the runs yielded a mean TMRCA less than the E-M34 Plaster data-set, if sample size was skewing the results I would expect >50% of the runs to have a mean TMRCA less than that of the E-M34 plaster dataset.

That said, the E-M34 Plaster data-set still had a relatively higher generations to coalescent than the E-M84 Haplozone dataset, E-M84 is a subclade of E-M34 and a high majority of haplotypes that belong to E-M34 also test positive for the E-M84 SNP (at least for the non-African E-M34 haplotypes that we know of).

Other than that, the new, and corrected, ordering of the markers did not have much impact in relative TMRCA terms between the Plaster and Haplozone/FTDNA data for the other lineages I had tested.

Tuesday, December 11, 2012

National Geographic fesses up on the origin of E-M35

In their second phase of the massive global scale genetic testing project, Geno 2.0: The Greatest Journey Ever Told, National Geographic has finally fessed up to the most parsimonious explanation to the origin of YDNA haplogroup E1b1b1, this is good news, even if it took 8 years to do so, i.e. about 8 years after the publishing of the first detailed paper on E-M35.

In the first phase of the Geneographic project, launched in 2005, E-M35's origin was explictly stated as the following :

"The man who gave rise to marker M35 was born around 20,000 years ago in the Middle East. His descendants were among the first farmers and helped spread agriculture from the Middle East into the Mediterranean region."
Original E-M35 National Geographic Description


You can read what it reads today in the screen shot below:
Current E-M35 Nat. Geographic Description

There is also the sentence, "Today, in keeping with its place of origin, this line is common among Afro-Asiatic speakers", could the part, 'in keeping with its place of origin', be also a 'nudge' at the very distinct,  and in my opinion, strong, possibility that Afroasiatic may have originated in East Africa as well ?
If so, this would be a first for a major outlet like Nat Geo and others, even though, renowned Afroasiatic experts like Greenberg, Ehret, Blench et. al had said this for decades.

Update: Another point that is odd in their new phylogeny seen above, is the ordering of some of the NRY SNPs leading up-to V12, the SNPs leading up-to P147 are in standard sequence, i.e the sequence M42 > M168 > M203 > M96 > P147, is common knowledge, however P177 is listed as downstream of P2, where common knowledge says it is the reverse, i.e. P147 > P177 > P2, instead of P147 > P2 > P177. Similarliy, M215 is not known to be a subclade of M35.1 but rather the reverse, so overall, their sequence should read as follows : M42 > M168 > M203 > M96 > P147 > P177 > P2 > M215 > M35.1 > M78 > V12. Unless off-course they have found some samples that upset the standard NRY SNP sequence leading upto E-V12 that we do not know about yet.

Tuesday, June 19, 2012

Finding the TMRCA of Ethiopian YDNA lineages using an ASD method.


I have been lately working on computing TMRCAs using an ASD or average square difference method on publicly available Y-STR haplotypes. The premise for finding the TMRCA using the ASD method is quite straight forward and easy to understand, a putative ancestral haplotype is calculated for a given dataset and the repeat of each sample at each marker in the dataset is subtracted from this ancestral haplotype, this result is then cumulated and divided by the number of samples and the marker specific mutation rate, the process is repeated for every single marker in the dataset and the mean is then multiplied by an assumed years per generation length, the formula below articulates this method:
TMRCA formula (ASD method)
 
Where;
N= Total number of Samples
Z= Total number of Markers
L0= Putative Ancestral Haplotype (Median or Modal repeats)
L= Individual sample haplotype repeats
m= Marker Specific Mutation Rate
G= Years / Generation

The biggest variable here, other than the sampling strategy of a given dataset, are the several marker specific mutation rates that are available. The process of selection of a correct mutation rate is an unsettled issue, I have therefore utilized 4 sets of mutation rates that were compiled by Paul Newlin, a collaborator at the E3b Project, these rates come from several different publications and you can read about them here for more detail:
 
  1. The Chandler Mutation Rates:
  2. Stafford Bayesian Mutation Rates:
    Essentially a compilation of other mutation rates
  3. Burgarella & Navascués Mutation Rates:
  4. Ballantyne Mutation Rates:
In order to have an analogously accurate comparison of the TMRCAs between the different publications, I had to weed out and intersect the available markers from above with markers that are found in the public domain. This essentially left me with the following 46 markers that intersected with all 4 of the above sets of rates as well as the 66 markers that are widely used:
406s1 , 19 , 388 , 389-1 , 389-2 , 390 , 391 , 392 , 393 , 426 , 436 , 437 , 438 , 439 , 442 , 444 , 446 , 447 , 448 , 450 , 454 , 455 , 456 , 458 , 460 , 472 , 481 , 487 , 490 , 492 , 511 , 520 , 531 , 534 , 537 , 557 , 565 , 568 , 572 , 578 , 590 , 594 , 617 , 640 , 641 and gatah4.

In addition, since the Chandler mutation rates had a complete intersection with the 66 widely used markers, an additional 66 marker Chandler set was independently used that included the following markers in addition to the 46 listed above:
385a , 385b , 459a , 459b , 449 , 464a , 464b , 464c , 464d , ycaiia , ycaiib , 607 , 576 , 570 , cdya , cdyb , 395s1a , 395s1b , 413a and 413b.
  
Haplogroups A, E and J, cover well over 90% of the YDNA lineages found in Ethiopia. More specifically within these haplogroups, I was more interested in finding the TMRCA for A-M13, E-M35 and J1-M267, as these lineages cover over 70% but under 80% of said lineages, whereas the remaining 20-30% of lineages found in Ethiopia belong to E1b1*(x E1b1b,E1b1a1), other types of E lineages like E2 and E*, and some specific clades that belong to haplogroups B,T and J2.

Wednesday, November 4, 2009

Cruciani et. al 2007

Cruciani et. al 2007, discusses the E1b1b1a (E-M78) sub lineage of E1b1b (E-M215) in further detail. The entire data for East Africa comes from the Cruciani et. al 2004 study, while for North Eastern Africa, in addition to the samples taken from the same paper, it includes some newer samples from; Libyan Jews, Libyan Arabs, Egyptian Berbers, Egyptians from Baharia and Egyptians from Gurna Oasis. The Phylogeny of E-M78, is further finely resolved into a new series of "V" sub clades as seen below (Taken from Figure 1)


A "Corridor for bidirectional  migrations between East Africa and North East Africa" is proposed as a result of the paper's findings:

"E-M78 belongs to clade E3b (E-M215). On the basis of robust phylogeographic considerations, an eastern African origin has been proposed for E-M215 (Underhill et al. 2001; Cruciani et al. 2004), with a coalescence time of 22.4 ky (95% C.I. 20.9-23.9 ky; recalculated from Cruciani et al. 2004, see Materials and Methods). A north-eastern African origin for haplogroup E-M78 implies that E-M215 chromosomes were introduced in north-eastern Africa from eastern Africa in the Upper Paleolithic, between 23.9 ky ago (the upper bound for E-M215 TMRCA in eastern Africa) and 17.3 ky ago (the lower bound for E-M78 TMRCA here estimated, fig. 1). In turn, the presence of EM78 chromosomes in eastern Africa can be only explained through a back migration of chromosomes that had acquired the M78 mutation in north-eastern Africa."

1) E1b1b1a (E-M78) frequencies in East African populations.

Above, it is clear that only sub clades E-V32 and E-V22 dominate in East African E-M78.
Notice here also the combination of the "Borana Oromo Kenya" (N=7) and the "Ethiopian Oromo" (N=25) samples taken from Cruciani et. al 2004, as a new group coined as "Borana/Oromo (Kenya/Ethiopia)" (N=32) in this study.

2) E1b1b1a (E-M78) frequencies in North East African populations.
                                      
"North Eastern Africa", has a much more richer subclade diversity of E-M78, than "Eastern Africa", which makes sense as to the current assumption of E-M78's origin in the Egypt/Sudan locality. See also Battaglia et al. (2008).

3) Microsatellite Networks for E-V12, E-V22 and E-V65.
 
(Taken from Fig. 3)
Microsatellite networks of haplogroups E-V12 (A); E-V22 (B); and E-V65 (C). In network (A), a dotted circle includes all of the E-V12 chromosomes carrying the V32 mutation. Branch lengths are proportional to the number of one-repeat mutations separating 2 haplotypes. Each circle area is proportional to the frequency of the sampled haplotype.

4) Comparing Cruciani '07 results to the E3b Project.


Notice on the comparison of the results of Cruciani '04 to the E3b project that there is an imbalance of E-M78 lineages (in favor of the E3b Project), above we can see that this imbalance is further characterized by the marked abundance  of E-V13 lineages in the E3b Project, this makes sense because E-V13 is for the most part found only in Europe, and Europeans have better financial and technological access to participate in private DNA testing.

5) Further reading and sources on E-M78.
Hassan et. al 2008:
"Y-Chromosome Variation Among Sundanese: Restricted Gene Flow, Concordance with Language, Geography, and History"

Sanchez et. al 2005:
"High Frequencies of Y Chromosome Lineages Characterized by E3B1, DYS19-11, DYS392=12 in Somali Males"

Battaglia et al. 2008:
"Y-chromosomal evidence of the cultural diffusion of agriculture in southeast Europe"

Wednesday, October 21, 2009

Cruciani et. al 2004

Cruciani et. al 2004, is perhaps the best comprehensive study that we currently have with respect to the signature paternal haplogroup that is most abundant among males of Eastern and Northern Africa, E1b1b.

The haplogroup is most likely of indigenous Eastern African origin according to the paper:

We obtained an estimate of 25.6 thousand years (ky) (95% CI 24.3–27.4 ky) for the TMRCA of the 509 haplogroup E3b chromosomes, which is close to the 30 +/-6 ky estimate for the age of the M35 mutation reported by Bosch et al. (2001) using a different method. Several observations point to eastern Africa as the homeland for haplogroup E3b—that is, it had (1) the highest number of different E3b clades (table 1), (2) a high frequency of this haplogroup and a high microsatellite diversity, and, finally, (3)the exclusive presence of the undifferentiated E3b* paragroup.

Note: E3b here refers to the Y- DNA lineage that is defined by E-M215 (as seen below taken from Figure 1) and is referred by current nomenclature as E1b1b.



1)  E1b1b (E-M215) frequencies in East African populations.



It's important to note here that after the publishing of this paper, a new SNP downstream of E-M35 and parallel to E-M78, E-M81, etc... was found by Henn et. al  2008, this new sub lineage of E1b1b1 (E-M35) was coined as being E1b1b1g (E-M293), therefore, some of the lineages labeled as E-M35* above could indeed turn out to be part of this newly discovered lineage, however, this is unlikely because E-M293 is more predominent in areas further south from the Horn of Africa according to Henn et. al.

2) E-M96 (xM215) + E-M215 with down stream clades


E-M96 is the upstream ancestor, 3rd node up (including the E-M215 node),  of E1b1b. It is the macro haplogroup  which includes some 70% of the male lineages found on the African continent.
The purpose of the above graphic is to  give a general sense of the ratio of E lineages without the E-M215 mutation to that of E lineages with the E-M215 mutation for populations in East Africa.
Notice that this ratio is at a Maximum (~5.3) for "Bantus from Kenya", while it is at a Minimum (0) for "Somalis", while the average for all the Eastern African populations sampled is ~0.5.
The paper states that the Kenyan Bantu data comes from "Human Genome Diversity Project/CEPH DNA panel (Cann et al. 2002)", while I have not (as of yet) seen the precise data, one can speculate that the majority of the E-M96 lineages that are non E-M215 found in the Kenyan Bantu probably belong to E1b1a (E-M2).
The paper also states that for the Ethiopian Jews (A.K.A Bete Israel) the data comes from "Cruciani et al. 2002", from this paper we can glean that the lineages described as E-M96 but not downstream of E-M215, all belong to E3* (x E-M2), which in current nomenclature would translate to E1b1* (E-PN2*), while noting that a slight chance exists that said lineages could belong to the later found subclade of E1b1 now known as E1b1c (E-M329).
Unfortunately, for the sampled populations (other than the Kenyan Bantu and Ethiopian Jews), no further data is given to clarify the composition of the E-M96 (x E-M215) lineages.

3) Breakdown of the 515 E-M215 lineages found in this study.


This paper serves as a benchmark for E1b1b, since it did the most extensive study of the lineage on a global level:
"We explored the phylogeography of human Y-chromosomal haplogroup E3b by analyzing 3,401 individuals from five continents."
Out of the 3,401 individuals, 515 or 15.1% were found to belong to E1b1b (E-M215) :
"Five hundred fifteen haplogroup E3b subjects were identified and further analyzed for the biallelic markers M34, M78, M81, M123, M281 (Underhill et al. 2000; Seminoet al. 2002), and V6."
It should be noted again that the 9% of E-M35* lineages found include those also found in Southern Africa among the Kung, Khwe and Bantu, but as stated above, these E-M35* lineages probably belong to the later found subclade E-M293 per the Henn et. al 2008 paper, in that event the share of E-M35* lineages found would probably drop by ~3%.

4) Microsatellite Networks for the main sub-lineages and E-M35*

(taken from Figure 2):
Microsatellite networks of E3b haplogroups. A, E-M35*. B, E-M78. C, E-M81. D, E-M34. Reduced-median and median-joining procedures (Bandelt et al. 1995, 1999) were applied sequentially. A haplogroup-specific weight proportional to the reciprocal of microsatellite variance was used in the construction of the networks. The E-M78 unweighted network (not shown) gave the same quadripartite structure. Unassigned chromosomes (B) showed an intermediate position between clusters "alpha" and "delta" in the unweighted network. Microsatellite haplotypes are represented by circles, with areas proportional to the number of individuals harboring the haplotype. Branch lengths are proportional to the number of one-step mutations separating two haplotypes.
The Alpha, Beta, Delta and Gamma clusters of E-M78 found in this study were investigated later in Cruciani et. al 2007 and were assigned to a new set of SNP's downstream of E-M78 known as the V-series. As a matter of fact,  the sole focus of Cruciani et. al 2007 with respect to E lineages was on E-M78.

5) Comparing Cruciani '04 results to the E3b Project



The E3b project, has an on going and a real time database of people who have tested positive for the E-M35 lineage and further downstream sub-lineages. As such, it is a good tool for testing the findings of Cruciani '04. Above, you can see such a test done by me comparing the relative frequency of E-M35 subclades found in Cruciani '04 to that of the E-M35 deep clade test results from the E3b project as of October 2009.

Further peer reviewed sources on E1b1b and it's subclades can be found below: