Showing posts with label South Africa. Show all posts
Showing posts with label South Africa. Show all posts

Friday, April 3, 2015

Revised Timeline and Distribution of the Earliest Diverged Human Maternal Lineages in Southern Africa

Abstract

The oldest extant human maternal lineages include mitochondrial haplogroups L0d and L0k found in the southern African click-speaking forager peoples broadly classified as Khoesan. Profiling these early mitochondrial lineages allows for better understanding of modern human evolution. In this study, we profile 77 new early-diverged complete mitochondrial genomes and sub-classify another 105 L0d/L0k individuals from southern Africa. We use this data to refine basal phylogenetic divergence, coalescence times and Khoesan prehistory.
Our results confirm L0d as the earliest diverged lineage (~172 kya, 95%CI: 149–199 kya), followed by L0k (~159 kya, 95%CI: 136–183 kya) and a new lineage we name L0g (~94 kya, 95%CI: 72–116 kya). We identify two new L0d1 subclades we name L0d1d and L0d1c4/L0d1e, and estimate L0d2 and L0d1 divergence at ~93 kya (95%CI:76–112 kya). We concur the earliest emerging L0d1’2 sublineage L0d1b (~49 kya, 95%CI:37–58 kya) is widely distributed across southern Africa. Concomitantly, we find the most recent sublineage L0d2a (~17 kya, 95%CI:10–27 kya) to be equally common. While we agree that lineages L0d1c and L0k1a are restricted to contemporary inland Khoesan populations, our observed predominance of L0d2a and L0d1a in non-Khoesan populations suggests a once independent coastal Khoesan prehistory. The distribution of early-diverged human maternal lineages within contemporary southern Africans suggests a rich history of human existence prior to any archaeological evidence of migration into the region. For the first time, we provide a genetic-based evidence for significant modern human evolution in southern Africa at the time of the Last Glacial Maximum at between ~21–17 kya, coinciding with the emergence of major lineages L0d1a, L0d2b, L0d2d and L0d2a.

Link (Open Access) 

Tuesday, February 25, 2014

mtDNA from Southern Africa

Reference mtDNA from Southern Africa from the pre-print "Migration and interaction in a contact zone: mtDNA variation among Bantu-speakers in southern Africa" (Thanks to Maju for the referral)

Friday, February 15, 2013

Gradient Maps for African ADMIXTURE components

Here below are gradient maps for my last African ADMIXTURE run, Africa_V2b, courtesy of a demo download of Mapviewer7 . The Kriging method was used for Gridding and 'Grid Z limits' mode was used for color mapping.

Sampled Population's Index

Sampled Population's Location

PCA for the FST distances
generated by ADMIXTURE  

West-Africa Cluster Freq.

Nilo-Saharan Cluster Freq.

East-Africa-2 Cluster Freq.

North-Africa Cluster Freq.

Khoi-San Cluster Freq.

Omotic Cluster Freq.

Mbuti-Pygmy Cluster Freq.

Biaka-Pygmy Cluster Freq.

Hadza Cluster Freq.

East-Africa-1 Cluster Freq.
Isometric view of the MDS plot
 for all Populations sampled


UPDATE (02/18/2013) : Below are gradient maps for the first African ADMIXTURE run, Africa_V1, courtesy of a demo download of Mapviewer7 . The same options as above were used both for gridding and color mapping.

Monday, October 8, 2012

YDNA from Southern Africa

Naidoo et. al (2010) reports YDNA from 3 different groups in Southern Africa with a fair amount of resolution.
Electropherogram and phylogeny

Here below are the frequencies found:

Friday, June 22, 2012

Intra African Genome-Wide Analysis, V2

See Also : Intra African Genome-Wide Analysis, V1


Population References and First Pass K10 Analysis



K2 - K10 Analysis

Tuesday, February 28, 2012

Intra African Genome-Wide Analysis


The primary purpose of studying Haplogroups (NRY and mtDNA) is to describe population movements, AKA Phylogeography . Autosomal DNA on the other hand, gives a rather ambiguous indication of a certain populations Paternal and Maternal history, since the chromosomes used undergo genetic recombination and can not be traced back to a single common ancestor. But still, there are drawbacks in just using NRY or mtDNA to study the history of a given population, and that is that they constitute only of a single Loci, which thereby reduce the effective population size relative to the Autosomes.

To this end, I have utilised publicly available Genome-Wide SNP data to get further insight into the population structure of Africa which may not be fully understood only from the data of uni-parental markers that we have. Perhaps the best published work out there with respect to African Autosomal Genome-wide data is that from Tishkoff (2009), this important paper found 14 ancestral Clusters in the African continent using the most diverse African dataset to date, however, the paper used Autosomal Microsatellites and a handful of SNPs.

On a publicly available dataset, I carried out two of the most popular approaches to help investigate population structure in Africa using Autosomal genome-wide data; (1) The non-parametric approach known as Principal Components or Multi Dimensional Scaling, which uses a Matrix whose elements are the quantification of the genetic similarity between pairs of individuals, and on which such a Matrix is used in order to perform a Principal Component Analysis upon, and (2) An explicit model based population structure analysis using the software ADMIXTURE, where individuals are assumed to come from one of K discrete populations and where population membership and allele frequencies are estimated using a Bayesian modeling strategy.

DATASET
A super set of the Data I used can be downloaded from here :http://dl.dropbox.com/u/23271596/ref.zip
The global Data Set, compiled by this blog author, contains publicly available data from 3970 individuals from around the world typed for 27,022 Autosomal SNPs, which can be found all over the 22 pairs of chromosomes (but not uniformly). I then utilized PLINK to perform the following on the above Data Set:
  1. Removed all Non-Continental African populations.
  2. Removed 18 Tunisians from Henn (2011) as previous analysis had shown independent cluster formation by this group, perhaps a sign of inbreeding.
  3. Removed 15 Morrocan Jews that came from Behar (2010) for the same reason as above.
  4. Kept SNPs above 99.46% genotyping success rate.
  5. Excluded SNPs in linkage disequilibrium (r2>0.5) with nearby markers in a window of 50 SNPs (advanced by 5 SNP).
  6. Added a handful of private African samples that took their genetic test with the Personal Genomics Company, 23andME. (The results of which I can not unfortunately publish in this post)

The above procedures left me with a core (public) Dataset of 1,065 Individuals from Africa and 26,129 SNPs for analysis. The complete SNPs typed for these individuals can be retrieved from: Behar (2010), Hapmap III, Henn (2011), HGDP and Xing (2010)
Furthermore, geographically, 362 were from East Africa, 304 from West Africa, 158 from North Africa, 142 from Central Africa and 99 from South Africa. Linguistically, the dataset contained 536 Niger Kordofanian speakers, 212 Nilo-Saharans , 211 AfroAsiatic speakers, 89 Khoisans and 17 Hadza.

Update: Reference Populations and Key:
 

MDS Analysis
The data for the MDS analysis was generated using PLINK, while the plots were generated using GNU OCTAVE. A 3 dimensional MDS plot for the dataset can be seen below, all populations are labelled according to their Median Co-ordinates.
Here, we can see that the first component, C1, separates East and North Africans from West/Central/South Africans, while the Second Component separates the divergent hunter gatherers (San,!kung, pygmies and Hadza from the rest), this may be more clearer on the two dimensional C1 vs C2 plot below,
 
The third Component C3, separates East Africans from all the rest, as more clearly seen on a C1 vs C3 plot below,

 
Model Based Analysis
The model based analysis was carried out for K=10 using ADMIXTURE, thus 10 clusters were generated from the Dataset, I took the liberty to name these clusters, some on a geographic basis, others on a linguistic basis and still others on a subsistence basis, there is obviously a lot of fluidity associated in naming a cluster, so it shouldn't be taken as something written in stone.

A PCA plot for the FST distances generated by ADMIXTURE for the 10 clusters can be seen below,

 
The extreme positioning of the 'Hadza' cluster is indeed striking, followed by the 'KhoiSan' and 'Pygmy' clusters. The 'West African', 'West-Central African' and 'Eastern Bantu' clusters are quite close to each other as can be expected. The divergence of the North African cluster from East Africa can be explained by the significant extra African Admixture North Africans have as evidenced by the amount of their direct maternal ancestries coming from Europe and the Near East, while a majority of their paternal Ancestry comes from East Africa (Namely, E1b1b).

Below are the Median proportions for the 10 clusters generated by ADMIXTURE for the 45 uniquely entered African populations categorised according to their 5 respective regions.


The unclear abbreviations above for the samples of EtA, EtO and EtT are respectively Ethiopian Amharas, Ethiopian Oromos and Ethiopian Tigrayans, these samples (as well as the Ethiopian Jews, AKA Beta Israel) come from Behar (2010), in addition, the EtO samples purportedly come from the southern most tip of Ethiopia close to the Kenyan border. The dominance of the North African cluster in Ethiopians is not much of a surprise, as it is well known that Ethiopia is a genetic conduit between East and North Africa.

Here, both the mbuti and biaka pygmies form completely independent clusters, which is not unexpected as they are some of the most divergent populations even on a global basis. Also to note, is the slight 'North African' Affinity of the Hema and the 'West African' affinity of the Bulala and Mada.


Many of the non-Khoisan South African populations in the Dataset show affinities to both the 'Eastern Bantu' and 'Central-West African' clusters in almost equal proportions, which is interesting.

As seen in the PCA plots of the FST distances, the 'Central-West African', the 'West African', as well as the 'Eastern Bantu' clusters are close. The Dogon population however shows the least amount of the 'Central-West African' cluster and is almost completely dominated by the 'West African' cluster, which the reverse is true for the Igbo and Yoruba. Similarly, the Fulani show almost none of the 'Central-West African' cluster but rather, are mostly dominated by the 'West African' cluster, with the difference from the Dogon being that the Fulani have a significant affinity with the 'North African' cluster rather than the 'Central-West African' one.

 
In the last graphic above, we can see a geographic affinity of North West Africans with West African based clusters and North East Africans with East African dominant clusters, as to be expected. As stated before however, the 'North African' cluster itself is likely a both ancient and recent synthesis of East African, European and Near Eastern Affinities.

Conclusion
I learned quite a bit on the population structure of Africa from this exercise but there is a lot more room left for improvement:
  1. The SNPs that are typed using almost all genotyping arrays are Eurasian biased, as they were first found in Europeans, as time goes on, more African specific SNPs will be discovered and their use in genome-wide analysis will change these results.
  2. More samples are needed, especially from both South and North Sudan, all along the Sahel belt, Tuaregs, different Omotic speakers from Ethiopia, populations from Mozambique and the South Eastern coast of Africa, as well as the South Western coast (Angola) and many many more. The inclusion of these samples will have an impact on these results.
  3. More dense SNPs (~200k) may also give slightly different results, although Sikora (2010) notes the following: “We can conclude that the common set of 2841 SNPs genotyped is an appropriate tool to study population structure in African populations; in general, world-wide patterns are evident and robust when using a minimum of 1000 SNPs.”
  4. Newer and more computer intensive methods for bridging the gap between model based and distance based Autosomal analysis have recently been published, it would be interesting to carry out an analysis of this dataset with these newer methods.