need help with a term paper

profilelauralopezn95
UnderstandinginfluencesofcultureandhistoryonmtDNAvariationandpopulationstructureinthreepopulationsfromAssamnortheastindia.pdf

O R I G I N A L R E S E A R C H A R T I C L E

Understanding influences of culture and history on mtDNA variation and population structure in three populations from Assam, Northeast India

Peter H. Rej1,2 | Ranjan Deka3 | Heather L. Norton4

1Department of Anthropology, University of Florida, Gainesville, Florida 32611 2Genetics Institute, University of Florida, Gainesville, Florida 32610 3Department of Environmental Health, University of Cincinnati Medical Center, Cincinnati, Ohio 45267 4Department of Anthropology, University of Cincinnati, Ohio 45221

Correspondence Peter H. Rej, Department of Anthropology, University of Florida, Turlington Hall, Room 1112, PO 117305, Gainesville, FL 32611- 7305. Email: [email protected]

Funding information Grant sponsorship: Charles Phelps Taft Research Center at the University of Cincinnati - Graduate Student Enrichment Award

Abstract

Objectives: Positioned at the nexus of India, China, and Southeast Asia, Northeast India is presumed to have served as a channel for land-based human migration since the Upper Pleistocene. Assam is the largest state in the Northeast. We characterized the genetic background of three populations and examined the ways in which their population histories and cultural practices have influenced levels of intrasample and intersample variation.

Methods: We examined sequence data from the mtDNA hypervariable control region and selected diagnostic mutations from the coding region in 128 individuals from three ethnic groups currently living in Assam: two Scheduled tribes (Sonowal Kachari and Rabha), and the non-Scheduled Tai Ahom.

Results: The populations of Assam sampled here express mtDNA lineages indicative of South Asian, Southeast Asian, and East Asian ancestry. We discovered two com- pletely novel haplogroups in Assam that accounted for 6.2% of the lineages in our sample. We also identified a new subhaplogroup of M9a that is prevalent in the Sonowal Kachari of Assam (19.1%), but not present in neighboring Arunachal Pra- desh, indicating substantial regional population structuring. Employing a large comparative dataset into a series of multidimensional scaling (MDS) analyses, we saw the Rabha cluster with populations sampled from Yunnan Province, indicating that the historical matrilineality of the Rabha has maintained lineages from Southern China.

Conclusion: Assam has undergone multiple colonization events in the time since the initial peopling event, with populations from Southern China and Southeast Asia hav- ing the greatest influence on maternal lineages in the region.

K E Y W O R D S

genetic structure, descent systems, migration, mitochondrial DNA

1 | INTRODUCTION

The populations of South Asia are characterized by extensive genetic, linguistic, and cultural diversity, resulting from the long history of human presence in the region. India is believed to be one of a limited number of points of dispersal

for anatomically modern humans (AMH) leaving Africa, pre- sumably arriving via a coastal route (Forster & Matsumura, 2005; Macaulay et al., 2005; Thangaraj et al., 2005). Studies of mitochondrial DNA (mtDNA) diversity in South Asia suggest that the major mtDNA haplogroups M, N, and R diversified in India beginning 65–55 thousand years ago (ka)

Am J Hum Biol. 2017;29:e22955. https://doi.org/10.1002/ajhb.22955

wileyonlinelibrary.com/journal/ajhb VC 2017 Wiley Periodicals, Inc. | 1 of 18

DOI: 10.1002/ajhb.22955

American Journal of Human Biology

(Kivisild et al., 1999, 2003; Macaulay et al., 2005; Thangaraj et al., 2006), corroborating archaeological evidence of the earliest presence of AMH in the area (Mellars, 2006; Sali, 1989). In the time since initial human colonization, these haplogroups underwent a period of rapid divergence, and today constitute the primary lineages observed in the region (Maji, Krithika, & Vasulu, 2008, 2009). MtDNA diversity has been subsequently molded by the effects of ongoing, pre- historic and historic migrations into and out of the subconti- nent (Chakrabarti, Kumar, Singh, & Dimitrova, 2012; Majumder, 2008).

The Himalayas and the Ganges River delta both provide substantial barriers to human population migration via land between South and East Asia (Dennell, 2009; Gayden et al., 2013). In between these dominant land features lies the Brah- maputra River Valley, which runs through the heart of Assam, the largest state in Northeast India. Genetic and lithic evidence support the idea that the region has experienced multiple major population movements since the middle Pleistocene, culminat- ing in the later half of the 13th century AD (Cann, 2001; Cavalli-Sforza, Menozzi, & Piazza, 1994; Chandrasekar et al., 2009; Gait, 1906; Hazarika, 2011; Nei & Roychoudhury, 1993; Reddy et al., 2007; Sharma, 2002). Due to Assam’s prominent position at a major geographic crossroads, it has undergone a number of cultural and demic transitions, and today exhibits a wealth of cultural and linguistic diversity. We are interested in understanding how demographic processes helped weave the cultural fabric of Assam, and how together they have shaped regional genetic structure.

As in the rest of the subcontinent, the autochthonous South Asians established a large portion of the mtDNA gene pool in Northeast India (Chandrasekar et al., 2009). How- ever, unlike the rest of India, in the millennia following the initial peopling event, the populations of the Northeast have been heavily influenced by multiple waves of invasion and migration from Southeast Asia and China, in addition to Western Eurasia (Chandrasekar et al., 2009; Chaubey, Met- spalu, Kivisild, & Villems, 2007; Peng et al., 2011; Reddy et al., 2007). Mitochondrial and Y-DNA evidence indicates that ancestors of modern Austro-Asiatic (A-A) speakers were the first to enter India from East Asia �15–20 ka (Reddy et al., 2007; Zhang et al., 2015). Today, however, the major- ity of East Asian mtDNA markers observed in the Northeast are found in speakers of a language from the Tibeto-Burman (TB) branch of the Sino-Tibetan (ST) language family (Chandrasekar et al., 2009; Reddy et al., 2007). It has been posited that TB languages originated in China, then spread to Southeast Asia via the Yellow River, and then subsequently into Northeast India and the Himalayas after the Last Glacial Maximum (LGM; van Driem, 2001). Although there are some alternative lines of Y-STR evidence indicating that TB speakers may have arrived in Northeast India and Bangla-

desh simultaneously or even before A-A’s (Gazi et al., 2013), most DNA studies indicate that this migratory event occurred during the early Holocene (Cordaux, Weiss, Saha, & Stoneking, 2004; Metspalu et al., 2004; Peng et al., 2011; Reddy et al., 2007).

In this study, we characterize mtDNA variation in 128 individuals from three non-caste population groups sampled from the Indian state of Assam: the Ahom, the Sonowal Kachari, and the Rabha. The Rabha and the Sonowal Kachari are two of �400 Scheduled tribal groups recognized by the Indian government (Majumder, 2008), 23 of which can be found in Assam (Long, 2012). “Sched- uled” is a legal classification that the national government reserves for lower castes and tribal populations in an effort to elevate the socio-economic status of these groups that are marginalized by the caste system (Chaubey et al., 2007). Historically, the Rabha and Kachari both speak dia- lects of the Bodo branch of the TB family (Burling, 2003), although, like many tribal populations, most individuals have lost their native language and now exclusively speak Assamese, the predominant Indo-European (IE) language in the region (Ripunjoy, 2013; van Driem, 2011). The Sonowals are one of many branches of the Bodo-Kachari people found in Northeast India, and are the third largest tribal population in Assam (�350,000 individuals; Kalita & Deb, 2004). Historically, they were a strictly endoga- mous population found in the plains regions of the Tinsu- kia, Dibrugarh, Sivasagar, Lakhimpur, and Dhemaji districts of Upper Assam (Figure 1; Das & Sengupta, 2003). In contrast, the Rabha are historically exogamous, inhabiting Lower Assam, along the southwestern border with Meghalaya, where they neighbor a number of Khasi speaking (A-A) tribes that share similar marriage practices (Burling, 2007; Deb, 2006). They are traditionally a matri- lineal/matrilocal society; however, this practice is no lon- ger commonplace (Raha, 1989).

The history of Assam also attests to more recent influxes of peoples from the East. One of the most recent major ethnic groups to settle along the Brahmaputra is the Tai Ahom. This Daic speaking population historically originating from the lands along the Irrawaddy River, which forms the border between Myanmar and modern-day Yunnan Province, crossed the Patkai hills and settled in the Brahmaputra Valley during the early 13th century (Fernquest, 2006; Gogoi, 1968; Gait, 1906). There, they established an empire that spanned roughly seven centuries, eventually falling at the hands of British imperialists (Gait, 1906). Historically, the initial Ahom conquest of the region was mediated by a force pri- marily consisting of young landless males who intermarried with local populations, and promptly converted to Hinduism (Gait, 1906).

The Ahom are one of the largest (formerly) Tai-Kadai/ Daic speaking populations in India (Diller, Edmondson, &

2 of 18 | American Journal of Human Biology REJ ET AL.

Luo, 2004). However their exact numbers are unknown, as they are classified by the Indian government in a general “ethnic Assamese” group, which includes all of the caste and non-Scheduled tribal peoples who speak Assamese as their primary language (Saikia, 2004). There is evidence of exten- sive cultural and linguistic assimilation with the Hindu caste populations of the region, although there have been recent movements towards cultural revival (Morey, 2015). Today there are approximately 1 million Ahoms in Assam, many of whom inhabit urban centers such as Dibrugarh and Sivasa- gar, the former capital of the Ahom Kingdom, in Upper Assam (Figure 1; Morey, 2015; Saikia, 2004).

In our study, we use mtDNA to highlight the ways in which the different population histories and cultural practices have influenced variation within and between the groups sampled. Specifically, we investigate the role that (pre) his- torical migrations and marriage practices have played in shaping mtDNA variation in all three samples. Furthermore, as Assam is the largest Northeast Indian state, it has undoubtedly played a role in migratory events between East and South Asia; we explore this position from an mtDNA perspective. We seek to place Northeast Indian mtDNA diversity in the broader context of Asian mtDNA diversity by exploring our three Assamese population samples in rela- tion to a large comparative dataset, which includes popula-

tion samples from China, India, and Southeast Asia, taking into account historical, cultural, geographic, and linguistic relationships.

2 | MATERIALS AND METHODS

2.1 | Populations studied

The sample consisted of 128 individuals from the Northeast Indian state of Assam: 35 Tai Ahom, 42 Sonowal Kachari, and 51 Rabha. Blood samples were initially collected as part of a 1984 collaborative venture between research universities in India, Germany, and the United States (Das, Das, Das, Walter, & Danker-Hopfe, 1985; Das & Deka, 1975; Das, Deka, & Flatz, 1987; Deka et al., 1988). These studies looked at allelic diversity in classical genetic markers (ABOs, MNS, Rhesus, Duffy, Diego, Hp Gc, Gm, Km, Tf, aP, EsD, AK, ADA, LDS, and HB-variants). The Ahom and Sonowal samples were collected from the Dibrugarh district in the eastern region of Upper Assam (Figure 1; Deka et al., 1988). The Rabha samples are from the Goalpara district of Lower Assam, which is in the Garo Hills along the border with Meghalaya (Figure 1; Deka, 1984). Subjects were of variable age and sex. Use of these de-identified samples for

FIGURE 1 Map of Assam. Rabha were sampled in Goalpara district, Ahom and Sonowal Kachari were sampled in Dibrugarh district, Assam

REJ ET AL. American Journal of Human Biology | 3 of 18

studies on genetic variation in human populations was approved by UC IRB (IRB# 03 09 16 013).

2.2 | DNA analysis

DNA was extracted from whole blood using phenol chloro- form extraction (Sambrook, Fritsch, & Maniatis, 1989). Amplification of the hypervariable region (HVR) was carried out using the polymerase chain reaction (PCR) and HVR pri- mers F15975 and R635 (Saiki et al., 1988). PCR products were purified for sequencing using the GeneJet Purification Kit (LifeTechnologies). Samples were characterized for the diagnostic Indian, East Asian, Southeast Asian, and Euro- pean RFLP-based haplogroups (M, N, D, F, U, and HV) using associated restriction enzymes. Primers, PCR, and restriction digestion conditions available upon request. The HVR was sequenced and individual sequence contigs were aligned to the revised Cambridge Reference Sequence (rCRS) (Andrews et al., 1999) using the Geneious software package (Geneious 6.0, 2012). All observed variable sites and their positions can be found in Supporting Information Table S1. Mitochondrial haplogroup assignment was con- firmed by genotyping diagnostic mutations in the coding region of the mtDNA molecule. Haplogroup frequencies were calculated using methods previously cited (Forster, Harding, Torroni, & Bandelt, 1996; Sahoo & Kashyap, 2006; Saillard, Forster, Lynnerup, Bandelt, & Norby, 2000; Yao, Kong, Bandelt, Kivisild, & Zhang, 2002) and were cat- alogued for the Ahom, Rabha, and Sonowal Kachari sam- ples. For researchers who prefer the Reconstructed Sapiens Reference Sequence (RSRS; Behar et al., 2012), all sample sequences were also aligned with the RSRS. This is summar- ized in Supporting Information Table S2. All subsequent analyses were completed using alignment with the rCRS.

2.3 | Statistical analyses

Haplotype frequencies, pairwise FST values, AMOVAs, and all intrapopulation summary statistics were calculated using MEGA version 6 (Tamura, Stecher, Peterson, Filipski, & Kumar, 2013) or Arlequin 3.5 (Excoffier & Lischer, 2010). To visualize haplotype diversity and haplogroup distribution, median joining networks were generated using PopART v1.7 (http://popart.otago.ac.nz). Two measures of nucleotide diversity were calculated. The first of these, p (Nei & Li, 1979), is based on the average number of nucleotide differ- ences between two sequences randomly drawn from a sam- ple. The second, h, is based on the proportion of segregating sites in the sample (Watterson, 1975). Two statistics based on the site frequency spectrum that are sensitive to changes in population size, Tajima’s D, and Fu’s Fs (Fu, 1997; Tajima, 1989), were also calculated. The significance levels of Tajima’s D and Fu’s Fs were assessed by comparing the

observed values with a distribution of 1,000 randomized per- mutations (Scozzari et al., 1999). Pairwise mismatch distri- bution plots were generated to help infer population growth.

Comparative HVR-1 data from tribal populations throughout India, as well as neighboring populations from China and Southeast Asia, (see Supporting Information Table S3) were incorporated into a series of AMOVAs, where we evaluated how geographic and linguistic affiliation affected mtDNA variation within and between Northeast India and each region. Comparative populations were chosen based on geographic proximity to Assam, shared language family, posited historical/prehistorical association with our populations, and the availability of published sequences. All HVR-1 sequences in our comparative data set were obtained from GenBank, and converted to fasta format using a custom Python script (available on request). Pairwise FST differences over the HVR were calculated between individuals in the sample and then incorporated into an multidimensional scal- ing (MDS) plot to investigate any patterning/grouping within the sample. FST differences were then calculated between all sample pairs in our dataset and the comparative HVR-1 data- set. These FST values were used to construct a series of MDS plots in an effort to place Assam in the broader context of Asian mtDNA haplotype diversity. All MDS plots were con- structed in the R computational environment (R version 2.14.2).

3 | RESULTS

3.1 | Haplogroup analysis

Macrohaplogroup M and its subhaplogroups were the most commonly observed haplogroups found in our three samples from Assam (Table 1; Figure 2), accounting for 69.5% of all individuals. Subhaplogroups of M33b and M9a1b1 make up �36% of all the haplogroups found in the Sonowal Kachari, at 14.3 and 21.4%, respectively. M33b1 was also observed in a single Ahom individual, but was not present in any of the Rabha individuals. Subhaplogroup M9a1b1 was found at low frequencies among both the Ahom and Rabha at 4.0 and 5.7%, respectively. Eight of the Kachari and one Rabha pos- sessing M9a1b1 were also typed for a C to A transversion at position 5178. A diagnostic SNP for haplogroup D, 5178 has not been previously identified in M9.

Haplogroup R and its subhaplogroups were the next most frequent among the members of our samples, encom- passing 20.3% of the individuals examined. Haplogroup F, a descendant of R that is most frequently found in Southeast Asians (Hill et al., 2007; Schurr & Wallace, 2002; Tolk et al., 2001), was observed in all three samples (Table 1). Haplogroup U, commonly associated with Western Eurasian ancestry (Kivisild et al., 1999), was not observed in any of

4 of 18 | American Journal of Human Biology REJ ET AL.

our samples. Excluding members of R, haplogroup N was the least common macrohaplogroup in our sample (10.2%). The only identifiable subhaplogroup was A, which was observed in all three population samples.

3.2 | Novel haplogroups

We identified three novel haplogroups in this study (Table 2). Incorporating the aforementioned transversion at 5178 into our suite of diagnostic SNPs, we classified the new M9a subhaplogroup as M9a1b1d, which dates back to 4.0 6 2.2 ka. The other two novel haplogroups possess HVR diagnos- tic motifs that have not been previously observed (Table 2). The novel M subhaplogroup observed in Assam was Rabha-

specific, found in roughly 10% of all Rabha individuals sampled. This novel haplogroup dates to 4.9 6 3.9 ka, and will be assigned M82 based on the next unoccupied number on the phylogenetic mtDNA haplogroup tree (van Oven & Kayser, 2009). We also found a novel R subhaplogroup that dated to 6.8 6 5.0 ka, which we labeled as R33. Of the popu- lations sampled, R33 was observed only among the Rabha, where it was found at a frequency of 5.9% (Table 1). Dating estimates of novel haplogroups based on rho statistics (For- ster et al., 1996; Soares et al., 2009) are included in Support- ing Information Table S4. More refined age estimates for the novel haplogroups could be ascertained by full mtDNA genome sequencing.

TABLE 1 Subhaplogroup frequencies (with counts in parentheses) in the Ahom, Sonowal Kachari, Rabha, and in the total Assamese sample

Subhaplogroup Ahom Sonowal Kachari Rabha Assam (total)

M 0.371 (13) 0.286 (12) 0.314 (16) 0.320 (41)

M3 0.057 (2) 0.000 (0) 0.020 (1) 0.023 (3)

M6a 0.029 (1) 0.000 (0) 0.000 (0) 0.008 (1)

M7b1a 0.000 (0) 0.000 (0) 0.039 (2) 0.016 (2)

M8 0.029 (1) 0.000 (0) 0.000 (0) 0.008 (1)

C 0.000 (0) 0.024 (1) 0.000 (0) 0.008 (1)

M9a 0.000 (0) 0.000 (0) 0.020 (1) 0.008 (1)

M9a1b 0.000 (0) 0.000 (0) 0.020 (1) 0.008 (1)

M9a1b1 0.057 (2) 0.024(1) 0.020 (1) 0.031 (4)

M9a1b1d 0.000 (0) 0.191 (8) 0.020 (1) 0.070 (9)

M11 0.000 (0) 0.024 (1) 0.000 (0) 0.008 (1)

M20 0.000 (0) 0.000 (0) 0.020 (1) 0.008 (1)

M30 0.057 (2) 0.000 (0) 0.000 (0) 0.016 (2)

M33b1 0.086 (3) 0.143 (6) 0.020 (1) 0.078 (10)

M48 0.000 (0) 0.024 (1) 0.000 (0) 0.008 (1)

M-novel 0.000 (0) 0.000 (0) 0.098 (5) 0.039 (5)

D 0.057 (2) 0.048 (2) 0.020 (1) 0.039 (5)

N 0.057 (2) 0.048 (2) 0.059 (3) 0.055 (7)

A 0.057 (2) 0.071 (3) 0.020 (1) 0.047 (6)

R 0.057 (2) 0.071 (3) 0.098 (5) 0.078 (10)

F 0.086 (3) 0.048 (2) 0.118 (6) 0.086 (11)

B4 0.000 (0) 0.000 (0) 0.039 (2) 0.016 (2)

R-novel 0.000 (0) 0.000 (0) 0.059 (3) 0.023 (3)

REJ ET AL. American Journal of Human Biology | 5 of 18

3.3 | Mitochondrial haplotype variation and demographic history

A total of 100 (78.1%) unique haplotypes were identified among the 128 individuals sequenced successfully. The Sonowal Kachari exhibited the most haplotype sharing both between and within populations, as indicated by their low haplotype diversity statistic (0.952; Table 3). This lower diversity can be visualized in the number of shared haplo- types in our network (Figure 3), as well as in the multimodal pairwise mismatch distribution (Figure 4). The two most common haplotypes (Haplotype H39 and Haplotype H6) observed among members of all three samples occur among

31% of Sonowal individuals (19 and 12%, respectively) and correspond with haplogroups M9a1b1d and M33b1. The next most common haplotypes (Haplotype H68 and Haplo- type H64) are associated with the novel M and R hap- logroups discovered in this study (Supporting Information Table S5; Table 2). The Rabha demonstrate the next lowest haplotype diversity in our sample (0.984), while the Ahom demonstrated the highest amount of diversity (0.998; Table 3), with all but three individuals possessing unique haplo- types. The heterogeneity of the Ahom can be observed in our haplotype network plot (Figure 3). None of the haplo- types observed in this study were shared by more than two population samples.

Table 3 summarizes diversity measures and selected neu- trality statistics for the populations considered here. Nucleo- tide diversity levels (p) range from 8.9 3 1023 to 1.1 3 1022, with highest levels among the Rabha, and lowest lev- els among the Sonowal Kachari. Both Tajima’s D (TD) and Fu’s Fs values were negative for all populations, although TD was not significant in any population. In contrast Fu’s Fs was significantly negative in all three groups (P < .01) indicating a history of population expansion. The raggedness index (r) ranged from 0.004 in the Sonowal Kachari to 0.015 in the Rabha, also supporting a history of expansion. Popula- tion growth in the Rabha, and in particular the Ahom is fur- ther supported by normally distributed pairwise mismatch plots (Figure 4). However, in the case of the Sonowal Kachari, the multimodal distribution plot and the significant deviance (P 5 .01) in goodness of fit (r-statistic5 0.046) reflect the genetic isolation of that group.

3.4 | Population structure

When only the three sampled populations are considered, our analysis of molecular variance (AMOVA) suggests that the majority of variation can be attributed to intra-sample variation, because variation between the sampled populations only accounts for 4.17% of the observed variation (Table 4). The average F-statistic over the entire control region was 0.042. Indicative of a significant degree of genetic variation between populations (P < .001). Our average F-statistic is comparable to other studies measuring regional HVR varia- tion in Northeast India (i.e., Reddy et al., 2007). Pairwise

FIGURE 2 Median joining haplogroup network generated using PopART v1.7. Number of diagnostic SNPs used to iden- tify haplogroups is indicated by hatchmarks on branches. Dashed line indicates separation of macrohaplogroup R. Blue indicates Rabha, red indicates Ahom, and black indicates Sonowal Kachari

TABLE 2 Diagnostic mutational motifs, compared to the CRS, of novel haplogroups identified in this study

Haplogroup Diagnostic HVR substitutions Populations

M82 A16316G-A73G-C150T-A263G-T293C Rabha

R33 A16051G-C16168T-T16311C-A73G-T146C-A263G Rabha

M9a1b1d C5178A-C16234T-T152C-C150T-A16158G Sonowal Kachari; Rabha

6 of 18 | American Journal of Human Biology REJ ET AL.

FST comparisons also showed significant (P < 0.05) differen- ces between all pairs of populations.

A UST comparison matrix between all individuals in our sample was incorporated into an MDS plot (Figure 5). There are no identifiable groupings of individuals based on population sampled; however, we observe a distinct clus- tering of individuals that had been assigned to autochtho- nous lineages. Included among them are a handful of

individuals who were assigned to the blanket group M. We presume that, with whole mitochondrial genomes, these individuals would also be assigned to one of the estab- lished autochthonous lineages. The predominant East Asian mtDNA lineages observed in our sample (M9a, F, A) all fall outside of the large cluster, but form their own distinct smaller clusters (Figure 5).

3.5 | Assam in the context of south and east Asia

Additional AMOVAs that incorporate published data were carried out to compare the effects of geography and language family on HVR-1 variation across South and East Asia (Table 5). We grouped our three Assamese samples with other Northeastern tribal population samples and compared them to samples from other parts of the subcontinent, China, and Southeast Asia (Table 5). A large proportion of the observed variation (84-94%) is accounted for by differences within the samples, regardless of whether samples are grouped by region or language, while only 2–5% of all the variation in our analyses could be attributed to between group differences. In all but one of our AMOVAs, our critical F-statistics were significant (P < .05). When grouped by region, the difference among groups (FCT) was not signif- icant when comparing Northeast Indian populations to those from Southeast Asia (Table 5).

MDS analysis using the same three-region strategy employed with AMOVA was performed to further investi- gate relationships between our samples from Assam and our comparative dataset (Figure 6A–D). In the MDS plot featur- ing tribal samples from across India (Figure 6A), we find that not only are the Northeast Indian populations clustering together, but other TB speaking samples from neighboring Nepal (Lepcha and Lachungpa) are included in this cluster. There is some separation between Kachari sampled in Aru- nachal Pradesh and our Assam sample, further reflecting the difference in haplogroup assignment between the two sampled Sonowal communities.

The second MDS plot compares tribal individuals sampled from throughout China to Northeast Indians (Figure

TABLE 3 Intrapopulation summary statistics

Population n k Haplotype diversity h

Nucleotide diversity p

Tajima’s D (P-value)

Fu’s Fs (P-value)

Raggedness (P-value)

Ahom 35 34 0.9983 1/2 0.0074

0.0097 1/2 0.0050

21.722 (0.016) 224.31* (<0.01) 0.0088 (0.34)

Sonowal Kachari

42 28 0.9524 1/2 0.0218

0.0089 1/2 0.0046

22.581 (0.035) 28.51* (0.01) 0.0044 (0.78)

Rabha 51 42 0.9843 1/2 0.0102

0.0106 1/2 0.0054

21.747 (0.015) 223.52* (<0.01) 0.0146 (0.71)

P-values: * 5 P < 0.01. Note: n, sample size; k, number of haplotypes.

FIGURE 3 Median joining haplotype network of mtDNA HVR. Mutations are shown as hatch marks on branches. Dashed line indicates separation of macrohaplogroup R. Major haplogroups are circled and labeled.

REJ ET AL. American Journal of Human Biology | 7 of 18

6B). At the origin of the plot, we see a grouping of Daic speaking samples from across Southern China. A handful of TB speaking groups from Yunnan Province are found inter- spersed in this cluster, but, for the most part, group with other TB speaking samples from Tibet and with the Rabha. TB samples from the rest of Northeast India, including the Sonowals, are dispersed roughly around x 5 20.1. The Ahom fall between the East and South Asian groups, though are closer to groups from Tibet than they are to any North- east Indian population.

When compared with populations sampled in Southeast Asia (Figure 6C), we observe a pattern similar to the previ- ous MDS plots. Northeast Indian populations, including the Sonowal Kachari, form a distinct cluster, seen in the upper right hand corner of the plot. However, the Ahom and Rabha do not fall into this cluster and appear much closer to the Southeast Asian populations included in this analysis.

When our three samples are compared to the entire bat- tery of comparative samples (Figure 6D), we see a large cen- tral cluster primarily consisting of East Asian TB speaking populations. Tight within this cluster are the Rabha, again located in a position that is distinct from the other Northeast/ Assamese populations. We also see separation of South Asian, East Asian TB, and East Asian Daic speaking groups

along the x-axis. The Chinese and Tibetan TB tribes fall between the Northeast Indian TB speakers and the Daic speakers. We also observe a split between Northeast and non-Northeast Indians along the Y-axis. The Sonowals are found among other TB speaking tribes, although closer to those of Arunachal Pradesh. They have virtually the same x- value as the Sonowal Kachari sampled from that province. The Ahom are also found among the TB populations, and appear to be far removed from other Daic speaking groups from Yunnan, their historical homeland and mother tongue.

4 | DISCUSSION

4.1 | mtDNA landscape of three of Assam’s ethnic groups

Previous studies assessing variation among classic alleles in the Sonowal Kachari and Rabha suggested that allele fre- quencies in both groups were similar to those previously reported in East Asian population samples (Das et al., 1985, 1987; Deka, 1984; Deka et al., 1988). The Ahom, however, often appeared intermediate to East Asians and Assamese Hindu Caste samples in their allele frequencies at these clas- sical loci (Das et al., 1987). We observe more complex

TABLE 4 AMOVA results of three Assam populations based on complete HVR sequence

Source of variation d.f. Sum of squares

Variance components

Percentage of variation

Among populations 2 18.752 0.143 4.17

Within populations 126 414.279 3.288 95.83

Total 128 433.031 3.431

Fixation Index (FST) 0.042

Significance (<0.001)

FIGURE 4 Mismatch distribution plots of pairwise nucleotide differences between mtDNA HVR sequences. A: Sonowal Kachari; B: Ahom; C: Rabha

8 of 18 | American Journal of Human Biology REJ ET AL.

patterns in the mtDNA sequence data reported here, likely reflecting cultural practices, regional variation, migration times, and varying levels of population admixture. We also identify three haplogroups that appear to be unique to Assam. These include novel M and R subhaplogroups that are most common in the Rabha, and one subhaplogroup of M9a1b1 that is found at an elevated frequency (19.2%) in the Sonowal Kachari. These new data help to provide a clearer picture of how some of Assam’s extant diversity was shaped.

4.2 | The Ahom

Of the populations sampled not only in this study, but also over the past 15 years of mtDNA studies in India, the Ahom stand out. They are unique among noncaste populations in India for a multitude of reasons: they are a historically Daic speaking population, they are relatively recent arrivals to the subcontinent, their history in this region is well-documented, and they have culturally assimilated with local Hindu Assamese populations.

The nature of the Ahom migration is reflected in the mtDNA diversity observed in our Ahom sample. When including Daic-speaking populations from Yunnan province, the historically documented language, and homeland of the Ahom, into inter-population MDS analyses, the Ahom always appear intermediate to the Northeast Indian and Daic

speaking Chinese samples, interspersed among TB speaking populations from Tibet and South China (Figure 6). We also do not observe a high proportion of mtDNA lineages com- monly associated with Tai-Kadai/Daic-speaking populations in this Ahom sample, as one might expect if the original Ahom migration to Assam included a large proportion of females. Haplogroup F is the most frequently observed R subhaplogroup in speakers of Tai-Kadai languages (Li et al., 2007), as well as among Southeast Asians (Bodner et al., 2011; Summerer et al., 2014). However, despite being believed to be the largest Daic population in India, hap- logroup F is not found at an elevated frequency in the Ahom sample. Like other Southeast Asian lineages including M20, M7b1a, and B4, haplogroup F is found at its highest fre- quency in Assam in the Rabha, and all are at frequencies similar to the matrilocal tribes of Meghalaya (Reddy et al., 2007). Thus, the appearance of F could be explained if it was already present in the Northeast at the time of the Ahoms arrival in the 13th century AD. If so, its existence in our Ahom sample can be explained in part via admixture between the Ahom and the populations already living in the region.

The distribution of mtDNA haplogroups present in the Ahom sampled also implies that the 13th century Ahom set- tlers admixed with local populations. In addition to Southeast Asian lineages, mtDNA haplogroups commonly associated

FIGURE 5 Plot of individual samples on the first two dimensions from the MDS analysis of a UST matrix generated by Arlequin 3.5

REJ ET AL. American Journal of Human Biology | 9 of 18

TABLE 5 AMOVA based on HVR-1 haplotypes

Source of Variation d.f Sum of squares

Variance compon ents

Percentage of variation

a) Indian Tribes

Regional Among groups 5 197.988 0.159 4.86

Among populations within groups 22 287.979 0.365 11.17

Within populations 840 2306.312 2.746 83.96

Total 867 2792.279 3.270

Fixation Indices FSC FST FCT

Significance 0.117* 0.160* 0.049*

<0.001 <0.001 <0.001

Language Family Among groups 5 179.185 0.144 4.36

Among populations within groups 21 289.382 0.397 12.04

Within populations 788 2173.218 2.758 83.61

Total 814 2641.784 3.300

Fixation Indices FSC FST FCT

Significance 0.125* 0.164* 0.044*

<0.001 <0.001 <0.012

b) China

Regional Among groups 6 191.770 0.080 2.28

Among populations within groups 43 362.194 0.119 3.39

Within populations 2171 7157.758 3.297 94.33

Total 2220 7711.749 3.495

Fixation Indices FSC FST FCT

Significance 0.035* 0.057* 0.023*

<0.001 <0.001 <0.001

Language Family Among groups 3 154.741 0.089 2.53

Among populations within groups 45 394.059 0.126 3.59

Within populations 2157 7101.719 3.292 93.89

Total 2305 7650.518 3.507

Fixation Indices FSC FST FCT

Significance 0.037* 0.061* 0.025*

<0.001 <0.001 <0.001

c) Southeast Asia

Regional Among groups 3 132.947 0.120 3.15

Among populations within groups 10 133.259 0.352 9.26

(continues)

10 of 18 | American Journal of Human Biology REJ ET AL.

with TB speakers (M9a1b), the initial peopling of India (M30, M3, M33b), and the West Eurasian migration (M6a) are all present in our Ahom sample. Only the Ahom sample included haplogroup M6a (23.1 1/2 7.7 ka; Thangaraj et al., 2006), albeit a singleton, suggesting that M6a is rare in Assam. M6a is found at low frequencies in Southeast India (<3%) and Oman (<1.5%; see Metspalu et al., 2004 for spa- tial distribution of M6a).

Other lineages commonly associated with Western Eura- sian populations, HV, W, and most surprisingly U, were completely absent in our sample. Haplogroup U is a highly diverse lineage that is frequently observed across India, aver- aging �13% (Roychoudhury et al., 2000) and up to >77% in certain Western Indian tribal populations (Kivisild et al., 1999; Roychoudhury et al., 2001). Haplogroup U is also indicative of the initial peopling of West India (Kivisild et al., 1999). Clinally distributed from west to east, descend- ants of subhaplogroup U2 have been found in neighboring Meghalaya at a frequency of 8% (Reddy et al., 2007). While we cannot say that U subhaplogroups are not present in Assam, we suspect that they are sufficiently uncommon to not be captured in our three Assamese samples.

As predicted, the Ahom are also the most genetically diverse of the three populations sampled here, and demon- strate recent population expansion (Table 3; Figure 4). Although this substantial diversity could partially be explained by non-Ahoms self-reporting as Ahom, we believe that it is more likely that this variation can be attributed to

the fact that this recently arriving (30–40 generations ago) and primarily male founding population experienced exten- sive admixture with local populations followed by subse- quent growth. This is supported by a history of assimilation with resident populations (Grimes & Diller, 2003; Saikia, 2004). Together, these factors may explain the lack of shared mitochondrial haplotypes within the Ahom and the higher levels of observed mtDNA diversity relative to the two TB speaking tribal groups studied here. Given the historical male bias in the migration of the Ahom to Assam (Gait, 1906), we would predict the opposite pattern in the Y- chromosome, that is, higher levels of within-sample haplo- type sharing and lower levels of diversity.

4.3 | The Sonowal Kachari

As expected from their history of endogamy (Das & Sen- gupta, 2003), the Sonowal Kachari are the most homogenous and are significantly different from all of the other popula- tions sampled in terms of both intrapopulation diversity and interpopulation divergence. They exhibit the lowest levels of intrasample genetic diversity among our three Assam sam- ples and display the highest levels of FST in pairwise popula- tion comparisons (FST values ranged from 6.62 3 10

23 – 4.14 3 1022). This pattern is also evident in the haplotypic data (Table 3; Figures 2 and 4), where only 66.7% of the haplotypes observed in the Sonowal Kachari were singletons

TABLE 5 (continued)

Source of Variation d.f Sum of squares

Variance compon ents

Percentage of variation

Within populations 583 1940.891 3.329 87.59

Total 596 2207.097 3.801

Fixation Indices FSC FST FCT

Significance 0.096* 0.124* 0.032

<0.001 <0.001 0.087

Language Family Among groups 4 134.810 0.010 2.63

Among populations within groups 9 131.396 0.360 9.50

Within populations 583 1940.891 3.329 87.87

Total 596 2207.097 3.789

Fixation Indices FSC FST FCT

Significance 0.098* 0.121* 0.026*

<0.001 <0.001 0.001

Resulting population structure is from assignment to geographical region or language family, and was measured between North- east Indians and (a) tribal populations from across India, as well as (b) Chinese, and (c) Southeast Asian populations. *A 0.05 significance level was assigned for all F-statistics. Reference to comparative populations are provided in Table S1.

REJ ET AL. American Journal of Human Biology | 11 of 18

as opposed to 97.1 and 82.4% among the Ahom and Rabha, respectively.

The prevalence of novel subhaplogroup M9a1b1d among the Sonowal Kachari is one of the most striking features of our haplogroup dataset (Table 1). Parental haplogroup M9a is a common East Asian lineage distributed at frequencies roughly around 9–18% in populations ranging from Southern China into Northeast India, Nepal, and Tibet (Chandrasekar et al., 2009; Metspalu et al., 2004; Peng et al., 2011; Reddy et al., 2007; Wang et al., 2012; Zhao et al., 2009). Peng et al. (2011) provide detailed spatial frequency distributions of haplogroup M9a’b and its sub-haplogroups. M9a is esti- mated to have emerged during the LGM (�18–23 ka), while subhaplogroup M9a1b1, along with M9a1a* and M9a1a2, are post-LGM lineages that arose during the Late Pleisto-

cene/early Holocene (�12–18 ka) (Peng et al., 2011). M9a1b1, M9a1a8, and M9a1a2 have been associated with Mesolithic cultures occupying Southern China (Bailandong Stage III in Guangxi) and Southeast Asia (Hoabinhian) at these times, and basal lineages still remain at high frequen- cies in these regions (Peng et al., 2011). A total M9a’b fre- quency of 11.7% in Assam is similar to those observed across Northeast India (8.7–11.7%; Chandrasekar et al., 2009; Reddy et al., 2007) and in Nepal (11.6%; Fornarino et al., 2009) and Tibet (19.2%; Qin et al., 2010; Zhao et al., 2009). This supports Peng et al. (2011) hypothesis that the current phylogeographic distribution of M9a1 subha- plogroups is indicative of a land-based migration from South China/Southeast Asia. Previous data, both genetic (Reddy et al., 2007) and archaeological (Sharma, 2003), suggested

FIGURE 6 Plots on the first two dimensions from the MDS analysis of the FST matrices generated via MEGA 6.0. References to comparative populations are provided in Table 1. A: Plot comparing Northeast Indian tribes (and the Ahom) to other Indian tribal populations; B: plot of Northeast Indians and neighboring populations in China; C: plot including Northeast Indians and popula- tions from Southeast Asia; D: MDS plot incorporating all of the populations from our comparative sample

12 of 18 | American Journal of Human Biology REJ ET AL.

that the Garo Hills were one of the regions colonized during this migration; our data expands the migration route/coloni- zation event to also include the Brahmaputra River Valley.

The M9a1b1d lineage prevalent in the Sonowal Kachari, and novel to Assam, likely originated from a small initial founder population of M9a1b1 that settled in Upper Assam as part of this greater post-LGM migration to the Indian sub- continent. It then evolved in situ by compiling an additional mutation, as detected in our sample, the transversion at mtDNA position 5178. Subsequent endogamy and genetic drift, followed by limited population expansion (as indicated by Tajima’s D, Fu’s Fs, and the raggedness index values reported in Table 3, as well as the pairwise mismatch distri- bution plot in Figure 4) likely resulted in the observed fre- quency in our Sonowal Kachari sample.

Along with M9a1b1d, M33b1 accounts for 35.8% of all the haplogroups identified in our Sonowal Kachari sample. M33b1 is a descendant of an ancient lineage presumed to be autochthonous to Northeast India. Chandrasekar et al. (2009) calculated that M33b emerged in the Northeast �44,000 ka. Apart from singletons in Maharashtra (Chan- drasekar et al., 2009) and Southern China (Kong et al., 2011), and two individuals in Myanmar (Li et al., 2015), M33b and its subhaplogroups are exclusive to Northeast India and Nepal (Chandrasekar et al., 2009; Fornarino et al., 2009; Reddy et al., 2007). M33b is found at high fre- quency (22%) in Pnar Khasi of the Jaintia Hills in Eastern Meghalaya, while only at 3% in the TB speaking Garo (Reddy et al., 2007). Similarly, it is found in the neighbor- ing Rabha at 2%.

An exceedingly rare subhaplogroup, M33b1 has only been reported at frequencies <2.2% (Gayden, 2012; Kong et al., 2011; Li et al., 2015). As with M9a1b1, we postulate that the high frequency of M33b1 in the Sonowal Kachari reflects a history of founder effect and subsequent endogamy. One hypothesis is that M33b1 emerged in Assam, and was introduced into TB speaking populations via admixture. The presence of M33b1singletons further west indicates limited introgression with Northeastern populations.

Interestingly, haplogroup frequencies differ between the Sonowal Kachari sampled here and those sampled in Aruna- chal Pradesh by Chandrasekar et al. (2009). Most notable is the sharp difference in frequencies between M33b in the Sonowal Kachari (35.8%) and the Sonowal from Arunachal Pradesh (2%). While the frequency of haplogroup M9 is sim- ilar in both Sonowal groups (�14%), haplotype M6 is also common in the Arunachal Pradesh sample (�14%), while being absent from our own. These differences are highlighted in our MDS plots (Figure 6), where in almost all cases, other Northeastern populations sampled in Arunachal Pradesh fall closer to Chandrasekar’s Sonowal sample than our Sonowal sample does. These findings go on to further reflect

the highly variable distribution of mtDNA lineages across Northeast India, where even among members of the same historically endogamous ethnic group, we see structuring based on state of residence. Alternatively, the disparate hap- lotype distributions between the two Sonowal Kachari groups could also be explained by independent migrations from East Asia into Northeast India.

Another autochthonous Northeastern haplogroup identified in the Sonowal Kachari is M48. Initially defined by Reddy et al. (2007), M48 is most concentrated in Meghalaya (24% in the Lyngngam). Its presence in Upper Assam indicates that the distribution of maternal lineages has been maintained in the Northeast since the initial colonization of the region. Like M33b1, this lineage is not present in the Rabha despite regional continuity and shared marriage practices.

4.4 | The Rabha

The exogamous practices of the Rabha are illustrated in our diversity measures, as they possess the highest diversity lev- els. However, despite the lack of autochthonous lineage M48, and the low frequency of M33b1, the Rabha in our sample bear some similarities in their mtDNA haplogroup assignments to the A-A tribes of Meghalaya sampled by Reddy et al. (2007). In particular, they share many markers commonly associated with Southeast Asian populations (i.e., F, M7a, B 2 �20%; Schurr & Wallace, 2002). These find- ings corroborate the conclusions of Chaubey et al. (2011) who found up to 1/3 of the maternal lineages in their Khasi sample to be of Southeast Asian origins. Unfortunately, Khasi mtDNA sequences were not freely available to be incorporated into our interpopulation analyses. Considering their close geographic proximity to the Rabha in the Garo Hills, and their shared historical exogamous marriage prac- tices, we would expect the Khasi to group with our Rabha sample.

The position of the Rabha in our MDS plots (Figure 6), clustered with populations from Southern China (Figure 6b, d) is another striking feature of our results. This likely reflects their marriage practices, but also the subsequent maintenance of regional mtDNA lineages following their early Holocene migration into the region.

Two completely novel haplogroups, one nested in M, the other in R, were identified in our Rabha sample. Novel hap- logroup M82 was only found in our Rabha sample, while novel haplogroup R33 was shared with individuals sequenced from Assamese speaking Hindu caste populations (Rej, 2013).

4.5 | mtDNA structure within Assam

Although haplogroup assignment drove clustering in our haplotype network (Figure 3) and sample-wide MDS plot

REJ ET AL. American Journal of Human Biology | 13 of 18

(Figure 5), according to our intrapopulation summary statis- tics (Table 3) and local AMOVA (Table 4), geography appears to play a role in structuring mtDNA variation. Despite both populations historically speaking TB languages, the Sonowal Kachari and the Rabha exhibit a much higher pairwise FST value (0.0414) than that observed between the Ahom and Sonowal Kachari. Although still significant (P 5 .027), the Ahom and Sonowal Kachari have a relatively low pairwise FST value (0.00662; Table 4). The Ahom always fall between the Rabha and Sonowals in all of our regional MDS plots (Figure 6A–D) as well. This distance between the Sonowals and the Rabha, may be explained by the geographic proximity of the former to the Ahom, as both populations were sampled in Dibrugarh. Infrequent Rabha– Sonowal mtDNA haplotype sharing can also be attributed to the endogamous marriage practices of the Sonowals and the matrilineality/exogamy of the Rabha, practice differences that are further illustrated by shared Y-chromosome microsa- tellites between the two populations (Su et al., 2000). This is confirmation of sex-biased admixture.

AMOVA indicates that 95.8% of the mtDNA variation observed in the Rabha, Ahom, and Sonowal Kachari occurs within the samples (Table 4). Similar to regional AMOVAs, high percentages of intrasample variation and significant mtDNA pairwise differences between samples have been a hallmark of South and East Asian population studies (Bodner et al., 2011; Langstieh, Reddy, Thangaraj, Kumar, & Singh, 2004; Li et al., 2007; Roychoudhury et al., 2001; Summerer et al., 2014), and Assam is no different. Significant (P < .05) pairwise differences were observed between all samples. As with our AMOVA data, the overall number of unique mtDNA HVR haplotypes observed in the combined Assam- ese sample (81.3%; Table 2) is similar to that observed in other sampled populations of South and East Asia (Bodner et al., 2011; Yao et al., 2002; Zimmermann et al., 2009).

4.6 | Regional mtDNA structure: Assam as a crossroads?

In this study, we also investigated where our sampled popu- lations from Assam fit in the greater scheme of Asian diver- sity by incorporating our comparative dataset into a series of AMOVAs and MDS analyses (Table 5; Figure 6). We observed significant population structuring. However, neither regional homeland nor language family is a better predictor of genetic variation among Asian populations, corresponding with evidence suggesting that language, geography and genetics all covary in Asian populations (Cordaux et al., 2003, 2004; Langstieh et al., 2004; Mittal et al., 2008; Roy- choudhury et al., 2001; Sahoo & Kashyap, 2006).

Overall, we find that very little variation (2–5%) can be accounted for by differences between population samples. Although small, these differences were nevertheless statisti-

cally significant in the AMOVAs comparing Northeast Indi- ans to population samples from other parts of India as well as to population samples from China. By contrast, levels of variation between Northeast Indians and samples from Southeast Asia were not statistically significant, potentially highlighting the presence of common Southeast Asian mtDNA lineages in the Rabha sample.

MDS using the full battery of comparative samples shows population grouping by both region and language family (Figure 6A). Apart from the odd outlier, we see Northeast populations (excluding the Ahom and Rabha) grouping together in the upper left portion of the plot, while non-Northeast Indians are in the bottom right. The Ahom are found interspersed among mostly TB speaking populations from Tibet and Southern China, while the Rabha appear closer to the origin of the plot, where Daic and TB speaking populations from Yunnan province are most heavily concen- trated (Figure 6A). Daic, A-A, and Austronesian populations fall in the other half of the plot from the Indian populations (x > 0). Similar divisions are observed when looking at the regional level (Figure 6B–D). All these data go to show that not only do language and geography play a role in mtDNA population structure, but so do history and cultural practices, consistent with other studies (Summerer et al., 2014; Tolk et al., 2001; Wang et al., 2012; Zhao et al., 2009; etc.). The haplogroup differentiation and separation of the Sonowal Kachari from our study and Chandrasekar in our MDS plots emphasize how even smaller geographic distances can influ- ence structure between members of the same ethnic group.

5 | CONCLUSIONS/FUTURE STUDIES

The Brahmaputra River Valley has undergone multiple colo- nization events in the �55,000 years since the initial peo- pling event, with populations from Southern China and Southeast Asia having the greatest influence on maternal lin- eages in the region. Evidence of each migratory event is apparent in the mitochondrial HVR data of the individuals sampled in this study, although these data only demonstrate a small subset of the complex history of this region. In addi- tion to effects associated with (pre)historic population move- ments, cultural practices have also left unique signatures in the HVR regions of our samples. Going forward, we are curious about how recent cultural shifts associated with the continuous modernization of Upper Assam will affect the mtDNA landscape of the newest generation of Sonowal Kacharis, Rabha, and Ahom.

Overall, these data highlight the diversity and uniqueness of the populations sampled, and provide a platform for fur- ther investigation. Future work incorporating a larger sample size and utilizing the whole mitochondrial genome could provide a higher resolution picture of overall mitochondrial

14 of 18 | American Journal of Human Biology REJ ET AL.

diversity patterns in Assam, as well as more information on haplogroup sharing, and potential answers to questions of cultural interaction and other demographic uncertainties, eliminating any ascertainment bias. As sex-biased migration has played an important role in some of the extant popula- tions in Assam, analyses of Y chromosome variation in these samples would provide a complement to the picture of maternal demographic history presented here, while whole genome data would help to better position Assam and North- east India in the overall scheme of Indian genetic diversity (Reich, Thangaraj, Patterson, Price, & Singh, 2009).

ACKNOWLEDGMENTS

We are sincerely grateful to the people of Assam for con- tributing to this study. We also thank the Charles Phelps Taft Research Center at the University of Cincinnati for providing financial support for this research.

AUTHOR CONTRIBUTIONS

PHR: designed the study, performed the experiments, ana- lyzed the data, wrote the manuscript. HLN: developed exper- imental approach, provided lab space, wrote the manuscript. RD: collected the samples and edited the manuscript.

REFERENCES

Andrews, R. M., Kubacka, I., Chinnery, P. F., Lightowlers, R. N., Turnbull, D. M., & Howell, N. (1999). Reanalysis and revision of the Cambridge reference sequence for human mitochondrial DNA. Nature Genetics, 23(2), 147.

Behar, D. M., Van, O. M., Rosset, S., Metspalu, M., Loogvali, E. L., Silva, N. M., . . . Villems, R. (2012). A “Copernican” reassessment of the human mitochondrial DNA tree from its root. American Journal of Human Genetics, 90, 675–684.

Bodner, M., Zimmermann, B., Rock, A., Kloss-Brandstatter, A., Horst, D., Horst, B., . . . Parson, W. (2011). Southeast Asian diversity: First insights into the complex mtDNA structure of Laos. BMC Evolutionary Biology, 11, 49.

Burling, R. (2003). The Tibeto-Burman languages of North- eastern India. In: G. Thurgood & R. LaPolla R (Eds.), The sino-tibetan languages. London: Routledge.

Burling, R. (2007). Language, Ethnicity, and Migration in North-Eastern India. South Asia: Journal of South Asian Studies, 3. 319–404.

Cann, R. L. (2001). Genetic clues to dispersal in human popu- lations: Retracing the past from the present. Science, 291 (5509), 1742V1748.

Cavalli-Sforza, L. L., Menozzi, P., & Piazza, A. 1994. History and geography of human genes. Princeton: Princeton Uni- versity Press.

Chakrabarti, B., Kumar, S., Singh, R., & Dimitrova, N. (2012). Genetic diversity and admixture patterns in Indian populations. Gene, 508(2), 250–255.

Chandrasekar, A., Kumar, S., Sreenath, J., Sarkar, B. N., Urade, B. P., Mallick, S., . . . Rao, V. R. (2009). Updating phylogeny of mitochondrial DNA macrohaplogroup m in India: Dispersal of modern human in South Asian corridor. PloS One, 4(10), e7447.

Chaubey, G., Metspalu, M., Choi, Y., Magi, R., Romero, I. G., Soares, P., . . . Kivisild, T. (2011). Population genetic structure in Indian Austroasiatic speakers: The role of land- scape barriers and sex-specific admixture. Molecular Biol- ogy and Evolution, 28(2), 1013–1024.

Chaubey, G., Metspalu, M., Kivisild, T., & Villems, R. (2007). Peopling of South Asia: Investigating the caste- tribe continuum in India. Bioessays, 29(1), 91–100.

Cordaux, R., Saha, N., Bentley, G. R., Aunger, R., Sirajuddin, S. M., & Stoneking, M. (2003). Mitochondrial DNA analy- sis reveals diverse histories of tribal populations from India. European Journal of Human Genetics 11(3), 253–264.

Cordaux, R., Weiss, G., Saha, N., & Stoneking, M. (2004). The northeast Indian passageway: A barrier or corridor for human migrations? Molecular Biology and Evolution, 21 (8), 1525–1533.

Das, B. M., Das, P. B., Das, R., Walter, H., & Danker-Hopfe, H. (1985). Anthropological studies in Assam, India. 1. Observations on five Mongoloid populations. Anthropolo- gischer Anzeiger; Bericht Uber Die Biologisch- Anthropologische Literatur, 43(3), 193–204.

Das, B. M., & Deka, R. (1975). Predominance of the haemo- globin E gene in a Mongoloid population in Assam (India). Humangenetik, 30(2), 187–191.

Das, B. M., Deka, R., & Flatz, G. (1987). Genetic variation of five blood group polymorphisms in ten populations of Assam, India. International Journal of Anthropology, 2, 325–340.

Das, B. M., & Sengupta, S. (2003). A Note on Some Mor- phogenetic Variables Among the Sonowal Kacharis of Assam. Anthropologist, 5, 211–212.

Deb, B. 2006. Ethnic issues, secularism, and conflict resolution in north east Asia. Delhi: Concept Publishing Company.

Deka, R. (1984). A genetic survey in four Mongoloid popula- tions of the Garo Hills, India. Anthropologischer Anzeiger; Bericht Uber Die biologischV Anthropologische Literatur, 42(1), 41–45.

Deka, R., Reddy, A. P., Mukherjee, B. N., Das, B. M., Bane- rjee, S., Roy, M., . . . Walter, H. (1988). Hemoglobin E distribution in ten endogamous population groups of Assam, India. Human Heredity, 38(5), 261V-66.

Dennell, R. (2009). The paleolithic settlement of Asia. Cam- bridge: Cambridge University Press.

Diller, A., Edmondson, J., & Luo, Y. 2004. The tai-kadai lan- guages. London: Routledge.

Excoffier, L., & Lischer, H. E. (2010). Arlequin suite ver 3.5: A new series of programs to perform population genetics analyses under Linux and Windows. Molecular Ecology Resources, 10(3):564V567),

REJ ET AL. American Journal of Human Biology | 15 of 18

Fernquest, J. (2006). Crucible of War: Burma and the Ming in the Tai Frontier Zone (1382-1454; pp. 27-81). SOAS bulle- tin of Burma research, 4(2), 27–81.

Fornarino, S., Pala, M., Battaglia, V., Maranta, R., Achilli, A., Modiano, G., . . . Santachiara-Benerecetti, S. A. (2009). Mitochondrial and Y-chromosome diversity of the Tharus (Nepal): a reservoir of genetic variation. BMC Evolutionary Biology, 9, 154.

Forster, P., Harding, R., Torroni, A., & Bandelt, H. J. (1996). Origin and evolution of Native American mtDNA varia- tion: a reappraisal. American Journal of Human Genetics, 59(4), 935–945.

Forster, P., & Matsumura, S. (2005). Evolution. Did early humans go north or south? Science, 308(5724), 965–966.

Fu, Y. X. (1997). Statistical tests of neutrality of mutations against population growth, hitchhiking and background selection. Genetics, 147(2), 915–925.

Gait, E. (1906). A history of Assam. Calcutta: Thacker, Spink, and Co.

Gayden, T. (2012). Genetic diversity in the himalayan popula- tions of nepal and tibet. Miami: Florida International University.

Gayden, T., Perez, A., Persad, P. J., Bukhari, A., Chennak- rishnaiah, S., Simms, T., . . . Herrera, R. J. (2013). The Himalayas: Barrier and conduit for gene flow. American Journal of Physical Anthropology, 151(2), 169–182.

Gazi, N. N., Tamang, R., Singh, V. K., Ferdous, A., Pathak, A. K., Singh, M., . . . Thangaraj, K. (2013). Genetic Struc- ture of Tibeto-Burman Populations of Bangladesh: Evaluat- ing the Gene Flow along the Sides of Bay-of-Bengal. PLoS One, 8(10), e75064. Oct 9

Gogoi, P. (1968). The Tai and the Tai kingdoms. Guwahati: Gauhati University Press.

Grimes, G., & Diller, A. (2003). Tai Languages. In W. Bright (Ed.), International encyclopedia of linguistics. New York: Oxford University Press.

Hazarika, M. (2011). Lithic industries with Palaeolithic elements in Northeast India. Quaternary International, 269, 48–58.

Hill, C., Soares, P., Mormina, M., Macaulay, V., Clarke, D., Blumbach, P. B., . . . Richards, M. (2007). A mitochondrial stratigraphy for island southeast Asia. American Journal of Human Genetics, 80(1), 29–43.

Kalita, D., & Deb, B. (2004). Some folk medicines used by the Sonowal Kacharis tribe of the Brahmaputra Valley, Assam. NPR 3:240V246.

Kivisild, T., Bamshad, M. J., Kaldma, K., Metspalu, M., Met- spalu, E., Reidla, M., . . . Villems, R. (1999). Deep com- mon ancestry of indian and western-Eurasian mitochondrial DNA lineages. Current Biology, 9(22), 1331–1334.

Kivisild, T., Rootsi, S., Metspalu, M., Mastana, S., Kaldma, K., Parik, J., . . . Villems, R. (2003). The genetic heritage of the earliest settlers persists both in Indian tribal and caste populations. American Journal of Human Genetics, 72(2), 313–332.

Kong, Q. P., Sun, C., Wang, H. W., Zhao, M., Wang, W. Z., Zhong, L., . . . Zhang, Y. P. (2011). Large-scale mtDNA screening reveals a surprising matrilineal complexity in east Asia and its implications to the peopling of the region. Molecular Biology and Evolution, 28(1), 513–522.

Langstieh, B. T., Reddy, B. M., Thangaraj, K., Kumar, V., & Singh, L. (2004). Genetic diversity and relationships among the tribes of Meghalaya compared to other Indian and Conti- nental populations. Human Biology, 76(4), 569–590.

Li, H., Cai, X., Winograd-Cort, E. R., Wen, B., Cheng, X., Qin, Z., . . . Jin, L. (2007). Mitochondrial DNA diversity and population differentiation in southern East Asia. Amer- ican Journal of Physical Anthropology, 134(4), 481V 488.

Li, Y. C., Wang, H. W., Tian, J. Y., Liu, L. N., Yang, L. Q., Zhu, C. L., . . . Zhang, Y. P. (2015). Ancient inland human dispersals from Myanmar into interior East Asia since the Late Pleistocene. Scientific Reports, 5, 9473.

Long, R. (2012). Assam. In S. Wolpert (Ed.), Encyclopedia of India (pp. 68–70). Detroit: Charles Scribner Sons.

Macaulay, V., Hill, C., Achilli, A., Rengo, C., Clarke, D., Meehan, W., . . . Richards, M. (2005). Single, rapid coastal settlement of Asia revealed by analysis of complete mito- chondrial genomes. Science, 308(5724), 1034–1036.

Maji, S., Krithika, S., & Vasulu, T. S. (2008). Distribution of Mitochondrial DNA Macrohaplogroup N in India with Special Reference to Haplogroup R and its Subha- plogroup U. International Journal of Human Geneitcs, 8, 85–96.

Maji, S., Krithika, S., & Vasulu, T. S. (2009). Phylogeo- graphic distribution of mitochondrial DNA macroha- plogroup M in India. Journal of Genetics, 88(1), 127–139.

Majumder, P. P. (2008). Genomic inferences on peopling of south Asia. Current Opinion in Genetics & Development, 18(3), 280–284.

Mellars, P. (2006). Going east: New genetic and archaeologi- cal perspectives on the modern human colonization of Eur- asia. Science, 313(5788), 796–800.

Metspalu, M., Kivisild, T., Metspalu, E., Parik, J., Hudjashov, G., Kaldma, K., . . . Villems, R. (2004). Most of the extant mtDNA boundaries in south and southwest Asia were likely shaped during the initial settlement of Eurasia by anatomically modern humans. BMC Genetics, 5, 26.

Mittal, B., Tripathy, V., Aruna, M., Reddy, A. G., Thanseem, I., Thangaraj, K., . . . Reddy, B. M. (2008). Mitochondrial DNA variation and substructure among the tribal popula- tions of Andhra Pradesh, India. American Journal of Human Biology: The Official Journal of the Human Biol- ogy Council, 20(6), 683–692.

Morey, S. (2015). Tai languages of Assam, a progress report - Does anything remain of the Tai Ahom language? In D. Brad- ley & M. Bradley (Eds.), Language endangerment and lan- guage maintenance: An active approach. London: Routledge.

Nei, M., & Li, W. H. (1979). Mathematical model for study- ing genetic variation in terms of restriction endonucleases.

16 of 18 | American Journal of Human Biology REJ ET AL.

Proceedings of the National Academy of Sciences of the United States of America, 76(10), 5269–5273.

Nei, M., & Roychoudhury, A. K. (1993). Evolutionary rela- tionships of human populations on a global scale. Molecu- lar Biology and Evolution, 10(5), 927–943.

Peng, M. S., Palanichamy, M. G., Yao, Y. G., Mitra, B., Cheng, Y. T., Zhao, M., . . . Zhang, Y. P. (2011). Inland post-glacial dispersal in East Asia revealed by mitochon- drial haplogroup M9a’b. BMC Biology, 9, 2.

Qin, Z., Yang, Y., Kang, L., Yan, S., Cho, K., Cai, X., . . . Li, H. (2010). A mitochondrial revelation of early human migrations to the Tibetan Plateau before and after the last glacial maximum. American Journal of Physical Anthro- pology, 143(4), 555–569.

Raha, M. (1989). Matriliny to patriliny: A study of the rabha society. Delhi: Gyan Publishing House.

Reddy, B. M., Langstieh, B. T., Kumar, V., Nagaraja, T., Reddy, A. N., Meka, A., . . . Singh, L. (2007). Austro- Asiatic tribes of Northeast India provide hitherto missing genetic link between South and Southeast Asia. PloS One, 2(11), e1141.

Reich, D., Thangaraj, K., Patterson, N., Price, A. L., & Singh, L. (2009). Reconstructing Indian population history. Nature, 461(7263), 489–494.

Rej, P. (2013). Measuring mitochondrial DNA diversity and demographic patterns of tribal and caste populations from the Northeast Indian State of Assam. Cincinnati: University of Cincinnati.

Ripunjoy, S. (2013). Indigenous knowledge on the utilization of medicinal plants by the Sponowal Kachari tribe of Dibrugarh District in Assam, North-East India. Interna- tional Research Journal of Biological Sciences, 2, 44–50.

Roychoudhury, S., Roy, S., Basu, A., Banerjee, R., Vishwana- than, H., Usha Rani, M. V., . . . Majumder, P. P. (2001). Genomic structures and population histories of linguisti- cally distinct tribal groups of India. Human Genetics, 109 (3), 339–350.

Roychoudhury, S., Roy, S., Dey, B., Chakraborty, P., Roy, M., Roy, B., . . . Majumder, P. P. (2000). Fundamental genomic unity of ethnic India is revealed by analysis of mitochondrial DNA. Current Science India, 79(9), 1182–1192.

Sahoo, S., & Kashyap, V. K. (2006). Phylogeography of mitochondrial DNA and Y- chromosome haplogroups reveal asymmetric gene flow in populations of Eastern India. American Journal of Physical Anthropology, 131(1), 84–97.

Saiki, R. K., Gelfand, D. H., Stoffel, S., Scharf, S. J., Higuchi, R., Horn, G. T., . . . Erlich, H. A. (1988). Primer-directed enzymatic amplification of DNA with a thermostable DNA polymerase. Science, 239(4839), 487–491.

Saikia, Y. 2004. Fragmented memories: Struggling to be Tai- Ahom in India. Durham, NC: Duke University Press.

Saillard, J., Forster, P., Lynnerup, N., Bandelt, H. J., & Norby, S. (2000). mtDNA variation among Greenland

Eskimos: The edge of the Beringian expansion. American Journal of Human Genetics, 67(3), 718–726.

Sali, S. (1989). The upper paleolithic and mesolithic cultures of Maharashtra. Pune: Deccan College.

Sambrook, J., Fritsch, E., & Maniatis, T. (1989). Molecular cloning: A laboratory manual. New York: Cold Spring Harbor Laboratory.

Schurr, T. G., & Wallace, D. C. (2002). Mitochondrial DNA diversity in Southeast Asian populations. Human Biology, 74(3), 431–452.

Scozzari, R., Cruciani, F., Santolamazza, P., Malaspina, P., Tor- roni, A., Sellitto, D., . . . Novelletto, A.. (1999). Combined use of biallelic and microsatellite Y-chromosome polymorphisms to infer affinities among African populations. American Jour- nal of Human Genetics, 65(3), 829–846.

Sharma, H. C. (2003). Prehistoric archaeology of the North- east. In T. Subba,& G. Ghosh (Eds.), The Anthropology of North-East India (pp. 11–30). New Delhi: Orient Longman Private Limited.

Sharma, S. (2002). Geomorphological contexts of the Stone Age record of the Ganol and Rongram vallyes in the Garo Hills, Meghalaya. Man and Environment, 27, 15–30.

Soares, P., Ermini, L., Thomson, N., Mormina, M., Rito, T., R€ohl, A., . . . Richards, M. B. (2009 ). Correcting for Puri- fying Selection: An Improved Human Mitochondrial Molecular Clock. American Journal of Human Genetics, 84(6), 740–759. Jun 12;

Su, B., Xiao, C., Deka, R., Seielstad, M. T., Kangwanpong, D., Xiao, J., . . . Jin, L. (2000). Y chromosome haplotypes reveal prehistorical migrations to the Himalayas. Human Genetics, 107(6), 582–590.

Summerer, M., Horst, J., Erhart, G., Weissensteiner, H., Schonherr, S., Pacher, D., . . . Brandstätter, A. (2014). Large-scale mitochondrial DNA analysis in Southeast Asia reveals evolutionary effects of cultural isolation in the multi-ethnic population of Myanmar. BMC Evolutionary Biology, 14, 17.

Tajima, F. (1989). Statistical method for testing the neutral mutation hypothesis by DNA polymorphism. Genetics, 123 (3), 585–595.

Tamura, K., Stecher, G., Peterson, D., Filipski, A., & Kumar, S. (2013). MEGA6: Molecular Evolutionary Genetics Analysis version 6.0. Molecular Biology and Evolution, 30 (12), 2725–2729.

Thangaraj, K., Chaubey, G., Kivisild, T., Reddy, A. G., Singh, V. K., Rasalkar, A. A., & Singh, L. (2005). Reconstructing the origin of Andaman Islanders. Science, 308(5724), 996.

Thangaraj, K., Chaubey, G., Singh, V. K., Vanniarajan, A., Thanseem, I., Reddy, A. G., & Singh, L. (2006). In situ origin of deep rooting lineages of mitochondrial Macroha- plogroup ’M’ in India. BMC Genomics, 7, 151.

Tolk, H. V., Barac, L., Pericic, M., Klaric, I. M., Janicijevic, B., Campbell, H., . . . Rudan, P. (2001). The evidence of mtDNA haplogroup F in a European population and its

REJ ET AL. American Journal of Human Biology | 17 of 18

ethnohistoric implications. European Journal of Human Genetics 9(9), 717–723.

van Driem, G. 2001. Languages of the Himalayas: An Ethno- linguistic Handbook of the Greater Himalayan Region: Brill.

van Driem, G. 2011. Lost in the sands of time somewhere north of the Bay of Bengal. In: Turin M, and Zeisler B, editors. Himalayan languages and linguistics: studies in phonology, semantics, morphology, and syntax. Boston: Brill.

van Oven, M., & Kayser, M. (2009). Updated comprehensive phylogenetic tree of global human mitochondrial DNA var- iation. Human Mutation, 30(2), E386–E394.

Wang, H. W., Li, Y. C., Sun, F., Zhao, M., Mitra, B., Chaudhuri, T. K., . . . Zhang, Y. P. (2012). Revisiting the role of the Himalayas in peopling Nepal: insights from mitochondrial genomes. Journal of Human Genetics, 57(4), 228–234.

Watterson, G. A. (1975). On the number of segregating sites in genetical models without recombination. Theoretical Population Biology, 7(2), 256–276.

Wen, B., Xie, X., Gao, S., Li, H., Shi, H., Song, X., . . . Su, B., et al. (2004). Analyses of genetic structure of Tibeto- Burman populations reveals sex-biased admixture in south- ern Tibeto-Burmans. American Journal of Human Genet- ics, 74(5), 856–865.

Yao, Y. G., Kong, Q. P., Bandelt, H. J., Kivisild, T., & Zhang, Y. P. (2002). Phylogeographic differentiation of mitochondrial DNA in Han Chinese. American Journal of Human Genetics, 70(3), 635–651.

Zhang, X., Liao, S., Qi, X., Liu, J., Kampuansai, J., Zhang, H., . . . Su, B. (2015). Y-chromosome diversity suggests

southern origin and Paleolithic backwave migration of Austro-Asiatic speakers from eastern Asia to the Indian subcontinent. Science Reports, 5, 15486. Oct 20

Zhao, M., Kong, Q. P., Wang, H. W., Peng, M. S., Xie, X. D., Wang, W. Z., Jiayang, . . . Zhang, Y. P. (2009). Mito- chondrial genome evidence reveals successful Late Paleo- lithic settlement on the Tibetan Plateau. Proceedings of the National Academy of Sciences of the United States of America, 106(50), 21230–21235.

Zimmermann, B., Bodner, M., Amory, S., Fendt, L., Rock, A., Horst, D., . . . Brandstatter, A. (2009). Forensic and phylogeographic characterization of mtDNA lineages from northern Thailand (Chiang Mai). International Journal of Legal Medicine, 123(6), 495–501.

SUPPORTING INFORMATION

Additional Supporting Information may be found in the online version of this article.

How to cite this article: Rej PH, Deka R, and Norton HL. Understanding influences of culture and history on mtDNA variation and population structure in three pop- ulations from Assam, Northeast India. Am J Hum Biol. 2017;29:e22955. doi:10.1002/ajhb.22955.

18 of 18 | American Journal of Human Biology REJ ET AL.

Copyright of American Journal of Human Biology is the property of John Wiley & Sons, Inc. and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use.