SYNTHESIS - Senior Seminar - Screening and Diagnostic Tools

profileRmedina1087
Lebersfeldetal.2021.pdf

Vol.:(0123456789)1 3

Journal of Autism and Developmental Disorders https://doi.org/10.1007/s10803-020-04839-z

O R I G I N A L PA P E R

Systematic Review and Meta‑Analysis of the Clinical Utility of the ADOS‑2 and the ADI‑R in Diagnosing Autism Spectrum Disorders in Children

Jenna B. Lebersfeld1 · Marissa Swanson1 · Christian D. Clesi1 · Sarah E. O’Kelley1

Accepted: 9 December 2020 © The Author(s), under exclusive licence to Springer Science+Business Media, LLC part of Springer Nature 2021

Abstract The Autism Diagnostic Observation Schedule, Second Edition (ADOS-2) and the Autism Diagnostic Interview, Revised (ADI-R) have high accuracy as diagnostic instruments in research settings, while evidence of accuracy in clinical settings is less robust. This meta-analysis focused on efficacy of these measures in research versus clinical settings. Articles (n = 22) were analyzed using a hierarchical summary receiver operating characteristics (HSROC) model. ADOS-2 performance was stronger than the ADI-R. ADOS-2 sensitivity and specificity ranged from .89-.92 and .81-.85, respectively. ADOS-2 accuracy in research compared with clinical settings was mixed. ADI-R sensitivity and specificity were .75 and .82, respectively, with higher specificity in research samples (Research = .85, Clinical = .72). A small number of clinical studies were identified, indicating ongoing need for investigation outside research settings.

Keywords Autism spectrum disorder · ADOS-2 · ADI-R · Meta-analysis · Diagnosis · HSROC

Introduction

Diagnostic evaluations are crucial for children with autism spectrum disorder (ASD) to access early intervention ser- vices and therapies. Determining the accuracy of meas- ures commonly used for ASD assessment is necessary to aid clinicians in making better and more accurate clinical diagnoses. A comprehensive evaluation for autism spectrum disorder (ASD) is most accurately conducted by a multidis- ciplinary team through the use of information from multi- ple sources, including a clinical observation of the child, an ASD-focused clinical interview with caregivers, and child and family history (Risi et al. 2006; Kim and Lord 2012; Stewart et al. 2014). The Autism Diagnostic Observation Schedule, Second Edition (ADOS-2; Lord et al. 2012a) and the Autism Diagnostic Interview, Revised (ADI-R; Rutter

et al. 2003; Howes et al. 2017; Penner et al. 2017) have high levels of diagnostic accuracy; however, both instruments require specialized training and experience to administer and score. Using both instruments together improves diagnostic accuracy (sensitivity .70-.98, specificity .80-.96) compared to each measure alone (Risi et al. 2006; Ventola et al. 2006; Kim and Lord 2012). The multidisciplinary team, often led by a clinical psychologist or physician (e.g., developmental/ behavioral pediatrician), takes the results of these measures as well as other information gathered during the evaluation and uses clinical judgement to render a final diagnosis. The ADOS-2 and ADI-R were initially developed as research tools and have been studied at length in the research litera- ture, with subsequent publication for use in clinical settings to aid diagnosis. Much of the literature published on the accuracy of the ADOS-2 and the ADI-R utilized evaluations from populations recruited specifically for research, and these are the studies on which the published psychometrics were based. However, research samples often utilize strict exclusion criteria, such as excluding children with behavio- ral challenges, intellectual disability, and genetic disorders, to create a more homogenous research sample. Therefore, results may not generalize to a clinical community sample (de Bildt et al. 2004; Tomanik et al. 2007; Neuhaus et al. 2017), and it is important to understand the accuracy of

Supplementary Information The online version of this article (https ://doi.org/10.1007/s1080 3-020-04839 -z) contains supplementary material, which is available to authorized users.

* Jenna B. Lebersfeld [email protected]

1 University of Alabama at Birmingham, 1720 7th Ave S, Birmingham, AL 35233, USA

Journal of Autism and Developmental Disorders

1 3

these measures when used in clinical practice compared to research settings.

Several statistical approaches are available and accepted in evaluating the accuracy of these measures in diagnosing ASD. Sensitivity (Se) is the likelihood that a child with a clinical diagnosis of ASD will score in the ASD range on the measure, and specificity (Sp) indicates the likelihood that a child without ASD will score in the non-ASD range on the measure. Positive predictive value (PPV) is the likeli- hood that a child who received an ASD classification on a measure truly has a diagnosis of ASD, and negative predic- tive value (NPV) is the likelihood that a child who scores in the non-ASD range on a measure will not receive a clinical ASD diagnosis. PPV and NPV are influenced by the preva- lence of the disorder in the sample whereas Se and Sp are not; therefore, Se and Sp are used to measure diagnostic test accuracy when comparing across samples. The accuracy of the ADOS-2 and ADI-R has been shown to be lower in clini- cal settings compared to the research context (de Bildt et al. 2009; Zander et al. 2016; Langmann et al. 2017; Zander et al. 2017; Kamp-Becker et al. 2018); however, the majority of these clinical studies were conducted in countries outside of the United States, including the Netherlands (de Bildt et al. 2009; Oosterling et al. 2010), Greece (Papanikolaou et al. 2009), Australia (Dereu et al. 2012; Gray et al. 2008), Germany (Kamp-Becker et al. 2018) and Sweden (Zander et al. 2015; Zander et al. 2017). Differences in sociocultural norms as well as translations of the originally published measures may have increased the error associated with these measures in clinical settings. Given that children referred for clinical evaluations in the community often have more com- plex presentations than samples recruited for and included in research studies, it was hypothesized that these two diag- nostic tools may be less accurate in clinical settings than the reported psychometrics from large scale studies conducted in the research setting. Therefore, the purpose of this system- atic review and meta-analysis was to determine the accuracy and clinical utility of the ADOS-2 and the ADI-R.

Methods

This systematic review and meta-analysis utilized methods outlined in the Preferred Reporting Items for Systematic Review and Meta-Analysis (PRISMA) of Diagnostic Test Accuracy studies guidelines (McInnes et al. 2018) and the Handbook for Diagnostic Test Accuracy Reviews (Deeks 2013) and was approved by the university Institutional Review Board. This protocol was registered with PROS- PERO 2018 (https ://www.crd.york.ac.uk/prosp ero/displ ay_ recor d.php?ID = CRD42018111589, Registration number: CRD42018111589).

Measures for Index Tests

Autism Diagnostic Observation Schedule, Second Edition (ADOS‑2; Lord et al. 2012a)

The ADOS-2 is a semi-structured, 45- to 60-minute obser- vation and interaction session with an evaluator and the child which is used to aid in the diagnosis of ASD. Only the ADOS-2 and its direct precursors were considered as accept- able index tests (i.e., ADOS-Toddler (Luyster et al. 2009; ADOS-G with revised algorithms (Gotham et al. 2007)), as they formed the basis for the WPS ADOS-2 publication. For ease of reference, these will be referred to collectively as the “ADOS-2.” Older ADOS versions were not considered eli- gible index tests for the purpose of this study (i.e., ADOS-G without the revised algorithms (Lord et al. 2000); PL-ADOS (DiLavore et al. 1995)). Published ADOS-2 sensitivity (Se) ranges from .60 to .95 and specificity (Sp) ranges from .75 to 1.00 (Lord et al. 2012a, b). A recent meta-analysis indicated pooled Se ranging from .77 to .90 and Sp ranging from .62 to .90 for the ADOS-2 (Dorlack et al. 2018).

Autism Diagnostic Interview, Revised (ADI‑R; Rutter et al. 2003).

The ADI-R is a semi-structured diagnostic interview given to a parent or caregiver by a trained clinician asking detailed questions about development and underlying behaviors asso- ciated with ASD. Each section of the algorithm has a raw score cut-off, and a child must meet or exceed the cut-off in all four sections to receive a classification of autism or not autism if any domain cut-off is not exceeded. Se and Sp were not published in the ADI-R manual; however, origi- nal research literature conducted prior to measure publica- tion indicated that Se varied widely and ranged from .19 to .88, and Sp was 1.00 (Cox et al. 1999; Gilchrist et al. 2001; Table 1). More recent literature suggests the Se of the ADI-R ranges from .53 to .92, and Sp ranges from .62 to .95 (Risi et al. 2006; Falkmer et al. 2013). Studies which used other algorithms such as the ADI-R Toddler diagnostic algo- rithms or those developed by the Autism Genetic Research Exchange (AGRE) were excluded given these algorithms are not yet published for clinical use.

Table 1 Sensitivity and specificity of published ADI-R algorithms

Se sensitivity, Sp specificity

Article n Se Sp

Cox 1999 ADI-R at 20 months, diagnosis at 42 months 45 .19 1.00 ADI-R at 42 months, diagnosis at 42 months 45 .48 1.00 Gilchrist 2001 53 .88 1.00

Journal of Autism and Developmental Disorders

1 3

Eligibility Criteria

Studies administering either one or both of the ADI-R and the ADOS-2 to children under 18 years for the purpose of an initial diagnostic evaluation in a clinical or research setting were eligible. A clinical setting was defined as a commu- nity setting where participants were not recruited specifically for research. A research setting included any studies which recruited participants for research. Some studies included participants from both community and research settings and were classified as such. Studies in which diagnostic tests were administered to confirm ASD diagnosis or assess treat- ment outcome were excluded. Data for all included articles were collected in the United States, Canada, or the United Kingdom and were published in English.

Reference Standard for Diagnosis

The reference standard for diagnosis was the final consensus diagnosis of a comprehensive evaluation for ASD (i.e., ASD or non-ASD), using the following conservative approach. The comprehensive evaluation must have included any ver- sion of the ADOS and any ASD-focused clinical interview. Papers using the ADOS-2 as the index test were required to include some type of ASD-focused clinical interview in the evaluation, but this interview did not necessarily have to be the ADI-R. For studies in which the ADI-R served as the index test, the evaluation must have included the administra- tion of any version of the ADOS (i.e., PL-ADOS, ADOS-G, ADOS-G with revised algorithms, ADOS-T, or ADOS-2) but did not need to include the ADOS-2 specifically. Papers were excluded in which the ADI-R was administered but no ver- sion of the ADOS was administered.

Studies which used another method for determining ASD or non-ASD diagnosis (e.g., pre-determined algorithm based on a combination of ADOS and ADI-R results) were not included, given that this type of methodology for determin- ing ASD diagnosis does not reflect clinical practice. Studies which did not report a final consensus clinical diagnosis and included only diagnoses reported by a parent, pediatrician, or educator, and/or other forms of ASD diagnosis were not included.

Study Design

Article eligibility included peer-reviewed original research with prospective, retrospective, cross-sectional, or longi- tudinal study designs. Case studies and case series were excluded. Review articles, meta-analyses, and grey litera- ture were not included, but citations within were reviewed.

Search Strategy

Searches were conducted in September 2018 from Psy- cINFO, ERIC, PubMed/MEDLINE, Cochrane Database of Systematic Reviews (including Cochrane Central Register of Controlled Trials (CENTRAL)), Journal of Autism and Developmental Disorders, Research in Autism Spectrum Disorders, Autism Research, and Autism. Google Scholar was used informally to identify keywords but not included in the formal search strategy. All articles published since the original publication date of each measure were considered (ADI-R - 2003, ADOS-2 with revised algorithms—2007). Detailed search terms are included in Online Appendix A.

Assessment of Methodological Quality

The QUADAS-2 (Quality Assessment of Diagnostic Accu- racy Studies-2, Whiting et al. 2011) is a tool used in sys- tematic reviews to evaluate risk of bias and applicability concerns in diagnostic test accuracy studies related to patient selection, index tests, reference standard, and flow and tim- ing. The QUADAS-2 tool for this study was adapted and operationalized from Vllasaliu et al. (2016).

Study Selection

Figure 1 reviews the process by which articles were selected for inclusion in the study. Citations from searches (n = 11,672) were exported into EndNote and duplicate articles (n = 2,591) were eliminated automatically. An additional 949 duplicate articles were identified manually. Therefore, 8,132 unique citations were reviewed. All titles, abstracts, and possibly relevant full-text articles were reviewed by two authors (JL and MS). Given differences in initial article eligibility identification between the two authors resulting in low initial agreement (i.e., 48 articles, 23% agreement), the inclusion criteria were clarified, and articles were re- reviewed by the same two authors yielding 62% agreement. Remaining discrepancies were rectified through discussion between these two authors and the last author (SO) as well as via outside review by two clinical psychologists with research backgrounds and expertise in ASD. These outside reviewers had 100% agreement with one another. These procedures resulted in 22 articles deemed appropriate for inclusion in the meta-analysis, with 14 articles included in the ADOS-2 analyses and 13 papers included in the ADI-R analyses. Despite the complexity of the inclusion criteria, the additional steps taken to rectify initial low agreement likely resulted in the inclusion of all appropriate papers in the meta-analysis.

Journal of Autism and Developmental Disorders

1 3

Fig. 1 PRISMA flow diagram

Journal of Autism and Developmental Disorders

1 3

Data Extraction

True positives (TP), false positives (FP), true negatives (TN), false negatives (FN), Se, and Sp for the ADOS-2 and/ or the ADI-R classifications were extracted by two authors (JL and CC). For some articles, these metrics were stated directly in the text or presented in supplementary materials. For articles in which these numbers were not directly stated, these statistics were calculated using the Review Manager (RevMan) software provided by Cochrane Library (Review Manager 2014). A total of 116 data points was extracted by JL and CC, and 112 data points were agreed upon (97%) across the 22 articles. Discrepancies were identified as errors due to referring to wrong text in the table (n = 2) or typo- graphical or calculation errors (n = 2).

Data Analysis

Articles were organized using the RevMan software. The hierarchical summary receiver operating characteristic (HSROC) model of Rutter and Gatsonis (Rutter 1995; Rut- ter and Gatsonis 2001) was conducted using the MetaDAS SAS macro (Takwoingi and Deeks 2010). This model pro- duces pooled Se and Sp and accounts for the correlation between Se and Sp across studies. Separate pooling of Se and Sp results in underestimation of these statistics, since it does not take into account the inherent trade-off between these statistics (Deeks 2001). Positive and negative predic- tive values are influenced by prevalence in the sample, which introduce heterogeneity and uncertainty. The chosen method for statistical analysis uses a Bayesian model to determine random effects and was preferred to fixed effects due to the large amount of heterogeneity commonly seen among diagnostic test accuracy studies. Additionally, the HSROC method is recommended when covariates are included in the model. This model also produces the Diagnostic Odds Ratio (DOR), a global estimate of overall test accuracy. The DOR is a summary of the diagnostic accuracy of a test and can be interpreted as how many times higher the odds are of a per- son with ASD to score in the ASD range on the diagnostic test compared to someone without ASD. DOR can be used to interpret and compare across tests and models.

Statistical analyses were conducted separately for the ADOS-2 and the ADI-R. The HSROC model was computed with and without the setting covariate to determine whether setting had an effect on diagnostic test accuracy. The setting covariate included three groups: clinical, research, and both.

For the ADI-R analysis, the model converged using these three groups. For the ADOS-2 analyses, having three groups did not allow the model to converge. Therefore, the “both” group was combined with the “research” group for the

ADOS-2 analyses of the setting covariate. Combining the “both” group with the “research group” was viewed as the more conservative approach compared with combining the “both” and “clinical” groups. If, as hypothesized, the admin- istration of the ADOS-2 in research settings was more accu- rate than clinical settings, including articles with clinical evaluations in the “research” group would dilute the accu- racy of the ASD diagnostic measures within the “research” setting and reduce the difference in accuracy of the ADOS-2 in clinical and research settings in this study.

Outliers and Sensitivity Analysis

Studies were plotted graphically on HSROC plots and visually inspected for outliers, with one article with low Sp identified as an outlier in the ADOS-2 analysis. Study characteristics were reviewed, and low sensitivity was likely due to the clinical population, which included many children with severe developmental and behavioral chal- lenges, resulting in many false positives on the ADOS-2. Although these children are often excluded from research studies, they present for clinical evaluations, and it is important to investigate the accuracy of diagnostic meas- ures in these populations. However, these study results may not generalize to other clinical settings given the sample characteristics. Therefore, the authors conducted analyses both with and without the outlier. Sp analyses were conducted by removing the outlier article and repeat- ing the analyses. Results were compared with and without the outlier to determine the effect of this specific study on the results, as discussed below. No outliers were identified for the ADI-R analysis.

The Gotham et al. (2007, 2008) papers provided sepa- rate Se and Sp estimates based on differing criteria from the Diagnostic and Statistical Manual for Mental Disorders, Fourth Edition (DSM-IV, American Psychiatric Association 2000) for two instances: Autism (i.e., Autistic Disorder) vs. Non-spectrum (NS) and ASD vs. NS. In the Autism vs. NS analysis, PDD-NOS and Asperger Disorder cases were excluded and ADOS-2 classifications of ASD were classified as non-spectrum. In the ASD vs NS condition, children with Autistic Disorder were excluded and ADOS-2 classifications of “autism spectrum” and “autism” were both considered classifications of ASD. For the purposes of the current study, including both estimates in a single analysis would result in the inclusion of the non-spectrum cases more than once, thus separate analyses were conducted for the Autism vs. NS and ASD vs. NS estimates for the Gotham et al. articles. Additionally, results were analyzed both with and without the outlier. For clarity, analytic approaches for the ADOS-2 are defined in Table 2.

Journal of Autism and Developmental Disorders

1 3

Results

Table  3 outlines study characteristics for the 22 articles included in the meta-analysis.

Quality of the Included Studies

Figure 2 displays metrics used for evaluating quality of the studies including risk of bias and applicability concerns.

Risk of bias was unclear or high risk for 12 of the 22 papers (54%), and there were concerns regarding the use of the reference standard (i.e., unclear or high risk for all articles). This was primarily due to the clinicians’ knowledge of the results of the index tests prior to the implementation of the reference standard, as opposed to using blind raters to come to a diagnostic conclusion. This is common practice in clinical settings, as the index tests (i.e., the ASD diagnos- tic measures) are inextricably linked and used as a primary source of information in the reference standard (i.e., ASD diagnostic evaluation and final clinical diagnosis) (Figs 3 and 4). Overall, there was low risk of bias from the index tests, flow and timing, and applicability of the findings to practice.

Diagnostic Accuracy of Measures

ADOS‑2

Estimates of overall Se (.89–.92) and Sp (.81–.85) of the ADOS-2 as well as individual estimates for identified articles

Table 2 ADOS-2 Analytical Approaches

Approach Outlier article Gotham et al. 2007, 2008

1 Included ASD vs. NS 2 Included Autism vs. NS 3 Excluded ASD vs. NS 4 Excluded Autism vs. NS 5 Included Excluded 6 Excluded Excluded

Table 3 Study characteristics

a 57 to 86% male b Diagnosis deferred n = 14 c 4 years for younger group, 9 years for older group d One participant diagnosis not reported

Study Test(s) Total N Sex Age Diagnosis

ADOS-2 ADI-R Male n Female n M or Range ASD n Non-ASD n

1. Baird 2006 X 255 223 32 12 years 158 97 2. Bishop et al. 2017 X X 289 203 86 8 years 142 126 3. Camodeca 2018 X 483 355 128 10 years 127 356 4. Dykens 2017 X 146 72 74 11 years 32 114 5. Gillentine 2017 X X 18 12 6 9 years 7 10 6. Gotham et al. 2007 X 1630 a a 41 to 104 months 1,351 279 7. Gotham et al. 2008 X 1282 923 359 37 to 118 months 1,068 214 8. Grzadzinski 2016 X X 212 176 36 9 years 164 48 9. Guthrie 2013 X 82b 64 18 19 months 56 12 10. Harris 2008 X 63 63 0 8 years 38 25 11. Havdahl 2016 X 389 288 101 c 255 163 12. Kim 2012 X 695d 353 160 33 months 491 203 13. Le Couteur 2008 X 101 81 20 36 months 77 24 14. Luyster 2009 X 206 158 48 15 to 26 months 59 147 15. Mazefsky 2006 X 78 56 22 4 years 59 19 16. Molloy 2011 X 584 507 77 3 to 9 years 329 255 17. Risi 2006 X 1039 818 221 27 to 94 months 881 158 18. Ventola 2006 X 45 37 8 26 months 36 9 19. Wiggins 2008 X 142 112 30 26 months 73 69 20. Wiggins 2015 X X 922 581 341 59 months 584 338 21. Ziats 2016 X X 18 14 4 14 years 8 10 22. Zwaigenbaum 2016 X 381 215 166 39 months 103 278

Journal of Autism and Developmental Disorders

1 3

Fig. 2 QUADAS-2 risk of bias and applicability concerns

Journal of Autism and Developmental Disorders

1 3

(Se =.85–1.00; Sp =.44–1.00) are presented in Table 4 and Fig. 5. These estimates were generally comparable to pub- lished algorithms (Table 5). Addition of the setting covariate

was significant (−2LL = 7.87, p < .05) when the Gotham et  al. (2007, 2008) papers were excluded (Table  4). The highest DOR was reported within clinical samples when the outlier was excluded and the Gotham et al. (2007, 2008) papers utilized the Autism vs. Non-Spectrum algorithms (Table 4). When all articles were included, the DOR was higher for research compared with clinical samples; how- ever, inclusion of the setting covariate was not significant (p =.071). Exclusion of the outlier had little effect on Se of the clinical sample but increased the Sp of the clinical sample from .80 to .90, which is higher than specificities reported in research samples (.81 and .83; Table 4).

Interpretation of the SROC plot (Fig. 3) for all three set- ting types (clinical, research, and both) when all articles were included in the analysis and the Gotham et al. (2007, 2008) ASD vs. NS accuracy estimates were used (Approach 1) suggests research samples have higher levels of accuracy compared with clinical samples and combined clinical and research samples. When the outlier (Sp =.44) was removed from the analysis (Fig. 4), and the ASD vs. NS accuracy estimates were used (Approach 3), visual inspection of the SROC curve suggests there was not a difference between accuracy of the ADOS-2 in research and clinical settings, and accuracy of the ADOS-2 for studies including both research and clinical evaluations was lower than either research or clinical settings individually.

ADI‑R

The ADI-R pooled Se was .75, Sp was .82, and individual articles ranged widely (Se =.33–1.00, Sp =.61–1.00, see Fig. 6 and Table 6).

Inclusion of the setting covariate in the model compared to the model without the covariate trended toward signif- icance (−2LL difference = 11.788, p = .067, see Fig. 7). Clinical and research samples had comparable Se (clinical = .71, research = .73) but articles utilizing both research and clinical samples had higher Se (.82). Sp was higher for research samples (.85) compared to clinical samples (.72) and those including both research and clinical evaluations in the study (.76, see Table 6 and Fig. 6).

Discussion

This study utilized a systematic review and meta-analysis to investigate the accuracy of the ADOS-2 and the ADI-R in clinical settings compared to research settings, and it was hypothesized that these measures would perform better in research settings given the heterogeneity and complexity of children referred for an ASD evaluation in clinical samples. ADOS-2 accuracy from the meta-analysis was comparable

Fig. 3 SROC plot of ADOS-2 by setting for Approach 1 (outlier included). Note: Size of shape indicates sample size

Fig. 4 SROC plot of ADOS-2 by setting for Approach 3 (outlier excluded). Note: Size of shape indicates sample size

Journal of Autism and Developmental Disorders

1 3

to accuracy reported in the published manual and was more accurate than the ADI-R in both research and clinical set- tings. For the ADI-R, the current meta-analysis painted a more nuanced picture than the literature cited in the pub- lished manual with overall Se of .75 and Sp of .82, and the ADI-R was less accurate in clinical studies compared to research-only studies or those utilizing both research and clinical samples.

For the ADOS-2, when comparing samples of children evaluated in clinical settings with those whose evalua- tions were completed in research settings (or which used a combination of clinical and research evaluations), analyses indicated Se was comparable across settings and Sp results were mixed. Some analyses indicated comparable or slightly lower Sp in clinical compared to research samples, whereas when an outlier was excluded, results showed that Sp in clin- ical samples was higher than research samples. This suggests

Table 4 Sensitivity and specificity of ADOS-2 overall and by evaluation setting

Se sensitivity, Sp specificity, DOR diagnostic odds ratio, −2LL −2 log likelihood difference, “–-” data not available * p < .05

Approach Overall Research or both Clinical −2LL p

n Se Sp DOR n Se Sp DOR n Se Sp DOR

1 14 .89 .81 36.5 11 .89 .80 34.8 3 .89 .80 31.0 7.02 .071 2 14 .92 .83 52.7 11 .92 .83 59.2 3 .89 .80 30.9 5.81 .120 3 13 .89 .83 42.3 11 .89 .81 36.3 2 .88 .90 71.1 3.23 .357 4 13 .92 .85 61.9 11 .92 .83 59.7 2 .88 .90 70.8 2.83 .418 5 12 .91 .81 47.0 9 .93 .81 53.8 3 .89 .80 31.0 7.87 .049* 6 11 .92 .84 55.7 – – – – – – – – 7.53 .057

Fig. 5 Forest plot of ADOS-2 by setting using the Gotham ASD vs. NS estimates

Table 5 Sensitivity and specificity of published ADOS algorithms

NVMA nonverbal mental age, ASD autism spectrum disorder, NS non-spectrum, AUT autism, Se sensitivity, Sp specificity, “–”data not available

Gotham et al. 2007 Gotham et al. 2008

Module and algorithm AUT vs. NS ASD vs. NS AUT vs. NS ASD vs. NS

Se Sp Se Sp Se Sp Se Sp

Module 1, no words, NVMA > 15 mo.

.95 .94 .82 .79 .86 .80 – –

Module 1, some words .97 .91 .77 .82 .89 .91 .95 1.00 Module 2, younger .98 .93 .84 .77 .94 1.00 .65 .88 Module 2, older .98 .90 .83 .83 – – – – Module 3 .91 .84 .72 .76 .82 .92 .60 .75

Journal of Autism and Developmental Disorders

1 3

that given the small number of studies identified that were conducted in solely clinical settings, a single article can have a large effect on results. Therefore, more research is needed to further examine ADOS-2 performance in clinical evaluations. Given current findings, Sp of the ADOS-2 may be more variable across clinical settings, whereas Se may remain relatively stable.

Sources of Heterogeneity

One limitation of this meta-analysis is that additional sources of heterogeneity were not investigated due to the limited number of eligible articles identified for inclusion. One consideration is the shifting definition of autism spec- trum disorder over time. Current diagnoses are based on cri- teria for Autism Spectrum Disorder outlined in the Diagnos- tic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5, American Psychiatric Association 2013), which conceptualizes ASD as a single disorder with differing levels of severity. The DSM-IV defined multiple types of autism spectrum disorders including Asperger’s Disorder; Autis- tic Disorder; and Pervasive Developmental Disorder, Not Otherwise Specified (PDD-NOS). However, these disorders could not be reliably differentiated, which led to the revi- sion of the diagnostic criteria in the DSM-5. The ADOS-2 revised algorithms reflect this change in conceptualization of ASD. However, the ADI-R has not yet been updated, and the ADI-R manual states the measure only reliably differentiates between those with Autistic Disorder (DSM-IV) and other non-spectrum conditions, not those with milder ASD symp- toms. These factors further complicate the already multifac- eted ASD diagnostic process. Within the research literature, ADI-R algorithms have been developed to capture milder presentations of ASD. However, these types of algorithms are used predominately for research, have not been published for clinical use, and are not widely used clinically. Given that a primary aim of this study was to investigate how

Fig. 6 ADI-R forest plot by setting

Table 6 Sensitivity and specificity of ADI-R Overall and by evalua- tion setting

Setting n Sens Spec DOR

Overall 13 .75 .82 13.6 Research 9 .73 .85 15.8 Both 2 .82 .76 15.9 Clinical 2 .71 .72 6.2

Fig. 7 ADI-R SROC plot by setting. Note: size of shape indicates sample size

Journal of Autism and Developmental Disorders

1 3

these measures function in clinical settings, articles utiliz- ing alternative and less disseminated diagnostic algorithms were not included in the meta-analysis. Wider clinical use of these types of algorithms may be beneficial in improving the diagnostic accuracy of the ADI-R. This distinction further emphasizes the importance of clinical expertise in accurate differential diagnosis of ASD from other non-ASD condi- tions and neurodevelopmental disorders which impact social communication (Maddox et al. 2017; Reaven et al. 2008).

Risk and Sources of Bias

The QUADAS-2 identified no studies with concerns of the applicability of the results to practice. This is likely due to the eligibility criteria of the studies included in the meta- analysis, which specified the types of measures and evalu- ations which were considered acceptable based on current clinical practice. However, nearly half of the articles had high risk of bias regarding patient selection, most often due to not enrolling a consecutive or random sample of par- ticipants in the study. Additionally, high or unclear levels regarding risk of bias of the reference standard were indi- cated for all studies. This is inherent in the nature of conduct- ing ASD evaluations, as the reference standard (i.e., outcome diagnosis) was almost always interpreted with knowledge of the index tests (e.g., ADOS-2), as is true in clinical practice. Clinicians making ASD diagnoses were therefore not blind to the results of the index tests; in fact, clinicians utilize the results of the index tests as part of the information used to make the final clinical diagnosis. Therefore, the reference standard is inherently influenced by the results of the index tests and cannot be interpreted separately. Although some research studies may consider utilizing techniques to miti- gate these concerns of bias, including having outside video reviewers or independent re-evaluations, this does not occur clinically. The accuracy of these measures in clinical prac- tice is predicated on the clinician administering and scor- ing the measures accurately, without outside confirmation. Given that a primary goal of this study was to investigate the utility of these measures in clinical practice, and the index test results and reference standard are inextricably linked in this type of evaluation, this bias is considered inherent in any comprehensive clinical evaluation for ASD.

Additionally, only peer-reviewed, published articles were considered for inclusion in this meta-analysis since the reli- ability of the information presented from other types of sources can be variable and difficult to determine. However, many other sources of potentially useful information were excluded. There is a clear publication bias within the inter- vention literature wherein studies with negative findings are often not accepted for publication, but this bias is less often observed in studies focused on diagnostics. There may be a publication bias regarding the level of training completed

by providers administering the ASD diagnostic measures. Two levels of training are available for the ADOS-2 and the ADI-R: clinical training, which is for professionals using the measure in clinical practice, and research training, which is designed for those who use the instrument for research. The clinical training is a prerequisite for the research train- ing. The majority of professionals utilizing the ADOS-2 and the ADI-R in clinical practice likely have completed the clinical training but have not attended the research training. However, peer-reviewed journals may favor publication of studies utilizing research-reliable clinicians. Therefore, the identified sensitivity and specificity in clinical settings in this study may overestimate the accuracy of these meas- ures when conducted by providers who only have completed the clinical training, but the effect of excluding non-peer- reviewed articles is not known.

An additional consideration is the decision to include only articles conducted in the United States, Canada, and the United Kingdom. Sociocultural and language factors are crucial to consider when conducting ASD evaluations. Although many language translations are available for both the ADOS-2 and the ADI-R, these measures were initially designed in English using Western sociocultural norms, and the vast majority of research and development of these meas- ures was conducted under similar parameters. Notably, the language of test administration was not reported for all but one of the studies included in this meta-analysis. Therefore, restricting the inclusion criteria to research conducted in the United States, Canada, and the United Kingdom was determined to be the best method available as a proxy to representing the sample for which these measures were ini- tially developed. It would be beneficial for future articles to directly specify the language in which the evaluations were conducted and the sociocultural background of the partici- pants and their families.

Conclusion

This systematic review and meta-analysis of the ADOS-2 and the ADI-R determined that the ADOS-2 is more accu- rate than the ADI-R. The ADOS-2 indicated high levels of sensitivity and specificity across settings, and it should be considered for any ASD evaluation. ASD diagnostic meas- ures may be less accurate in clinical compared to research settings, but more research utilizing solely clinical popula- tions is needed.

Acknowledgements The authors would like to thank the UAB Libraries Reference Department for their support in formalizing and improving the search strategy for this project, Sarah Ryan, Ph.D. and

Journal of Autism and Developmental Disorders

1 3

Cassandra Newsom, Ph.D. for serving as article reviewers, and Dustin Long, Ph.D. for assistance with biostatistical analyses. This research was supported in part by the Health Resources and Services Adminis- tration (HRSA) Maternal and Child Health Bureau (MCH) Leadership Education in Neurodevelopmental and Related Disabilities (LEND; PI: Biasini), UAB Civitan International Science Center and Foundation for Children with Intellectual and Developmental Disabilities McNulty Scientist Award (O’Kelley), and the UAB Civitan-Sparks Clinics.

Author Contributions JL designed the project and wrote the protocol supervised by SO. JL and MS conducted the literature searches and determined article eligibility. JL and CD completed data extraction, and JL conducted the statistical analysis. JL wrote the first draft of the manuscript and SO provided substantial edits and guidance. All authors have approved the final manuscript.

References

Review Manager (RevMan) [Computer program]. Version 5.3. Copen- hagen: The Nordic Cochrane Centre, The Cochrane Collaboration, 2014.

American Psychiatric Association (2000). Diagnostic and statistical manual of mental disorders (4th ed., Text Revision). Washington, DC: Author.

American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). Arlington, VA: Author.

Baird, G., Simonoff, E., Pickles, A., Chandler, S., Loucas, T., Meldrum, D., & Charman, T. (2006). Prevalence of disorders of the autism spectrum in a population cohort of children in South Thames: the Special Needs and Autism Project (SNAP). Lancet, 368, 210–15.

Camodeca, A. (2018). Utility of three N-Item scales of the child behavior checklist 6–18 in autism diagnosis. Research in Autism Spectrum Disorder, 51, 75–85. https ://doi.org/10.1016/j. rasd.2018.04.004.

Bishop, S. L., Huerta, M., Gotham, K., Havdahl, K. A., Pickles, A., Duncan, A., et  al. (2017). The Autism Symptom Interview, School-Age: A brief telephone interview to identify autism spec- trum disorders in 5-to-12-year-old children. Autism Research, 10(1), 78–88. https ://doi.org/10.1002/aur.1645.

Cox, A., Klein, K., Charman, T., Baird, G., Baron-Cohen, S., Swetten- ham, J., et al. (1999). Autism spectrum disorders at 20 and 42 months of age: stability of clinical and ADI-R diagnosis. Journal of Child Psychology and Psychiatry, and Allied Disciplines, 40(5), 719–32.

De Bildt, A., Sytema, S., Ketelaars, C., Kraijer, D., Mulder, E., Volk- mar, F., & Minderaa, R. (2004). Interrelationship between autism diagnostic observation schedule-generic (ADOS-G), autism diag- nostic interview-revised (ADI-R), and the diagnostic and statis- tical manual of mental disorders (DSM-IV-TR) classification in children and adolescents with mental retardation. Journal of Autism and Developmental Disorders, 34(2), 129–137.

De Bildt, A., Sytema, S., van Lang, N. D. J., Minderaa, R. B., van Engeland, H., & de Jonge, M. V. (2009). Evaluation of the ADOS revised algorithm: the applicability in 558 Dutch children and adolescents. Journal of Autism and Developmental Disorders, 39(9), 1350–8. https ://doi.org/10.1007/s1080 3-009-0749-9.

Deeks, J. J. (2001). Systematic reviews of evaluations of diagnostic and screening tests. British Medical Journal, 323, 157–162.

Deeks J. J., Wisniewski S., & Davenport C. (2013). Cochrane Hand- book for Systematic Reviews of Diagnostic Test Accuracy Version

1.0.0. The Cochrane Collaboration, 2013. http://srdta .cochr ane. org/.

Dereu, M., Roeyers, H., Raymaekers, R., Meirsschaut, M., & Warreyn, P. (2012). How useful are screening instruments for toddlers to predict outcome at age 4? General development, language skills, and symptom severity in children with a false positive screen for autism spectrum disorder. European Child Adolesc Psychiatry, 21(10), 541–551.

DiLavore, P. C., Lord, C., & Rutter, M. (1995). The pre-linguistic autism diagnostic observation schedule. Journal of Autism and Developmental Disorders, 25(4), 355–379.

Dorlack, T. P., Myers, O. B., & Kodituwakku, P. W. (2018). A com- parative analysis of the ADOS-G and ADOS-2 algorithms: preliminary findings. Journal of Autism and Developmental Disorders, 1–12.

Dykens, E. M., Roof, E., Hunt-Hawkins, J., Dankner, N., Lee, E. B., Shivers, C. M., et al. (2017). Diagnoses and characteristics of autism spectrum disorders in children with Prader-Willi syn- drome. Journal of Neurodevelopmental Disorders, 9(18), 1–12. https ://doi.org/10.1186/s1168 9-017-9200-2.

Falkmer, T., Anderson, K., Falkmer, M., & Horlin, C. (2013). Diag- nostic procedures in autism spectrum disorders: A systematic literature review. European Child & Adolescent Psychiatry, 22(6), 329–40. https ://doi.org/10.1007/s0078 7-013-0375-0.

Gilchrist, A., Green, J., Cox, A., Burton, D., Rutter, M., & Le Couteur, A. (2001). Development and current functioning in adolescents with Asperger syndrome: a comparative study. Journal of Child Psychology and Psychiatry, and Allied Disciplines, 42(2), 227–40.

Gillentine, M. A., Berry, L. N., Goin-Kochel, R. P., Ali, M. A., Ge, J., Guffey, D., et al. (2017). The cognitive and behavioral phe- notypes of individuals with CHRNA7 duplications. Journal of Autism and Developmental Disorders, 47(3), 549–562. https :// doi.org/10.1007/s1080 3-016-2961-8.

Gotham, K., Risi, S., Dawson, G., Tager-Flusberg, H., Joseph, R., Carter, A., et al. (2008). A Replication of the autism diagnos- tic observation schedule (ADOS) revised algorithms. Journal of the American Academy of Child & Adolescent Psychiatry, 47(6), 642–651. https ://doi.org/10.1097/CHI.0b013 e3181 6bffb 7.

Gotham, K., Risi, S., Pickles, A., & Lord, C. (2007). The Autism Diag- nostic Observation Schedule: Revised algorithms for improved diagnostic validity. Journal of Autism and Developmental Dis- orders, 37(4), 613.

Gray, K. M., Tonge, B. J., & Sweeney, D. J. (2008). Using the Autism Diagnostic Interview-Revised and the Autism Diagnostic Obser- vation Schedule with young children with developmental delay: evaluating diagnostic validity. Journal of Autism and Develop- mental Disorders, 38(4), 657–667.

Grzadzinski, R., Dick, C., Lord, C., & Bishop, S. (2016). Parent- reported and clinician-observed autism spectrum disorder (ASD) symptoms in children with attention deficit/hyperactivity disor- der (ADHD): implications for practice under DSM-5. Molecular Autism, 7(7), 1–12. https ://doi.org/10.1186/s1322 9-016-0072-1.

Guthrie, W., Swineford, L. B., Nottke, C., & Wetherby, A. M. (2013). Early diagnosis of autism spectrum disorder: stability and change in clinical diagnosis and symptom presentation. Journal of Child Psychiatry, 54(5), 582–590. https ://doi.org/10.1111/jcpp.12008 .

Harris, S. W., Hess, D., Goodlin-Jones, B., Ferranti, J., Bacal- man, S., Barbato, I., et  al. (2008). Autism profiles of males with fragile X syndrome. American Journal on Intellectual and Developmental Disabilities, 113(6), 427–438. https ://doi. org/10.1352/2008.113:427-438.

Havdahl, K. A., von Tetzchner, S., Huerta, M., Lord, C., & Bishop, S. L. (2016). Utility of the child behavior checklist as a screener for autism spectrum disorder. Autism Research, 9(1), 33–42. https :// doi.org/10.1002/aur.1515.

Journal of Autism and Developmental Disorders

1 3

Howes, O. D., Rogdaki, M., Findon, J. L., Wichers, R. H., Charman, T., King, B. H., et al. (2017). Autism spectrum disorder: Consensus guidelines on assessment, treatment and research from the British Association for Psychopharmacology. Journal of Psychopharma- cology, 32(1), 3–29. https ://doi.org/10.1177/02698 81117 74176 6.

Kamp-Becker, I., Albertowski, K., Becker, J., Ghahreman, M., Lang- mann, A., Mingebach, T., Poustka, L., Weber, L., Schmidt, H., Smidt, J., Stehr, T., Roessner, V., Kucharczyk, K., Wolff, N., & Stroth, S., (2018). Diagnostic accuracy of the ADOS and ADOS-2 in clinical practice. European Child \& Adolescent Psychiatry, 1–15.

Kim, S. H., & Lord, C. (2012). Combining information from multi- ple sources for the diagnosis of autism spectrum disorders for toddlers and young preschoolers from 12 to 47 months of age. Journal of Child Psychology and Psychiatry, 53(2), 143–151.

Langmann, A., Becker, J., Poustka, L., Becker, K., & Kamp-Becker, I. (2017). Diagnostic utility of the autism diagnostic observa- tion schedule in a clinical sample of adolescents and adults. Research in Autism Spectrum Disorders, 34, 34–43.

Le Couteur, A., Haden, G., Hammal, D., & McConachie, H. (2008). Diagnosing autism spectrum disorders in pre-school children using two standardised assessment instruments: The ADI-R and the ADOS. Journal of Autism and Developmental Disorders, 38, 362–372. https ://doi.org/10.1007/s1080 3-007-0403-3.

Lord, C., Risi, S., Lambrecht, L., Cook, E. H., Leventhal, B. L., DiLavore, P. C., et al. (2000). The autism diagnostic observation schedule – generic: A standard measure of social and communi- cation deficits associated with the spectrum of autism. Journal of Autism and Developmental Disorders, 30(3), 205–223.

Lord, C., Rutter, M., DiLavore, P. C., Risi, S., Gotham, K., & Bishop, S. (2012). Autism diagnostic observation schedule: ADOS-2. Los Angeles, CA: Western Psychological Services.

Lord, C., Rutter, M., DiLavore, P. C., Risi, S., Gotham, K., & Bishop, S. (2012b). ADOS-2. Autism Diagnostic Observation Schedule. Manual (Part I): Modules 1-4. Western Psychological Services Los Angeles, CA.

Luyster, R., Gotham, K., Guthrie, W., Coffing, M., Petrak, R., Pierce, K., et al. (2009). The Autism diagnostic observation schedule – toddler module: A new module of a standardized diagnostic measure for autism spectrum disorders. Journal of Autism and Developmental Disorders, 39(9), 1305–1320.

Maddox, B. B., Brodkin, E. S., Calkins, M. E., Shea, K., Mullan, K., Hostager, J., et  al. (2017). The accuracy of the ADOS-2 in identifying autism among adults with complex psychiatric conditions. Journal of Autism and Developmental Disorders, 47(9), 2703–2709. https ://doi.org/10.1007/s1080 3-017-3188-z.

Mazefsky, C., & Oswald, D. P. (2006). The discriminative ability and diagnostic utility of the ADOS–G, ADI–R, and GARS for children in a clinical setting. Autism, 10(6), 533–549. https :// doi.org/10.1177/13623 61306 06850 5.

McInnes, M. D., Moher, D., Thombs, B. D., McGrath, T. A., Bossuyt, P. M., Clifford, T., et al. (2018). Preferred reporting items for a systematic review and meta-analysis of diagnos- tic test accuracy studies: the PRISMA-DTA statement. JAMA, 319(4), 388–396.

Molloy, C., Murray, D. S., Akers, R., Mitchell, T., & Manning- Courtney, P. (2011). Use of the autism diagnostic observation schedule (ADOS) in a clinical setting. Autism, 15(2), 143–162. https ://doi.org/10.1177/13623 61310 37924 1.

Neuhaus, E., Beauchaine, T. P., Bernier, R. A., & Webb, S. J. (2017). Child and family characteristics moderate agreement between caregiver and clinician report of autism symptoms. Autism Research, 11(3), 476–487.

Oosterling, I. J., Roos, S., de Bildt, A., Rommelse, N., de Jonge, M., Visser, J., et al. (2010). Improved diagnostic validity of the ADOS revised algorithms: a replication study in an independent

sample. Journal of Autism and Developmental Disorders, 40(6), 689–703. https ://doi.org/10.1007/s1080 3-009-0915-0.

Papanikolaou, K., Paliokosta, E., Houliaras, G., Vgenopoulou, S., Giouroukou, E., Pehlivanidis, A., et al. (2009). Using the autism diagnostic interview-revised and the autism diagnostic obser- vation schedule-generic for the diagnosis of autism spectrum disorders in a Greek sample with a wide range of intellectual abilities. Journal of Autism and Developmental Disorders, 39(3), 414–420.

Penner, M., Anagnostou, E., Andoni, L. Y., & Ungar, W. J. (2017). Systematic review of clinical guidance documents for autism spectrum disorder diagnostic assessment in select regions. Autism, 22(5), 517–527.

Reaven, J. A., Hepburn, S. L., & Ross, R. G. (2008). Use of the ADOS and ADI-R in children with psychosis: Importance of clinical judgment. Clinical Child Psychology and Psychiatry, 13(1), 81–94. https ://doi.org/10.1177/13591 04507 08634 3.

Risi, S., Lord, C., Gotham, K., Corsello, C., Chrysler, C., Szatmari, P., et al. (2006). Combining information from multiple sources in the diagnosis of autism spectrum disorders. Journal of the American Academy of Child and Adolescent Psychiatry, 45(9), 1094–103.

Rutter, C. M. (1995). Regression methods for meta-analysis of diag- nostic test data. Academic Radiology, 2, S48–S56.

Rutter, C. M., & Gatsonis, C. A. (2001). A hierarchical regression approach to meta-analysis of diagnostic test accuracy evalua- tions. Statistics in Medicine, 20(19), 2865–2884.

Rutter, M., Le Couteur, A., Lord, C., et al. (2003). Autism diagnos- tic interview-revised. Los Angeles, CA: Western Psychological Services, 29, 30.

Stewart, J. R., Vigil, D. C., Ryst, E., & Yang, W. (2014). Refin- ing best practices for the diagnosis of autism: A comparison between individual healthcare practitioner diagnosis and trans- disciplinary assessment. Nevada Journal of Public Health, 11(1), 1.

Takwoingi, Y. & Deeks, J. (2010). MetaDAS: A SAS macro for meta- analysis of diagnostic accuracy studies. User Guide Version 1.3. 2010 July. http://srdta .cochr ane.org/.

Tomanik, S. S., Pearson, D. A., Loveland, K. A., Lane, D. M., & Shaw, J. B. (2007). Improving the reliability of autism diag- noses: Examining the utility of adaptive behavior. Journal of Autism and Developmental Disorders, 37(5), 921–928.

Ventola, P. E., Kleinman, J., Pandey, J., Barton, M., Allen, S., Green, J., et al. (2006). Agreement among four diagnostic instruments for autism spectrum disorders in toddlers. Journal of Autism and Developmental Disorders, 36(7), 839–47.

Vllasaliu, L., Jensen, K., Hoss, S., Landenberger, M., Menze, M., Schütz, M., et  al. (2016). Diagnostic instruments for autism spectrum disorder (ASD). The Cochrane Library.

Whiting, P. F., Rutjes, A. W. S., Westwood, M. E., Mallett, S., Deeks, J. J., Reitsma, J. B., et al. (2011). QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy stud- ies. Annals of Internal Medicine, 155(8), 529–536. https ://doi. org/10.7326/0003-4819-155-8-20111 0180-00009 .

Wiggins, L. D., Reynolds, A., Rice, C. E., Moody, E. J., Bernal, P., Blaskey, L., et  al. (2015). Using standardized diagnostic instruments to classify children with autism in the Study to Explore Early Development. Journal of Autism and Develop- mental Disorders, 45, 1271–1280. https ://doi.org/10.1007/s1080 3-014-2287-3.

Wiggins, L. D., & Robins, D. L. (2008). Brief Report: excluding the ADI-R behavioral domain improves diagnostic agreement in toddlers. Journal of Autism and Developmental Disorders, 38, 972–976. https ://doi.org/10.1007/s1080 3-007-0456-3.

Zander, E., Sturm, H., & Bӧlte, S. (2015). The added value of the combined use of the autism diagnostic interview-revised and the

Journal of Autism and Developmental Disorders

1 3

autism diagnostic observation schedule: Diagnostic validity in a clinical Swedish sample of toddlers and young preschoolers. Autism, 19(2), 187–199.

Zander, E., Willfors, C., Berggren, S., Choque-Olsson, N., Coco, C., Elmund, A., et al. (2016). The objectivity of the Autism Diag- nostic Observation Schedule (ADOS) in naturalistic clinical set- tings. European Child Adolescent Psychiatry, 25(7), 769–780.

Zander, E., Willfors, C., Berggren, S., Coco, C., Holm, A., Jifält, I., et al. (2017). The interrater reliability of the autism diagnostic interview-revised (ADI-R) in clinical settings. Psychopathol- ogy, 50(3), 219–227.

Ziats, M. N., Goin-Kochel, R. P., Berry, L. N., Ali, M., Ge, J., Guffey, D., et al. (2016). Genetics in Medicine, 18(11), 1111– 1118. https ://doi.org/10.1038/gim.2016.9.

Zwaigenbaum, L., Bryson, S. E., Brian, J., Smith, I. M., Roberts, W., Szatmari, P., et al. (2016). Stability of diagnostic assessment for autism spectrum disorder between 18 and 36 months in a high- risk cohort. Autism Research, 9, 790–800. https ://doi.org/10.1002/ aur.1.

Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

  • Systematic Review and Meta-Analysis of the Clinical Utility of the ADOS-2 and the ADI-R in Diagnosing Autism Spectrum Disorders in Children
    • Abstract
    • Introduction
    • Methods
      • Measures for Index Tests
        • Autism Diagnostic Observation Schedule, Second Edition (ADOS-2; Lord et al. 2012a)
        • Autism Diagnostic Interview, Revised (ADI-R; Rutter et al. 2003).
      • Eligibility Criteria
        • Reference Standard for Diagnosis
        • Study Design
      • Search Strategy
      • Assessment of Methodological Quality
      • Study Selection
      • Data Extraction
      • Data Analysis
        • Outliers and Sensitivity Analysis
    • Results
      • Quality of the Included Studies
      • Diagnostic Accuracy of Measures
        • ADOS-2
        • ADI-R
    • Discussion
      • Sources of Heterogeneity
      • Risk and Sources of Bias
    • Conclusion
    • Acknowledgements
    • References