do summary for this article one page?
Epidemiologic Reviews Copyright © 1999 by The Johns Hopkins University School of Hygiene and Public Health All rights reserved
Vol.21, No. 2 Printed in U.S.A.
Spatial Analytical Methods and Geographic Information Systems: Use in Health Research and Epidemiology
Dale A. Moore and Tim E. Carpenter
INTRODUCTION
Person, place, time: these are the basic elements of outbreak investigations and epidemiology. Historically, however, the focus in epidemiologic research has been on person and time, with little regard for the implica- tions of place or space even though disease mapping has been done for over a hundred years. The develop- ment of geographic information systems (GISs) over the last 20 years has provided a more powerful and rapid ability to examine spatial patterns and processes. This, in turn, has fostered the discussion of such policy- relevant issues as health services and planning (1), as well as the use of GISs for epidemiologic investigations and disease surveillance.
Methods of spatial analysis are given little, if any, introduction in modern epidemiology texts, and few epi- demiologists have ventured further than making dot or choropleth maps, or listing geographic units along with rates of disease in tabular form (2-6). Spatial patterns are frequently intricate and complex, while most spatial methods used by epidemiologists can capture or identify only gross, simplistic patterns (2). Epidemiologists understand that disease processes have an historical (time) component, and formal methods of time series and hazard analyses are well-developed to study them. Few, however, recognize that every epidemic also has a geography (space). Some epidemiologists may not be aware that evaluation of the spatial distribution of mea- sures of disease risk may provide etiologic insight.
The logic of using geography to study disease or health care is derived from appreciation of factors caus- ing non-uniformity of disease distribution (2). These fac-
Received for publication May 20, 1998, and accepted for publica- tion July 6, 1999.
Abbreviations: AIDS, acquired immunodeficiency syndrome; GAM, geographic analysis machine; GIS, geographic information system; HIV, human immunodeficiency virus.
From the Veterinary Medical Teaching and Research Center, Department of Population Health and Reproduction (Moore), and the Department of Medicine and Epidemiology (Carpenter), School of Veterinary Medicine, University of California, Davis, Davis, CA.
Reprint requests to Dale A. Moore, VMTRC, University of California, Davis, School of Veterinary Medicine, 18830 Road 112, Tulare CA, 93274.
tors include physical and environmental factors; social, economic and cultural factors; and genetic factors. For example, diseases may be associated with environmen- tal pollution, linked to individual or group behaviors, or associated with a genetic predisposition. In turn, all of these factors may have spatial distributions influencing the extent and intensity of a particular disease.
GISs, when combined with spatial analytical meth- ods, may be helpful in the study of health care and health care delivery (7). The literature on the geogra- phy of health care can be divided into three areas of research: the spatial properties of delivery systems and their accessibility, the concomitant utilization and planning of health care services, and the spatial struc- ture of disease patterns in both static and dynamic form (8).
In the last decade, computerization of spatial data, through the use of GISs, has emerged as a tool for health care research and epidemiology (7). No single review to date has covered both spatial analytical tech- niques and modern geographic information systems (2, 7, 9, 10). The purpose of this presentation is to review both topics by highlighting some notable classic stud- ies and some more recent examples of spatial analyti- cal methods and geographic information systems in human and animal health research. Several kinds of analytical techniques will be described, as well as the role GISs have had in improving understanding of the spatial aspects of health care and disease research.
HISTORICAL PERSPECTIVE
Geography is concerned with the identification and explanation of spatial structure, pattern, and process, and with the analysis and explanation of the links between humans and the environment (2). Since epi- demiology is the study of the distribution and determi- nants of diseases and injuries in populations, and is concerned with the frequency and type of illnesses or injuries in groups of people and factors that influence their distribution, geography has a logical fit in many epidemiologic studies (4). For example, John Snow demonstrated the importance of the geography of dis-
143
144 Moore and Carpenter
ease events. When Snow mapped cholera cases in 1854 for the city of London, nebula-like spatial clus- ters, with distance-decay effects, were readily appar- ent. His maps led to the hypothesis that one particular water supply was the source of the outbreak. Without knowing the bacterial cause or means of transmission of cholera in the mid nineteenth century, he was able to quell the outbreak once he understood the spatial aspects (3, 11).
Another classic and important spatial investigation was that of Burkitt's lymphoma in Africa. On a "tumour safari" in east and central Africa in the 1950s, Denis Burkitt recognized that an unusual jaw tumor of children was limited in its distribution to a particular equatorial region and altitude range. On this basis, he suspected a vector-borne viral disease. Although the disease was not a vector-borne one, his investigations led to the discovery of the first tumor identified to be caused by a virus (Epstein-Barr virus). He was par- tially correct about the vector-borne aspect because the tumor is apparently the result of synergism between Epstein-Barr virus and malaria (12, 13).
THE NATURE OF SPATIAL DATA
In both of the above classic studies, geography, or location, was important in understanding the etiology of the disease. However, to use spatial methods, one must understand the nature of spatial data (14, 15). One unique aspect of spatial data is that the spatial component is based on two continuous dimensions, one in the horizontal direction (easting) and one in the vertical direction (northing) (14). Another aspect is the problem of spatial dependence, analogous to temporal dependence. Nearby locations are likely to possess similar attributes; or in other words, everything is related to everything else and near things are more related than distant things (16). These features of spa- tial data create needs for special analytical techniques and should be considered every time a project involv- ing geography (i.e., location) is attempted.
The primary issue in using geographic and spatial analytical techniques in epidemiology and health research is one of recognizing the spatial structure of a process, whether it is a cluster of health events or a spatial pattern of disease diffusion over time. The con- nections among people, among animals, and the way in which such interrelation are embedded in a complex of environmental variables may well define or struc- ture the space. A particular spatial structure includes the individuals affected and how they are connected in communities, as well as the dynamics of these com- munities and their organization into larger units. The geography of a disease can give valuable clues into understanding of how culture, environment, and
behavior interact with health and disease. The kinds of spatial issues that might concern health researchers include spatial patterns of morbidity and mortality; factors associated with these patterns; disease diffu- sion and disease etiology; spatial distribution, location, and regionalization of health care resources; access to, and utilization of, resources and factors related to resource distribution; spatial aspects of the interaction between disease and access to health care; and risk assessment for environmental toxicology (7).
SPATIAL ANALYTICAL METHODS
The study of spatially-related objects or characteris- tics can be divided into the description of locational characteristics that differentiate areas (exploratory analytical techniques) and the analysis of spatial inter- relations (explanatory analytical techniques) (17). Some common spatial techniques used in health research include disease mapping, clustering tech- niques, diffusion studies, identification of risk factors through map comparisons and regression analysis (7). In this presentation of spatial analytical techniques, we will briefly cover some important considerations and reviews available for disease mapping, a number of techniques for detection of disease clusters for both point and areal data, techniques for "relative spaces", aspects of diffusion studies, techniques for interpolat- ing and smoothing spatial data, and some studies and techniques for identification of spatial risk factors. Table 1 provides a list of some of the methods dis- cussed in this paper including cluster detection, dis- persion methods, and interpolation techniques; the advantages, disadvantages and use of each; and some computer software (where applicable) for some of the techniques covered in the text.
Disease mapping
Mapping disease includes mapping point locations of cases, incidence rates by area, and standardized rates. Although disease rates are routinely mapped, the representation of diseases on maps varies depending on the investigators' objectives. How the map is devel- oped will dictate what information can be gleaned from it. Sources of information about, and examples of, mapping disease and representation of disease events are provided in atlases by Cliff et al. (18), Smallman-Raynor et al. (19), and Pickle et al. (20).
The map is a good communication device, but can also be misleading (15, 21). The eye can detect pat- terns in noisy data displayed on maps, but maps are not good at representing complex relations between response and explanatory variables (21). For example, figure 1 is a cumulative dot-density map of raccoon
Epidemiol Rev Vol. 2 1 , No. 2, 1999
Spatial Analytical Methods/Geographic Information 145
rabies cases in Pennsylvania from 1982 to 1996 (22). Although a pattern of disease diffusion is visually apparent, spatial analytical techniques are necessary to understand the complex nature of the spatial trend.
Detection of clusters
A cluster, in epidemiology, is a number of health events situated close together in space and/or time. Although there are those who warn about devoting much energy to the detection of clusters, much of the literature in spatial epidemiology deals with this topic (23). Several reviews of cluster detection techniques are available (24-32). Cuzick and Edwards (33) describe three general methodological approaches for detecting clustering: methods based on cell counts, on autocorrelative adjacencies of cells with high counts, and on distance between events. In this review, we will highlight techniques for detecting clusters in point data and areal data.
Point patterns. Clusters are commonly identified from point-pattern data (34). The objective of examin- ing point patterns is to recognize when events are sys- tematically organized or structured compared with events distributed at random (35). The simplest point- pattern analysis is visual inspection of a dot map which displays a geographic distribution of events (figure 1). Statistical tools, such as variance to mean ratio, may be needed to evaluate subjective impressions. The usual null hypothesis in spatial analysis of disease data is that the number of disease events in a given area is proportional to the number at risk (28). In the absence of clustering, points are either randomly or uniformly distributed in space (33). The analysis usually assumes a homogeneous Poisson process over the study region, implying independence of the spatial events.
Nearest neighbor analysis uses inter-event distances to develop impressions of the strength of the clustering of point data (35). Significance tests, using distance to nearest neighbors, test for complete spatial random- ness. This test may be used as a preliminary procedure in an analysis of event data and describes the geo- graphic distribution of a set of points according to their spacing (36). Nearest neighbor analysis has been severely criticized, however, because it fails to distin- guish between homogeneous and random patterns and because different results will be obtained if different sized areas are analyzed with the same data (36). Thus, the scale of the area plays a critical role in identifying clusters.
The occurrence of more than a single process may occur; for example, when pairs of points tend to cou- ple together while another process may or may not be present. The resulting interpretation may miss the pair-
wise, or higher order, clustering. This could be avoided by examining, in addition to the nearest neighbor, dis- tances between the first closest points and the second, or higher order, closest point. The calculation of the expected distance to the kth nearest neighbor was derived by Thompson (37).
In order to adjust for a non-homogeneous, or non- Poisson distribution of the population at risk, a number of alternatives to the nearest neighbor method have been used. Bithell (38) estimated a relative-risk func- tion for childhood leukemia in Cumbria, England, comparing first and second nearest neighbor distances for cases and randomly selected controls. A similar approach was taken by Gatrell and Bailey (39) who estimated k functions for randomly selected cases and controls of childhood leukemia in Lancashire, England. They examined the difference plots of these two functions against distance. Significant clustering was identified when peaks exceeded an analytical or simulated confidence interval. Glaser (40) used two alternative techniques, one which adjusted for popula- tion density algebraically (41) while the second pro- duced a transformed map where the population was uniformly distributed (42) to examine clustering of Hodgkin's disease in the San Francisco Bay area. A common approach to control for the distribution of the population at risk was developed by Cuzick and Edwards (33). They used a variation of the kth nearest neighbor approach whereby cases are examined with respect to their number of nearest neighbors that are also cases. The expected number of case nearest neigh- bors is based on the number and proportion of cases in the case-control selection. In contrast to the traditional nearest neighbor test, the Cuzick-Edwards test consid- ers the relative and not the actual distance between points. All these techniques control for the bias toward clustering one finds with the basic nearest neighbor test.
Quadrat analysis (or cell count method) is another method to test for complete spatial randomness and addresses the issue of point density (35, 36). A grid, typically a square or circle, is placed, either randomly or in a fixed sequence, over a map, and the number of events is counted in each grid. For example, in figure 2, a random allocation of 190 points is shown distrib- uted among 100 cells of a 10 x 10 matrix. Given the mean number of points per grid (X) is 1.9, the random- ness of that distribution may be tested by assuming the points follow an underlying Poisson distribution, using a chi-square test comparing the observed versus expected frequency distribution of cell counts. In this example, there are 14 cells containing no points. Given the mean of 1.9 and assuming the points followed a Poisson distribution, the expected probability of a cell
Epidemiol Rev Vol. 2 1 , No. 2, 1999
TABLE 1. Select statistical techniques and computer software for identifying clusters, diffusion patterns, and methods of interpolation
Test/format Date type Primary strength/use Primary limitation Software Use in epidemiology/health
Area Joint counts
Ohno
Gerry's c
Moran's /
Poisson
Point NN*
MhNN
Geary's c
Moran's /
Cuzick-Edwards
K-function
Geographic analysis machine
Dichotomous
Rank or categorical
Continuous
Continuous
Intergers
Case points
Same as NN
Continuous
Continuous
Case-control
Case-control
Dichotomous
o 5; (b 3
5) 5 o. ro
p
!°1999
Spatial scan
Line analysis
Trend surface
Spatial adaptive filtering
Expansion method
Case data
Simulation
Emperical data; time to first report
Emperical data and simulation
Simulation
Clustering
Identification of large-scale clusters
Identification of large-scale or local area clusters
Identification of large-scale clusters
Identification of large-scale clusters
Identification of large-scale or local area clusters Ignores adjacencies
Low power
Low power; ignores adjacencies
May ignore close, nonadjacent areas
May ignore close, nonadjacent areas
Identification of local area clusters sensitive to area selection and heterogenous population distribution
Identification of local area clusters
Identification of local area clusters
Identification of local area clusters
Identification of local area clusters, adjusts for nonrandom distribution of population at risk
Identification of local area clusters, adjusts for population distribution
Adjusts for population distribution
Adjusts for population distribution; identifies primary and secondary clusters
Ignores underlying distribution of population at risk, ignores multiple (>2) case clusters
Ignores underlying, distribution of population risk
SpaceStat(111), TSpStat(131)
CLUSTER (132), TSpStat
SpaceStat
CAST (133), Space-Stat, Statl (134), TSpStat
CLUSTER, TSpStat
CAST, Statl, TSpStat
CAST, SpaceStat, Statl, TSpStat
Ignores underlying distribution SpaceStat of population at risk
Ignores underlying distribution of population at risk
Ignores temporal occurrence of events
May ignore temporal occurrence of events
Computer intensive; statistical properties
Computer intensive
CAST, SpaceStat, Statl, TSpStat
CAST, Statl, TSpStat
SPLANCS (S-PLUS) (135)
Geographic analysis machine (46)
SaTScan (136)
Dispersion (diffusion)
Compares disease fronts to random walk to detect a pattern of movement; can get rate and pattern spread (7, 72)
Can model time to first occurrence, model pattern, and rate of spread (86)
Used to forcast case-numbers using population as predictor (71, 137)
Model diffusion in space and time
Requires programming
Does not account for spatial autocorrelation
Requires programming
ArcView Special Analyst (99)
Need understanding of parameter relations (88)
Clustering of endodontic offices (50)
Areal cancer mortality (52)
Bone cancer mortality (61)
Cancer (52, 62), stroke (63), Lyme disease (64)
Heart disease (43), lupus, (44), asthma (45)
Leukemia/lymphoma (33)
Leukemia (39), Hodgkin's disease (41)
Leukemia (46)
Breast cancer (48), leukemia (47)
Simulation of rabies diffusion (72)
Variola minor (84), cholera (85), rabies (86)
Forecasting AIDS* cases (71)
Spatial Analytical Methods/Geographic Information 147
•= in £j ^
I i l l li
sat
S
I 8 1
s!
I c
.§
C
o
§
2?
E
CD a.
gj.
1 E
f a ^ 3 CO
ft
co o a § a co
II | 1 I
© O )
S »
.E
Is o c
if I
A,1982
C, 1989
D,1996
FIGURE 1. Cumulative dot density maps of cases of racoon rabies in Pennsylvania, 1982-1996. (Reprinted with permission from: Moore DA. Spatial diffusion of raccoon rabies in Pennsylvania, USA. Prev Vet Med 1999;40:19-32).)
Epidemiol Rev Vol. 2 1 , No. 2, 1999
148 Moore and Carpenter
100 90 80
P 70
» 60 | 50 S 40 c 30
20 10 0
•
*
•
•
•
• ,
• •
• *
* •
* 4
* *
4>
»
• •
•v • •
«
• W
4>
• *
« •
• •
• «
• •
• •
«
ft 4
» • •
• • • •
•
•
•
•
• • 4>
•
*
•
• •
0 10 20 30 40 50 60 70 80 90 100 easting (X)
FIGURE 2. Hypothetical Poisson distribution of points in grid to illustrate a quadrat analysis {X = 1.9).
containing no points is 0.1496. Therefore, the expected number of cells containing no points is 14.96. Expected cell counts are determined for the remaining point frequencies and a chi-square value calculated to determine the deviation from randomness demon- strated by this approach. In this example, the chi- square value was calculated to be 2.841 with a/? value of 0.585. Epidemiologic applications of this approach include cluster investigations for standardized mortal- ity ratio for heart disease (43), systemic lupus erythe- matosus (44), and asthma mortality (45). An important issue is the choice of an appropriate grid shape and size. A shortcoming of this approach is that it fails to consider relative cell location.
Openshaw et al. (46) used a quadrat-type analysis in a geographic analysis machine (GAM) to detect clus- tering of childhood leukemia cases and graphically dis- play the results. They estimated the likelihood of find- ing a certain number of leukemia cases given the population at risk within the quadrat area. The objec- tive of this analytical tool was to detect clusters worthy of further investigation. Besag and Newell (25) pre- sented an alternative to the GAM to detect clusters of a rare disease over a large area, subdivided into smaller units. They examined each case in order to determine if its respective area centroid formed the center of a case cluster of predetermined size. The result was that each individual test focused on only the local structure of the pattern without an attempt to compensate for an apparent cluster once detected. This technique was used to identify apparent clusters of acute lymphoblas- tic leukemias diagnosed in children under 15 years of age in several authority districts in England.
A similar technique, the spatial scan test, iteratively searches for case clusters (26, 47). This method scans a large area, with a circular window, without previ- ously specifying window location or size. Once a clus- ter is identified, the test determines the significance, adjusting for the inherent problem of multiple testing. The scan test may identify secondary as well as pri- mary clusters and orders them according to their like- lihood ratios, and has been applied to breast cancer (48) and leukemia (47) cluster investigations.
Areal data. Frequently, spatial information is unavailable for point data, the data are grouped or summarized as area or regional data, or the focus is on identifying clusters on a larger scale or area. Several tests have been designed to identify areal clustering. While they traditionally evaluate the level of similarity of adjacent areas, they differ in the type of data they analyze: continuous, dichotomous, categorical.
For dichotomous, areal data, the degree of clustering or dispersion can be quantified by measuring the num- ber of total and dissimilar "joins" between areas, or join count method (49). A join is said to occur if two areas are adjacent to one another. A dissimilar join occurs if two adjacent areas are different, e.g., black versus white or above versus below the median. Clustered areas have relatively fewer dissimilar joins, dispersed have more, and random have an intermediate number of joins. The joins test was used to identify a pattern of spatial clustering of endodontic office loca- tions in the United States (50). This method, however, suffers from low power, presumably due to the loss of information compared with other area tests (51).
Epidemiol Rev Vol.21, No.2, 1999
Spatial Analytical Methods/Geographic Information 149
The Ohno method was developed to evaluate clus- tering of categorical or ranked areal data. The original application was for cancer mortality data in Japan (52). Areas are compared with respect to concordance (i.e., identical), adjacent areas are said to have concordant values if they have the same category or rank and are discordant otherwise. The number of adjacencies is counted and the number of concordant adjacencies cal- culated and compared with the expected number, based on the frequency distribution of each category. The level of significance of the difference between observed and expected concordant adjacencies is cal- culated for each category and tested using the chi- square test. A chi-square value may also be calculated for the entire distribution of categories. A potential shortcoming of the Ohno method is that all dissimilar joins are treated alike, i.e., adjacencies with ranks 1 and 2 are considered as different as those with ranks 1 and 5. If these relative differences are important, the Ohno method may not be suitable. A more appropriate test for rank data in adjacent areas evaluates what is referred to as the non-parametric rank adjacency sta- tistic, D. It is a measure of the average absolute differ- ence in ranks of all adjacent areas. Applications of this test have been made for detection of areal clustering of cancer (53-56).
Two techniques, Geary's c and Moran's I, are com- monly used to compare continuous data (57, 58). These techniques are similar in that they compare adja- cent area values in order to assess the level of large scale clustering. Clustering may be identified as the result of an unexpectedly large number of adjacent areas having either relatively large or small values. Jumars et al. (59) recommended the application of both, since autocorrelation (clustering) may be detected by one while missed by the other. Others show that tests based on Moran's / are consistently more powerful than those based on Geary's c (51, 54, 55). In addition, Walter (60) found that while Geary's c and Moran's / techniques may respond to localized clusters of high risk, they have negligible power in detecting highly localized hot spots. Although only limited use of Geary's test has been reported (61), Moran's test has been frequently applied to a variety of epidemiologic problems to examine areal clusters, including cancers (56, 62), stroke-mortality rates (63), and Lyme disease (64).
A review of the power associated with several of these adjacency clustering techniques was performed by Walter (55). He reported that Moran's / consistently had a higher power than Geary's c, which exceeded that of the rank adjacency test, D. The conclusion was that the power of the D test was severely limited to the nature of its nonparametric data. This limitation is
even more noticeable in the reduced power of the join count method which used dichotomous data (51).
Hungerford (49) demonstrated the use of second- order analysis with data on seroprevalence of anaplas- mosis in cattle in Illinois. The value at each point (in her study, county centroids) was compared with an expected value if all points and values were randomly distributed. Since the measured distance between points with similar values was smaller than expected, clustering was suggested for the data. Second order spatial analysis determines the degree of spatial depen- dence among variables as a function of distance between points or areas. This technique was used to analyze the spatial association of swine pseudorabies virus prevalence among counties in Illinois (65). The investigators studied clustering of county swine pseudorabies virus prevalence rates and compared these rates with geographic clustering of swine herds. Counties with high swine pseudorabies virus preva- lence rates clustered more than the observed clustering of counties with large numbers of swine herds.
Spatial autocorrelation analysis is an additional technique used to detect disease patterns. It is defined as the relation among values of a single variable that is attributable to the geographic arrangement of areal units on a map (36). A good introduction to spatial autocorrelation is given by Goodchild (66). Spatial autocorrelation is a measure of interdependence between values of a variable at different geographic locations and can be used to identify the degree of spa- tial clustering (64). Spatial correlograms are series of Moran's / statistics which can be evaluated at greater and greater distances from the areas to determine where spatial effects are maximized. This technique was used in a study of risk factors for anaplasmosis (49). A spatial correlogram is a function that shows the correlation among sample points (for some variable) separated by distance h. Correlation usually decreases with distance until it reaches or approaches zero. It describes the autocorrelation in a variable by comput- ing some index of covariance for a series of lag dis- tances (67).
In a study of geographic relations among county lung cancer mortality rates, Kennedy (68) was able to demonstrate the important influence of neighboring counties (local effects) on the lung cancer mortality rate for men and the more regional influence on the lung cancer mortality for women. First- to fifth-order neigh- bors mortality rates were weighted by the geographic relation with each county. Residual plots indicated that problems with autocorrelation were overcome by this autoregressive model.
Monte Carlo techniques are probabilistic methods and can be used for simulation purposes where spatial
Epidemiol Rev Vol. 2 1 , No. 2, 1999
150 Moore and Carpenter
data are not independent and areal units may be of dif- ferent sizes. Monte Carlo techniques are tools used to solve various problems by construction of some ran- dom process (69). Hierarchical clusters of "high risk" areas can be developed by ranking disease rates for spatial units from high to low. Adjacencies among high-ranking units are counted and can then be com- pared with the results of a Monte Carlo simulation which would establish the probabilities for the occur- rences of these adjacencies (70).
Relative spaces
The term "relative spaces" refers to events or factors that might be related in something other than simple, or untransformed, geographic space. These spaces may be communication space, commuter space, air passen- ger space, or any other space that appears relevant to the analysis (71-73).
Multidimensional scaling is a technique which has been used for problems with or without a traditional "geographic" issue. It can be used to identify relation- ships among individuals in two, three, or more dimen- sional spaces. The classic example of multidimen- sional scaling was done by Cliff et al. (74, 75) using measles outbreak data in Iceland and the United States. In addition to studies of infectious diseases, the tech- nique was used to identify important attributes by which people used to judge mental health facilities (76). Transformations of geographic space (distances between urban centers of varying population) have been used to simplify complex hierarchical diffusion processes using gravity model mapping (77).
The Markov process was used to model acquired immunodeficiency syndrome (AIDS) transmission in the New York, New York, metropolitan region (78). A Markov process is where individuals randomly move among a fixed set of states through a "transition". In Gould's study of AIDS, the probability of infection was related to "commuter space", or the amount of commuter traffic that could be carrying infected per- sons and their viral baggage to different boroughs and counties of the New York metropolitan region. From this procedure, it was shown that the structure of the AIDS space could be modeled in this "relative" space and that technology helped to shape the course of dif- fusion of the virus.
Diffusion studies
Diffusion can be visualized by creating a series of maps of disease or events. An example is Wallace's report on diffusion of tuberculosis in New York City (79) or figure 1. However, to address the complex spatial-temporal dynamics of a diffusion process, meth-
ods to model diffusion have been developed. Two gen- eral approaches have been used to model diffusion processes: stochastic and deterministic (80). A sto- chastic model has elements which include probability; deterministic models do not allow for chance. There are three types of diffusion processes: purely conta- gious, purely hierarchical (where the disease or health practice jumps from one place to another based on some hierarchy, such as population density), and mixed hierarchical. It is important to understand the spatial "backcloth" upon which a disease diffuses in order to model the process effectively. For a primer on spatial diffusion modeling and the elements that char- acterize diffusion phenomena, the reader is referred to Morrill et al. (80).
The primary theory of spatial diffusion was put forth by Torsten Hagerstrand in the 1950s (80). The method he developed to model spatial diffusion used a Monte Carlo technique to simulate the diffusion process. The first model assumed random adoption over space. The second model introduced the mean information field, a 5 x 5 grid providing the probabilities of adoption upon contact with an earlier adopter. The third model incor- porated resistance barriers to diffusion (80). The back- drop or surface on which the diffusion takes place could be the human or animal population density at any particular location on a grid map, and can have placed upon it geographic and other potential barriers to diffusion. The purpose of the model is to imitate or simulate patterns of diffusion.
Gilg (34) studied the 1970-1971 fowl pest epidemic in England, and identified diffusion of this infectious disease from the east to the west. For every data cell, he calculated time curves which highlighted the "local" epidemic. The time curves graphically demon- strated differences in the epidemic in different loca- tions, combining both space and time on one map.
Line analysis has been used to compare a disease's "front" of movement with a random walk. The actual direction of movement is compared with chance move- ments to detect any pattern (7). Vectors or lines which indicate magnitude and direction can be used to describe disease flow through an area. Following the spread of fox rabies westward from Poland after the second world war, a model for the spatial spread of rabies was developed in the United Kingdom (72). This model was a deterministic one which was applied to a rabies-free area. It predicted the wave speed of the disease and estimated the width of an intervention zone to prevent spread. The emphasis of the model was on the rate of spread and not necessarily the determi- nants or possible deterrents of disease spread.
A diffusion model for fox rabies was recently pro- posed by Jeltsch et al. (81). The model involved a grid
Epidemiol Rev Vol. 2 1 , No. 2, 1999
152 Moore and Carpenter
Spatial adaptive filtering is a technique which has been used to forecast the number of AIDS cases in dif- ferent geographic areas, thereby predicting the pattern of diffusion (87). A spatial filter is a functional expres- sion between some variable, such as the number of people with AIDS in a county, zip code area, or census tract, and some predictor variable, such as population (87). Since human geography is complex, the structure of human space, and, hence, AIDS space, is uncovered through this technique. Kabel (71) used the technique to forecast people with AIDS in county / using JS pop- ulation (71). The forecasts were averaged and used in a negative feedback expression to make better predic- tions. The adaptation comes when comparing fore- casted numbers with actual numbers in the counties.
Gould (73) used a technique to model diffusion of human immunodeficiency virus (HIV) in the United States using air passengers as a surrogate for the hier- archical relations between 102 major urban centers. A 102 x 102 air passenger origin-destination matrix was scaled by total numbers of passengers and represented the probability of interaction between each pair. HIV was "injected" into the system at a particular place, and at each time period or iteration the new results of the interactions were used in a probabilistic "smear- ing". When "injected" at New York, Los Angeles, California, Miami, Florida, and Houston, Texas, the correlation with the actual number of cases was very high. His conclusion was that spatial dynamics accounted for about 80 percent of the variation in loca- tion of HIV in these major urban centers (with about half the national population).
The expansion method has been used to demonstrate that parameters of a temporal equation for a disease epi- demic may themselves be functions of geographic vari- ables, such as population density, and by substitution or "expansion" of the original temporal equation, it can be used to model diffusion processes in time and space (87). The expansion method involves specifying a model for understanding relations, redefining some of the parameters of that initial model by expansion equa- tions into functions of variables, replacing the expanded parameters into the initial model, and using these expressions to capture the "drift" or variation over geo- graphic space (88). In other words, this method does not presuppose invariance of the model parameters in the temporal domain. Instead, the contextual variations of the relations are taken into account. Expansion method- ology can be applied to any form of model.
Interpolation and smoothing
Some widely-used spatial techniques do not test hypotheses but, instead, are used to interpolate new
data points, "smooth" data, or filter signals from noise. Filtering signals from noise is an important issue for epidemiologic data. Because of limitations of passive surveillance data, some means to find major geo- graphic or temporal signals in data can be challenging. Trend-surface analysis or some other filtering tech- nique may prove useful in these situations. The size of the filter and perhaps the geographic scale or resolu- tion of the data must be considered because of their effect on the detectable pattern in the diffusion process. Trend-surface analysis is one means of inter- polation (see discussion above). Other methods include kriging, splining, inverse weighted distance method, and construction of Thiessen polygons (89).
In a study of the distribution of antibodies to Chlamydia pneumonia (strain TWAR), a Finnish research group used trend-surface analysis as a smooth- ing technique to serve as a "filter" from which to extract signals from noise to find the regional trend of antibody prevalence (90). After the noise was filtered, residual differences were calculated. Residuals outside the con- fidence limits were considered to be locally important.
Kriging is a technique used to estimate point values by using surrounding, known point values (91). Kriging is a method of spatial prediction using a weighted moving average interpolation to produce the optimal spatial linear prediction (92). The weights reflect the distances between the location for which a value is being predicted and the locations with mea- sured values. It has been used in geostatistics as an interpolation method and is considered the best linear unbiased estimator of the characteristic under study where it best reflects the minimum mean square error. It minimizes the variance of the estimation errors. Kriging results in a marked smoothing effect with high original values tending to be underestimated and low values overestimated. The kriged values will be less variable than the original ones.
Kriging has been used in several epidemiologic studies. Peak weeks in rotavirus detections from US laboratories were interpolated using kriging and then mapped (93). Peak activity varied by location, and there were differences in variability of detection by location. Investigators found this technique useful for visualizing geographic and temporal trends. In another infectious disease study, the spatial and temporal dis- tribution of Anopheles gambiae mosquitoes in houses in a village in Ethiopia was monitored (94). Using kriging techniques, investigators demonstrated cluster- ing at the edges of the village and the changing pattern over time. The spatial patterns of infant mortality and birth defects rates in Des Moines, Iowa, were described by a contoured surface based on kriging (95). Kriging was also applied to data from a 6-week
Epidemiol Rev Vol. 2 1 , No. 2, 1999
Spatial Analytical Methods/Geographic Information 153
influenza-like illness outbreak in France (96) (figure 4). Estimated case numbers per medical practitioner were obtained through the procedure and mapped in a series of weekly maps. The maps, interpreted in suc- cession, depicted the spatial pattern and density of cases over time.
Environmental studies have used kriging estima- tions. Spatial distribution of air pollution in Prague (Czech Republic) was estimated by kriging and multi- ple regression modeling (97). Investigators found that school levels of nitrous oxide were positively related to symptoms of wheezing/whistling in children although home levels had a negative association. Kriging was also used in a study of the mid-Atlantic area of the United States to interpolate hourly ozone data from air quality monitoring stations (98). The advantages of the kriging process are that the predicted values are not constrained by the borders of the geo- graphic units, as in trend-surface analysis, and that the procedure can deal with problems of missing data.
Splining is a method to fit a minimum curvature sur- face through input points (67, 89, 99). The purpose is
to connect points with a smooth, continuous line if the data are strings of coordinate pairs. Splines are a large class of piecewise functions used to represent curves in two or three dimensions (67). This technique is not appropriate if there are large changes in a surface within a short horizontal distance because it can over- shoot estimated values. It is used widely in interpola- tion and is important in computer software to generate computer displays.
Other useful smoothing approaches include non- parametric and parametric techniques, Bayes' and empirical Bayes' methods. The Bayesian approach requires specification of the mean and variance of the disease risk, for example, to summarize prior beliefs about risk distribution (100). With the empirical Bayes' approach the values for the unknown parame- ters are estimated based on the observed data (101).
Risk factor identification
Areal patterns have been used to identify potential risk factors for disease incidence or prevalence. The
Year 1989
Week 48 Week 49 I 1 I I
200 400
Week 51 Week 52
Week 50
Week 01
FIGURE 4. Example of the use of kriging to interpolate weekly numbers of cases of influenza-like illness per practitioner in France. (Reprinted with permission: Carrat F, Valleron A. Epidemiologic mapping using the "Kriging" method: application to an influenza-like illness in France. Am J Epidemiol 1992;135:1293-300.)
Epidemiol Rev Vol.21, No. 2, 1999
154 Moore and Carpenter
most common areal-pattern displays are choropleth maps. Categories for the numbers of cases, rates of dis- ease, or standardized mortality ratios can be displayed by colors, patterns, or other areal features. These tech- niques are widely used in epidemiology and health care research. For areal data, the kappa statistic can be used to measure the degree of similarity between dis- tributions, and is corrected for areal overlap due to chance (49).
Synoptic mapping is a technique which produces contours by interpolating values for areas between data points and values at data points (102). Using this tech- nique to map wildlife rabies cases, Pool and Hacker (103) demonstrated coincident rabies identification in wildlife species within the seven biotic provinces of Texas (the explanatory variable). Time-series analyses with autocorrelation functions were calculated for each biotic province. Periodicity in skunk rabies dif- fered among the different provinces and could be related to habitat differences, thus giving some expla- nation for variation in rabies risk in different areas.
Spatial plots or "surfaces" developed from indepen- dent and dependent variables can be used for visual map comparisons (7). Correlation analysis, or "eco- logic correlation", is the most commonly used statisti- cal method of map comparison (7). Health care resource or disease rates for different spatial units can be compared using Pearson's product moment or Spearman's rank correlation statistics. Ecologic corre- lation was used in a study of the geographic distribu- tion of alcohol treatment facilities in Oklahoma (104). An index of service comprehensiveness was correlated with need, urbanization, income, and attitudes toward alcohol use by county. The coefficient of areal corre- spondence is the ratio of the area over which two phe- nomena are located together to the total area covered by the two phenomena. The method of areal corre- spondence is best used in a static system and not for a process which is diffusing. This technique has not been used to any extent in medical geography, epidemiol- ogy, or health care research (7).
Almost all variables available for a geographic risk factor analysis are likely surrogates for other variables not attainable from data or not yet identified (15). For risk factor analysis, particularly in an exploratory sense, these geographic variables may be used to iden- tify descriptive relations, serve as a basis for future research, and be used in the search for causal relations. Identification and classification of geographic corre- lates has been simplified by the development of geo- graphic information systems.
Techniques of general linear regression or hierarchi- cal modeling are other approaches to evaluating poten- tial associations in spatially-dependent data. Cook and
Pocock (105) used cardiovascular mortality data in a multiple regression model to examine potential explanatory variables. Because ordinary least squares regression models assume independent, uncorrelated residuals, using spatial data in these models is likely to violate that assumption due to spatial autocorrelation and may overstate the significance of the coefficients. Cook and Pocock illustrated a way that residuals from ordinary least squares regression can be used to suggest a suitable parameterization for a spatially-correlated error structure.
Harries (106) demonstrated the use of map overlay and a cluster analysis (as opposed to cluster detection) to classify Baltimore, Maryland, neighborhoods with respect to their levels of violence and social stressors. This was done to provide information to public health and other policy makers for making resource alloca- tion decisions. Twenty-four variables were used as "social stressors" for each of the 1,357 block groups in the city and surrounding area. Cluster analysis, a tech- nique used to combine observations that are more sim- ilar to each other, identified five clusters which were mapped by block group. The locations of these clusters indicated areas of highest poverty and violence.
In health care research, the expansion method was used to evaluate federal policies on the supply of physicians in the United States and their geographic distribution (107). The models took into account fac- tors in physicians' location decisions such as the pro- fessional climate, social amenities, and market factors, as well as the influence of federal policies. Location behavior was shown to be dominated by professional and social amenity factors, but as the supply of pri- mary care physicians increased, market forces became dominant.
USE OF GEOGRAPHIC INFORMATION SYSTEMS
A GIS is an integrated set of computer hardware and software tools to capture, store, edit, organize, analyze, and display spatially-referenced data (35). Georefer- encing is usually expressed in terms of Cartesian coor- dinates (northing, easting, elevation), latititude/ longitude, postal codes, or various area divisions. GISs can be used to generate maps and data needed for risk factor identification, as well as to perform some spatial analytical techniques useful in health care research and epidemiology. In this section, we will highlight some of the advantages and limitations of GISs in health research and epidemiology and describe some examples of the use of GISs.
The advantages of GISs include the ability to handle repetitive tasks and quickly compare spatial data from various sources and different spatial areas. The speed of data manipulation, the ability to handle large volumes of
Epidemiol Rev Vol. 2 1 , No. 2, 1999
Spatial Analytical Methods/Geographic Information 155
data, the enforcement of standardization, and the ability to ask "what if' questions are distinct advantages. Some functions include reclassification of data, Boolean searches, creation of buffer zones, handling changes in map projections or scales, referencing attribute data, and providing detailed cartographic output. Another advan- tage of GISs is the ability to use data from remote sens- ing, such as digital satellite imagery of the earth's sur- face, for physical environmental analyses.
Historic disadvantages of GISs included mainte- nance and use of the system, which required special expertise. Currently, GIS programs are easier to learn and use. Although many functions of a GIS can be car- ried out manually on very small datasets (9), larger datasets necessitate a geographic database manager. The benefits of GIS technology to public and environ- mental health have been reviewed (10). GIS applica- tions, such as automated atlases for disease rates and the detection of disease clusters, are already in use.
Health data often come from official or government sources. There are numerous problems associated with these data such as underreporting, coding errors, and diagnostic errors (108). The data can be inaccurate, incomplete, and unreliable (1). In the context of this review, those problems are important, but the spatial data problems with these datasets need primary con- sideration. Issues of spatial detail, which is generally lacking in official statistics, spatial incompatibility or inconsistency among sources, and spatial referencing are all important considerations when using health data collected for purposes other than use in a GIS (9). Usefulness of GISs to health information management will depend on data quality and spatial attributes. All spatial analyses within a GIS must be done at an appropriate and consistent scale, although defining that scale is not always easy (109).
A final limitation to GIS data is that they usually represent static points in time and do not adequately represent spatial-temporal information, which may be of particular interest in infectious disease epidemiol- ogy. Map resolution, map format or map projection, availability, expense, quality, and age of the data or map need to be considered when using spatially- referenced datasets. The question of data scale is important to identify significant relations. What is apparent at one scale may not be apparent at another (1). In a critical review of the use of GISs in veteri- nary epidemiology, Paterson (110) cautioned readers about the seductiveness of the visually attractive out- put from a GIS. He warned that there is a tendency for users of GISs and readers of the output to be seduced by attractive graphics and forget rules of data man- agement, analysis, presentation, and interpretation of results.
For the most effective use of GISs in epidemiology and health research, a GIS must be have the ability to conduct or be linked to spatial analytical and modeling programs. In the past, this was a major weakness of most GIS software programs (9). Many statistical analyses using spatial data can now be done within the GIS environment, although more advanced or complex techniques may require data analysis within a statisti- cal software package with export to the GIS for map displays. Some GIS programs have been linked to sta- tistical packages, such as the SpaceStat Extension (111) for Arc View® (112).
GISs can be used as tools for health planning of pri- mary care delivery and access based on geography, locality, distance, or population structure. Demographic information, in the way of census data, is already devel- oped for GIS applications at a resolution of census tract or zip code. These data in the United States are linked with geographic base data and are known as the TIGER files (113). Numbers of individuals, location, and demographic information linked to location are avail- able for a wide variety of uses.
Recently, investigators in both Great Britain and Sweden have reported using GISs as tools for assess- ing access of populations to primary health care facil- ities (114, 115). The GIS was used to apply many dif- ferent sets of geographically-linked datasets (map layers) in order to investigate the problem of primary care delivery. Bullen et al. (114) used GISs to develop map layers which contained the important criteria for defining local general practice catchment areas. These included locations of the medical practices and factors which act as physical or psychologic barriers, such as commuting distances or time, to use of the practice. Central to the definition of the practice localities was the use of a large matrix of patient-to-practice flows based on post-coded data. Clear geographic patterns of patient allegiance to practices were delineated. The GIS allowed the development of alternative regional- izations using both visualization and statistical analy- ses for practice locality profiling. In the Swedish study, one function of a GIS, the ability to calculate dis- tances, was used to identify access to primary health care using property registries and the location of pri- mary health care facilities (115). Hyndman et al. (116) used a GIS to study where to locate mammography and abdominal aortic aneurysm screening clinics in Perth, Australia. Using small area census demographic data and geocodes for household water meters, her group was able to investigate the effect of distance from residence on screening clinic use.
Mapping programs have also been useful in devel- oping variables for multivariate analysis as well as for visualization. MacKinnon et al. (117) studied drunk
Epidemiol Rev Vol. 2 1 , No. 2, 1999
156 Moore and Carpenter
driving arrests and other alcohol-related problems and found that they were locationally related to alcohol retail outlet density in Los Angeles County, California.
In addition to chronic disease problems and health care, GISs have been valuable in studying infectious diseases. Lyme disease and other vector-borne diseases are good examples of diseases studied with a GIS. The geographically static nature of the disease vectors makes it easier to map their locations and make sense of habitat determinants. In an analysis of the distribution of Lyme disease in Wisconsin, a GIS was used to asso- ciate county-level data on tick-distribution, human pop- ulation density, Lyme disease case distribution, and pro- portion of wooded areas to help explain the distribution of the disease in the state (64). The GIS allowed inves- tigators to obtain location data and measure distances between locations. Measures of spatial autocorrelation and local spatial statistics were used to identify clusters of disease cases or nonrandom patterns of environmen- tal risk factors.
In infectious disease epidemiology, GISs are often used in combination with a statistical modeling tech- nique such as logistic regression. In a Lyme disease study of environmental risk factors, a GIS was used for a study of one county in Maryland to generate envi- ronmental risk factors such as land use/land cover, for- est distributions, soils, elevation, and watersheds (118, 119). These variables were then used in a logistic regression analysis to model risk factors for cases of Lyme disease in certain areas.
Pseudorabies virus infection in swine results in quarantine of the herd in most states in the United States. GISs have been used to assess spatial and tem- poral aspects of pseudorabies epidemiology. In one study, Norman et al. (120) used buffer analysis to determine the radius around a quarantined herd that captured surrounding herds that subsequently became quarantined. The study emphasized the importance of spatial proximity and the contagious nature of the spread of this disease. Pseudorabies researchers in Minnesota used GISs in a similar way (121).
Remote sensing and a GIS have been used to assess site-specific risk for the intermediate snail host for fas- cioliasis (cattle liver flukes) in Louisiana (122-124). Environmental variables were elucidated from the GIS and used in a regression analysis of the farm risk index against the log of the highest worm egg counts in one study. Soil maps and snail habitat were evaluated in another. Remote sensing and a GIS were also used in a project on Guinea worm disease eradication (125).
Surveillance and monitoring of infectious diseases are other uses for GISs in applied epidemiology. A national computerized surveillance system of Anopheles mosquito breeding sites and imported cases
of malaria was established in 1992 in Israel (126). It was developed to identify the risk of malaria, given an introduced case, in order to target appropriate control measures. GIS mapping is currently being used to identify risk of Vibrio vulnificus infections in oyster bays in Louisiana (127). Bovine tuberculosis control is being affected by the use of GISs in southwest England (128). A GIS and spatial statistics were used in a dengue investigation in Puerto Rico (129). In this study, Morrison et al. demonstrated a very rapid tem- poral and spatial progression of dengue in the commu- nity, and concluded that control measures would best be applied to the entire municipality rather than local areas around affected households. GISs have also been used to map and analyze information on African try- panosomiasis, leishmaniasis, schistosomiasis, and food-borne trematode infections (130).
CONCLUSIONS
There are a number of considerations when deciding to use a spatial analytical technique or GIS in epi- demiology and health research. Problems in spatial analyses in epidemiology include the need to aggre- gate disease occurrences over space and time, which may gain data stability but loses information; the accu- racy of death certificate and health information for diagnoses; the choice of a suitable rate standardization procedure; the choice or scale and data classes when making maps of disease rates; and the problem of eco- logic fallacies (7). All data are subject to problems of incompleteness and inaccuracies due to measurement error (2).
The choice of an analytical technique is of concern because of the nature of spatial data, scalar influences, spatial dependence, and spatial autocorrelation. For example, the identification of disease patterns is dependent on the scale selected. The selection of one scale may mask or ignore spatial variations in another. Replication of findings in a number of different scales would tend to confirm an hypothesis. In addition, spa- tial studies of human conditions are often hampered in highly mobile societies by in-migration and out- migration (2). To identify clusters, for example, near- est neighbor analysis depends on the scale at which the investigator draws the data. Openshaw's solution to identify disease clusters was to generate and test all possible geographic hypotheses relevant to a particu- lar problem through the development of the GAM, an automated modeling system (46). Spatial autocorrela- tion may pose a problem when using least squares regression. When using a GIS, several concerns about data have been outlined, but also include database availability in advance of the hypotheses under study (46). The construction of post-data models is an
Epidemiol Rev Vol. 2 1 , No. 2, 1999
Spatial Analytical Methods/Geographic Information 157
important problem since data resolution, data quality, and other attributes will dictate what can eventually be examined.
In general, when choosing a spatial analytical tech- nique, the investigator needs to determine what type of data will be analyzed (continuous, dichotomous or cat- egorical, and point or areal), what process will be examined (detection of clusters, interpolation, or diffu- sion), what spatial data are available, what scale should be chosen for the analysis, the availability of commercial software for the technique (versus the requirement for custom programming), and the reported power of the test.
When selecting a GIS program, the investigator should identify program cost, what spatial datasets are readily available in the program's format, what spatial analytical tools are available with the program (some have software extensions to help with analysis), whether the program is readily compatible with statis- tical software packages or output data from statistical software, how easy the program is to learn or operate,
at what scale geocoding can be done within the pro- gram, what techniques the program uses for data inter- polation (e.g., kriging, splining), if the program will measure distances if needed, whether the output maps are of sufficient quality, how much hard-disk space and memory the program requires, and what kind of support (in the way of educational programs or trouble-shooting support) is available from the com- pany. Table 2 provides a few select websites of poten- tial use to epidemiologists regarding general information about GISs and some spatial datasets, and their avail- ability to investigators.
Spatial analytical techniques and GISs have many current uses in epidemiology and health research. Network analysis for examination of patient referrals, modeling diffusion patterns, identification of environ- mental risks and risk assessment, detection of clusters, Monte Carlo simulations to assess point or areal pat- terns of disease, or any assessment of "distance" are techniques which can be used to explore the spatial aspects of health care, injury, and disease. GISs have
Table 2. Selected websites for geographic information systems and spatial, population, and health databases potentially useful to epidemiologists
Website Site definition
http://info.er.usgs.gov/research/gis/title.html
http://www.census.gov
The US Geological Survey site for information on available spatial data such as digital elevation models (DEMs) and digital line graphs (DLGs)
US Bureau of the Census site for digital map databases such as the TIGER/Line® data; has sample data sets and shows availability of state data
http://www.clarklabs.org
http://www.cdc.gov/nchswww/datawh/datawh.htm
http://www.esri.com
http://www.mapinfo.com
http://www.geo.ed.ac.uk/home/giswww.html
http://www.idrc.ca/library/document/gis.html
http://www.king.ac.uk/geog/gis/intro.htm
Site for the IDRISI raster geographic information system and imaging processing software; has a demonstration program, teaching manuals, and support; currently used in risk/hazards research and for disease incident data and environmental health
The National Center for Health Statistics Data Warehouse with links to health and mortality data
Site for products: Arclnfo, ArcView, AtlasGIS; also links in ArcData Online for downloadable, free geographic data such as the US Bureau of the Census TIGER/Line® 1995 data, the Federal Emergency Management Agency's flood data, topographic maps, and US street data
For Maplnfo product information and available data
For an index of GIS* resources (other sites like this are available and numerous)
GIS, Health, and Epidemiology: an Annotated Resource Guide; a bibliography of useful references
Good overview of GIS, glossary of terms, and acronyms
*GIS, geographic information systems.
Epidemiol Rev Vol. 21, No. 2, 1999
158 Moore and Carpenter
become more user-friendly and users have access to many more spatially-referenced datasets.
GISs will be increasingly helpful to epidemiologists to vizualize, manage, and analyze large volumes of data. They can help to better define populations and environ- mental exposures with perhaps better specificity. They have already found use in disease surveillance pro- grams. In public health, sentinel physicians located across geographic space with location and disease infor- mation linked to a GIS can provide data to help rapidly identify the location of emerging disease problems and diffusion patterns of disease. Epidemiologists should look forward to better links between GISs and statistical or analytical programs. When working with case- locations for public health or research, however, they will find themselves in debates about patient confiden- tiality and the spatial scale at which to report health data versus the need for research or the public's right to know exactly where diseases are occurring.
Future directions in GISs involve spatial data min- ing and visualization. Currently, most mapping pro- grams produce high-quality, detailed, but static, two- dimensional maps. Dynamic mapping, spatio-temporal GIS visualization, and exploratory spatial data analysis are new areas in GISs that epidemiologists might find useful. When coupled with environmental or other spatially-related risks, interpolation of data and simu- lation of historic or future events might provide insight into disease diffusion processes and predictions. Just as GISs have redefined our ability to work with spatial data in epidemiology, these newer techniques will help to redefine our analysis of epidemiologic, spatially- referenced data.
ACKNOWLEDGMENTS
The first author would like to acknowledge an important mentor, Dr. Peter Gould, from the Department of Geography at The Pennsylvania State University, for pointing her in the right geographic direction and showing her the way. Many thanks go to the reviewers for their helpful suggestions in revising this manuscript.
REFERENCES
1. Matthews SA. Epidemiology using a GIS: the need for cau- tion. Comput Environ Urban Syst 1990;14:213-21.
2. Mayer JD. The role of spatial analysis and geographic data in the detection of disease causation. Soc Sci Med 1983;17:1213-21.
3. Martin SW, Meek AH, Willeberg P. Veterinary epidemiology: principles and methods. Ames, IA: Iowa State University Press, 1987.
4. Mausner JS, Kramer S. Epidemiology: an introductory text. 2nd ed. Philadelphia, PA: Saunders, 1985.
5. Rothman KJ. Modern epidemiology. Boston, MA: Little, Brown, 1986.
6. Kleinbaum DG, Kupper LL, Morgenstern H. Epidemiologic research: principles and quantitative methods. New York, NY: Van Nostrand Reinhold, 1982.
7. Gesler W. The uses of spatial analysis in medical geography: a review. Soc Sci Med 1986;23:963-73.
8. Keams RA, Joseph AE. Space in its place: developing the link in medical geography. Soc Sci Med 1993;37:711-17.
9. Twigg L. Health based geographical information systems: their potential examined in the light of existing data sources. Soc Sci Med 1990,30:143-55.
10. Scholten HJ, de Lepper MJC. The benefits of the application of geographical information systems in public and environ- mental health. World Health Stat Q 1991;44:160-70.
11. Snow J. The case books of Dr. John Snow. In: Medical history: supplement no. 14. London, England: Wellcome Institute for the History of Medicine, 1994.
12. Burkitt DP. A 'tumour safari' in east and central Africa. Br J Cancer 1962; 16:379-86.
13. Burkitt DP. Geography of a disease: purpose and possibilities from geographical medicine. In: Rothschild HR, ed. Biocultural aspects of disease. New York, NY: Academic Press, 1981:133-51.
14. Goodchild MR Geographical information science. Int J Geogrlnf Syst 1992;6:31-45.
15. Openshaw S. Spatial analysis and geographic information systems: a review of progress and possibilities. In: Scholten HJ, Stillwell JCH, eds. Geographical information systems for urban and regional planning. Boston, MA: Kluwer Academic Publishers, 1990:153-63.
16. Tobler W. Cellular geography. In: Gale S, Olsson G, eds. Philosophy in geography. Dordrecht, The Netherlands: Reidel, 1979:379-86.
17. Douven W, Scholten HJ. Spatial analysis in health research. In: de Lepper MJC, Scholten HJ, Stern RM , eds. The added value of geographical information systems in public and environmental health. Boston, MA: Kluwer Academic Press, 1995:117-33.
18. Cliff AD, Haggett P. Atlas of the disease distribution: analyt- ical approaches to epidemiological data. Oxford, England: Basil Blackwell, 1988.
19. Smallman-Raynor MR, Cliff AD, Haggett P. Atlas of AIDS. Oxford, England: Blackwell Publishers, 1992.
20. Pickle LW, Mungiole M, Jones GK, et al. Atlas of United States mortality. Hyattsville, MD: US Department of Health and Human Services, Centers for Disease Control and Prevention, National Center for Health Statistics, 1997.
21. Westlake A. Strategies for the use of geography in epidemio- logic analysis. In: de Lepper MJC, Scholten HJ, Stern RM , eds. The added value of geographical information systems in public and environmental health. Boston, MA: Kluwer Academic Press, 1995:135-44.
22. Dean JA, Burton AH, Dean AG, et al. Epi map: a mapping program for IBM-compatible microcomputers. Atlanta, GA: Centers for Disease Control and Prevention, 1993.
23. Rothman KJ. A sobering start for the cluster busters confer- ence. Am J Epidemiol 1990;132(suppl):S6-13.
24. Alexander F, Cartwright RA, McKinney PM. A comparison of recent statistical techniques of testing for spatial cluster- ing: preliminary results. In: Elliott P, ed. Methodology of enquiries into disease clustering. London, England: London School of Hygiene and Tropical Medicine, 1988:23-33.
25. Besag J, Newell J. The detection of clusters in rare diseases. J R Stat Soc 1991;154:143-55.
26. Kulldorff M, Nagarwalla N. Spatial disease clusters: detec- tion and inference. Stat Med 1995; 14:799-810.
27. Waller LA, Lawson AB. The power of focused tests to detect disease clustering. Stat Med 1995;14:2291-308.
28. Waller LA, Jacquez GM. Disease models implicit in statisti-
Epidemiol Rev Vol. 2 1 , No. 2, 1999
Spatial Analytical Methods/Geographic Information 159
cal tests of disease clustering. Epidemiology 1995;6:584-90. 29. Elliott P, Martuzzi M, Shaddick G. Spatial statistical methods
in environmental epidemiology: a critique. Stat Methods MedRes 1995 ;4:137-59.
30. Marshall RJ. A review of methods for the statistical analysis of spatial patterns of disease. J R Stat Soc 1991;154:421—41.
31. Jacquez GM, Waller LA, Grimson RC, et al. The analysis of disease clusters. Part I: State of the art. Infect Control Hosp Epidemiol 1996;17:319-27.
32. Jacquez GM, Grimson RC, Waller LA, et al. The analysis of disease clusters. Part II: Introduction to techniques. Infect Control Hosp Epidemiol 1996; 17:385-97.
33. Cuzick J, Edwards R. Spatial clustering for inhomogeneous populations. J R Stat Soc 1990;52:73-104.
34. Gilg AW. A study in agricultural disease diffusion: the case of the 1970-71 fowl pest epidemic. Trans Inst Br Geogr 1973; 59:77-97.
35. Bailey TC, Gatrell AC. Interactive spatial data analysis. New York, NY: Wiley, 1995.
36. Griffith DA, Arnrhein CG. Statistical analysis for geogra- phers. Englewood Cliffs, NJ: Prentice Hall, 1991.
37. Thompson HR. Distribution of distance to Mh neighbour in a population of randomly distributed individuals. Ecology 1965;37:391^.
38. Bithell JF. An application of density estimation to geograph- ical epidemiology. Stat Med 1990;9:691-701.
39. Gatrell AC, Bailey TC. Interactive spatial data analysis in medical geography. Soc Sci Med 1996;42:843-55.
40. Glaser SL. Spatial clustering of Hodgkin's disease in the San Francisco Bay area. Am J Epidemiol 1990;132(suppl): S167-77.
41. Whittemore AS, Friend N, Brown BW, et al. A test to detect clusters of disease. Biometrika 1987;74:631-5.
42. Selvin S, Merrill D, Schulman J, et al. Transformations of maps to investigate clusters of disease. Soc Sci Med 1988; 26:215-21.
43. Miyawaki N, Chen SC. A statistical consideration on the mapping of mortality. Soc Sci Med 1981;15D:93-101.
44. Hopkinson ND, Muir KR, Oliver MA, et al. Distribution of cases of systemic lupus erythematosus at time of first symp- tom in an urban area. Ann Rheum Dis 1995;54:891-5.
45. Weiss KB, Wagener DK. Geographic variations in US asthma mortality: small-area analyses of excess mortality, 1981-1985. Am J Epidemiol 1990;132(suppl):S107-15.
46. Openshaw S, Charlton M, Wymer C, et al. A Mark 1 Geographical Analysis Machine for the automated analysis of point data sets. Int J Geogr Inf Syst 1987;l:335-58.
47. Hjalmars U, Kulldorff M, Gustafsson G, et al. Childhood leukemia in Sweden: using GIS and a spatial scan statistic for cluster detection. Stat Med 1996,15:707-15.
48. Kulldorff M, Feuer EJ, Miller BA, et al. Breast cancer clus- ters in the northeast United States: a geographic analysis. Am J Epidemiol 1997;146:161-70.
49. Hungerford LL. Use of spatial statistics to identify and test significance in geographic disease patterns. Prev Vet Med 1 9 9 1 1 1 2 3 7 1 2
50. Wright SE. The spatial distribution and geographic analysis of endodontic office locations at the national scale. J Endod 1994;20:500-5.
51. Cliff AD, Ord JK. Spatial processes, models and applica- tions. London, England: Pion, 1981.
52. Ohno Y, Aoki K, Aoki N. A test of significance for geo- graphic clusters of disease. Int J Epidemiol 1979;8:273-80.
53. Colonna M, Esteve J, Menegoz F. Detecting spatial autocor- relation of cancer risk when population density is heteroge- neous. (In French). Rev Epidemiol Sante Publique 1993 ;41: 235-40.
54. Walter SD. The analysis of regional patterns in health data. I. Distributional considerations. Stat Med 1992;136:730-41.
55. Walter SD. The analysis of regional patterns in health data. II. The power to detect environmental effects. Am J Epidemiol 1992;136:742-59.
56. Walter SD, Birnie SE, Marrett LD, et al. The geographic variation of cancer incidence in Ontario. Am J Public Health 1994;84:367-76.
57. Geary R. The contiguity ratio and statistical mapping. Incorporated Statistician 1954;5:115—45.
58. Moran PAP. The interpretation of statistical maps. J R Stat SocB 1948;19:243-51.
59. Jumars P, Thistle D, Jones M. Detecting two dimensional spa- tial structure in biological data. Oncologia 1977;28:109-23.
60. Walter SD. Assessing spatial patterns in disease rates. Stat Med 1993;12:1885-94.
61. Shafer S. Mapping bone cancer death rates in Pennsylvania counties. Soc Sci Med 1980;14D:ll-15.
62. Glick B. The spatial autocorrelation of cancer mortality. Soc Sci Med [Med Geogr] 1979;13D: 123-30.
63. Lanska DJ, Kryscio R. Geographic distribution of hospital- ization rates, case fatality, and mortality from stroke in the United States. Neurology 1994;44:1541-50.
64. Kitron U, Kazmierczak JJ. Spatial analysis of the distribution of Lyme disease in Wisconsin. Am J Epidemiol 1997; 145: 558^66.
65. Austin CC, Weigel RM. Factors affecting the geographic dis- tribution of pseudorabies (Aujesky's disease) virus infection among swine herds in Illinois. Prev Vet Med 1992;13: 239-50.
66. Goodchild MF. Spatial autocorrelation. Norwich, CN: Geo Books, 1985.
67. Davis JC. Statistics and data analysis in geology. New York, NY: Wiley, 1986.
68. Kennedy S. A geographic regression model for medical sta- tistics. Soc Sci Med 1988;26:119-29.
69. Buslenko NP. The Monte Carlo method: the method of sta- tistical trials. International series of monographs in pure and applied mathematics. Vol 87. New York, NY: Pergamon Press, 1966.
70. Abler R, Adams JS, Gould P. Spatial diffusion: meshing space and time. In: Spatial organization: the geographer's view of the world. Englewood Cliffs, NJ: Prentice-Hall, 1990:389^51.
71. Kabel JA. Geographic perspective on AIDS in the United States: past, present, and future. PhD dissertation. University Park, PA: The Pennsylvania State University, 1992.
72. Kallen A, Arcuri, P, Murray JD. A simple model for the spa- tial spread and control of rabies. J Theor Biol 1985;116: 377-93.
73. Gould P. Spreading HIV across America with an air passen- ger operator. In: Proceedings of the International Symposium on Computer Mapping in Epidemiology and Environmental Health. Tampa, FL: World Computer Graphics Foundation and The University of South Florida, 1997:159-62.
74. Cliff AD, Haggett P, Ord JK, et al. Spatial diffusion: an his- torical geography of epidemics in an island community. Cambridge, England: Cambridge University Press, 1981.
75. Cliff AD, Haggett P, Smallman-Raynor MR, et al. The appli- cation of multidimensional scaling methods to epidemiolog- ical data. Stat Methods Med Res 1995;4:102-23.
76. Smith CJ, Hanham RQ. Any place but here! mental health facilities as noxious neighbors. Prog Geogr 1981;33:326-34.
77. Gould P. Epidemiologie et sante. In: Bailly A, Ferras R, Pumain, eds. Geographie et le Monde Contemporain. (In French). Paris, France: Editions Economica, 1992.
78. Gould P, Wallace R. Spatial structures and scientific para- doxes in the AIDS pandemic. Geogr Ann 1994;76:105-16.
79. Wallace D. The resurgence of tuberculosis in New York City: a mixed hierarchically and spatially diffused epidemic. Am J Public Health 1994;84:1000-2.
80. Morrill R, Gaile GL, Thrall GI. Spatial diffusion. Newbury Park, CA: Sage Publications, 1988.
81. Jeltsch F, Muller MS, Grimm V, et al. Pattern formation trig- gered by rare events: lessons from the spread of rabies. Proc R Soc Lond B Biol Sci 1997;264:495-503.
82. Rogers EM. Network analysis of the diffusion of innova-
Epidemiol Rev Vol. 2 1 , No. 2, 1999
160 Moore and Carpenter
tions. In: Holland PW, Leinhardt S, eds. Perspectives on social network research. New York, NY: Academic Press, 1979:137-64.
83. Unwin DJ. An introduction to trend-surface analysis. In: CATMOG 6: An introduction to trend-surface analysis. Norwich, CN: Geo Abstracts, 1975:3-35.
84. Angulo JJ, Haggett P, Megale P, et al. Variola minor in Braganca Paulista County, 1956: a trend-surface analysis. Am J Epidemiol 1977;105:272-8.
85. Kwofie KM. A spatio-temporal analysis of cholera diffusion in Western Africa. Econ Geogr 1976;52:127-35.
86. Moore DA. Spatial diffusion of raccoon rabies in Pennsylvania, USA. Prev Vet Med 1999;40:19-32.
87. Gould P, Kabel J, Gorr W, et al. AIDS: predicting the next map. Interfaces 1991;21:80-92.
88. Casetti E, Jones JP. An introduction to the expansion method and its applications. In: Casetti E, Jones JP, eds. Applications of the expansion method. London, England: Routledge, 1992:1-9.
89. Briggs DJ. Mapping environmental exposure. In: Cuzick J, English D, et al., eds. Geographical and environmental epi- demiology: methods for small area studies. New York, NY: Oxford University Press, 1992:158-76.
90. Karvonen M, Tuomilehto J, Naukkarinen A, et al. The preva- lence and regional distribution of antibodies against Chlamydia pneumoniae (strain TWAR) in Finland in 1958. Int J Epidemiol 1992,21:391-8.
91. Oliver MA, Webster R. Kriging: a method of interpolation for geographical information systems. Int J Geogr Inf Syst 1990;4:313-32.
92. Cressie NAC. Statistics for spatial data. New York, NY: Wiley, 1991.
93. Torok TJ, Kilgore PE, Clarke MJ, et al. Visualizing geo- graphic and temporal trends in rotavirus activity in the United States, 1991 to 1996. National Respiratory and Enteric Virus Surveillance System Collaborating Laboratories. Pediatr Infect Dis J 1997;16:941-6.
94. Ribeiro JM, Seulu F, Abose T, et al. Temporal and spatial dis- tribution of anopheline mosquitoes in an Ethiopian village: implications for malaria control strategies. Bull World Health Organ 1996;74:299-305.
95. Rushton G, Krishnamurthy R, Krishnamurti D, et al. The spatial relationship between infant mortality and birth defect rates in a US city. Stat Med 1996;15:1907-19.
96. Carrat F, Valleron AJ. Epidemiologic mapping using the "kriging" method: application to an influenza-like illness epidemic in France. Am J Epidemiol 1992;135:1293-300.
97. Pikhart H, Prikazsky V, Bobak M, et al. Association between ambient air concentrations of nitrogen dioxide and respira- tory symptoms in children in Prague, Czech Republic: pre- liminary results from the Czech part of the SAVIAH Study. Small Area Variations of Air Pollution and Health. Cent Eur J Public Health 1997;5:82-5.
98. Georgopoulos PG, Purushothaman V, Chiou R. Comparative evaluation of methods for estimating potential human expo- sure to ozone: photochemical modeling and ambient moni- toring. J Expo Anal Environ Epidemiol 1997;7:191-215.
99. Using the ArcView spatial analyst. Redlands, CA: Environmental Systems Research Institute, 1996.
100. Devine OJ, Parrish RG. Monitoring the health of a popula- tion. In: Stroup DF, Teutsch SM, eds. Statistics in public health: qualitative approaches to public health problems. New York, NY: Oxford University Press, 1998:59-91.
101. Devine OJ, Louis TA, Halloran ME. Empirical Bayes meth- ods for stabilizing incidence rates before mapping. Epidemiology 1994;5:622-30.
102. Smith JS, Yager PA, Bigler WJ, et al. Surveillance and epi- demiologic mapping of monoclonal antibody-defined rabies variants in Florida. J Wildl Dis 1990;26:473-85.
103. Pool GE, Hacker CS. Geographic and seasonal distribution of rabies in skunks, foxes, and bats in Texas. J Wildl Dis 1982;18:405-18.
104. Smith JC. Locating alcoholism treatment facilities. Econ Geogr 1983;59:368-85.
105. Cook DG, Pocock SJ. Multiple regression in geographical mortality studies, with allowance for spatially correlated errors. Biometrics 1983;39:361-71.
106. Harries K. Social stress and trauma: synthesis and spatial analysis. Soc Sci Med 1997;45:1251-64.
107. Foster SA, Gorr W, Wimberly FC. A comparison of drift analysis and the expansion method: the evaluation of federal policies on the supply of physicians. In: Casetti E, Jones JP, eds. Applications of the expansion method. London, England: Routledge, 1992:94-114.
108. Lilienfeld DE, Stolley PD. Morbidity statistics. In: Lilienfeld DE, Stolley PD, eds. Foundations of epidemiology. Rev ed. New York, NY: Oxford University Press, 1994:101-50.
109. Briggs DJ, Elliott P. The use of geographical information sys- tems in studies on environment and health. World Health Stat Q 1995;48:85-94.
110. Paterson AD. Problems encountered in the practical imple- mentation of geographical information systems (GIS) in vet- erinary epidemiology. Presented at the meeting of the Society of Veterinary Epidemiology and Preventive Medicine, Reading, United Kingdom, March 29, 1995.
111. Anselin L, Bao S. Exploratory spatial data analysis linking SpaceStat and Arcview. In: Fischer M, Getis A, eds. Recent developments in spatial analysis. Berlin, Germany: Springer- Verlag, 1998:35-59.
112. ArcView GIS, 3.1. Redlands, CA: Environmental Systems Research Institute, 1998.
113. TIGER/Line files, 1995. Washington, DC: US Department of Commerce, 1996.
114. Bullen N, Moon G, Jones K. Defining localities for health planning: a GIS approach. Soc Sci Med 1996;42:801-16.
115. Kohli S, Sahlen K, Sivertun A, et al. Distance from the pri- mary health center: a GIS method to study geographical access to health care. J Med Syst 1995; 19:425-36.
116. Hyndman J, Holman CD, Jamrozik K. The effect of spatial definition on the allocation of clients to screening clinics. Soc Sci Med 1997;45:331^0.
117. MacKinnon DP, Scribner R, Taft KA. Development and applications of a city-level alcohol availability and alcohol problems database. Stat Med 1995;14:591-604.
118. Glass GE, Schwartz BS, Morgan JM III, et al. Environmental risk factors for Lyme disease identified with geographic information systems. Am J Public Health 1995;85:944-8.
119. Glass GE, Morgan JM, Johnson DT, et al. Infectious dis- eases: epidemiology and GIS: a case study of Lyme disease. Geogr Inf Syst 1992;2:65-9.
120. Norman HS, Sischo WM, Pitcher P, et al. Spatial and tempo- ral epidemiology of pseudorabies virus infection. Am J Vet Res 1996;57:1563-8.
121. Marsh WE, Damrongwatanapokin T, Larntz K, et al. The use of a geographic information system in an epidemiological study of pseudorabies (Aujesky's disease) in Minnesota swine herds. Prev Vet Med 1991;ll:249-54.
122. Zukowski SH, Hill JM, Jones FW, et al. Development and validation of a soil-based geographic information system model of habitat of Fossaria bulimoides, a snail intermediate host of Fasciola hepatica. Prev Vet Med 1992;11:221-7.
123. Malone JB, Fehler DP, Loyacano AF, et al. Use of LAND- SAT MSS imagery and soil type in a geographic information system to assess site-specific risk of fascioliasis on Red River Basin farms in Louisiana. Ann N Y Acad Sci 1992;653:389-97.
124. Zukowski SH, Wilkerson GW, Malone JB Jr. Fasciolosis in cattle in Louisiana. II. Development of a system to use soil maps in a geographic information system to estimate disease risk on Louisiana coastal march rangeland. Vet Parasitol 1993;47:51-65.
125. Clarke KC, Osleeb JP, Sherry JM, et al. The use of remote sensing and geographic information systems in UNICEFs dracunculiasis (Guinea worm) eradication effort. Prev Vet
Epidemiol Rev Vol. 2 1 , No. 2, 1999
Spatial Analytical Methods/Geographic Information 161
Med 1991;ll:229-35. 126. Kitron U, Pener H, Costin C, et al. Geographic information
system in malaria surveillance: mosquito breeding and imported cases in Israel, 1992. Am J Trap Med Hyg 1994; 50:550-6.
127. Hugh-Jones M, Wilson K, Scheffler S. GIS mapping of expected Vibrio vulnificus levels on southern Louisiana oys- terbays. Presented at the meeting of the Society for Veterinary Epidemiology and Preventive Medicine, Glasgow, Scotland, March 27, 1996.
128. Clifton-Hadley RS. The use of a geographical information system (GIS) in the control and epidemiology of bovine tuberculosis in south-west England. Presented at the meeting of the Society for Veterinary Epidemiology and Preventive Medicine, Exeter, United Kingdom, March 31, 1993.
129. Morrison AC, Getis A, Santiago M, et al. Exploratory space- time analysis of reported dengue cases during an outbreak in Florida, Puerto Rico, 1991-1992. Am J Trap Med Hyg 1998;58:287-98.
130. Mott KE, Nuttall I, Desjeux P, et al. New geographical approaches to control of some parasitic zoonoses. Bull World Health Organ 1995;73:247-57.
131. Carpenter TE. TSpStat, time-space statistics: a spreadsheet add-in. Davis, CA: School of Veterinary Medicine, University of California, Davis, 1999.
132. System for epidemiologic analysis. Cluster 3.1. Atlanta, GA: US Department of Health and Human Services, Public Health Service, Agency for Toxic Substances and Disease Registry, 1993.
133. Applied Biomathematics. CAST, cluster analysis in space and time, 2.0. Setauket, NY: Applied Biomathematics, 1993.
134. Jacquez GM. Stat!: statistical software for the clustering of health events. Ann Arbor, MI: BioMedware, 1994.
135. Rowlingson BS, Diggle PJ. SPLANCS: a spatial point pat- tern analysis code in S-PLUS. Comput Geosci 1993; 19: 627-55.
136. Kulldorf M, Rand K, Williams G. SaTScan: program for the space and time scan statistic, 1.0. Bethesda, MD: National Cancer Institute, 1996.
137. Gould P. Dynamic structures of geographic space. In: Brunn SD, Leinbach TR, eds. Collapsing space and time: geo- graphic aspects of communications and information. London, England: Harper Collins Academic, 1991:3-30.
138. Geostatistical environmental assessment software. GEO- EAS, 1.2.1. Las Vegas, NV: US Environmental Protection Agency, Office of Research and Development, Environmental Monitoring Systems Laboratory, 1991.
139. Geographic resources analysis support system. GRASS, 5.0. Waco, TX: Baylor University, 1999.
Epidemiol Rev Vol. 21, No. 2, 1999