Foundations of Geographic Information Systems Final Project
Foundations of Geographic Information Science
GIS5103 – Fall 2019
Week 5
Data Input, Throughput, & Quality
By the end of this week you should be able to:
Describe methods of data input
Describe Digitizing sources ‐ the cartographic base
Contrast Digitizing methods within a GIS (Tablet vs. “Heads Up”)
Describe Data editing (correcting digitizing errors, re‐projections
and generalizations)
Define accuracy and precision
Describe the importance of data standards and metadata
List sources of errors in a GIS project
Differentiate spatial and attribute errors
Describe techniques for coping with errors in Data
The Data Stream
The focus of this week’s lecture component is on creating GIS or spatial data from a variety of
sources. We will also be looking at how GIS data is updated.
Digitizing GIS Data
As we have discussed over the past several weeks one of the primary features of a GIS is its ability to integrate data from a variety of sources. Depending on how the data is to be used the need to convert the data or extract features from it may arise.
Fortunately there are a number of techniques and tools available to assist in acquiring data for
a new GIS is straightforward these days. Some of the sources for this data can include some of the following:
GPS (Global Positioning Systems) has become a major source of new GIS data. Data acquired through GPS can be highly spatially accurate (1cm horizontal) depending upon the type of receiver used.
Digital map images such as scanned maps, LiDAR , satellite data, and aerial photographs (Ortho photos) are now often used as a cartographic base for digitizing; this is another way to create new GIS data layers.
Remotely sensed satellite data are becoming an important source of GIS data as the cost of data falls, free data becomes available and new types of data emerge.
The following examples show the different spatial scales commonly found
with remote sensing data. Examples are also provided that show how they can be used to digitize temporal changes or changes to features over time (urbanization) or after an event (Natural Disaster).
Satellite data are raster data. The amount of spatial detail represented by one pixel of a raster
image determines its spatial resolution. So, an image with one meter spatial resolution means
that each pixel in the image represents one square meter on the ground. With improvements in
cameras and digital image acquisition the cost of acquiring aerial images has decreased. It is now possible to obtain imagery with a 6” pixel resolution for a few thousand dollars.
30 meters
10 meters
5 meters
The Different Spatial Resolutions
Image Resolutions Continued
2 meters
1 foot
In cartographic applications, higher spatial resolutions are often favored. However, for applications such as weather forecasting or to monitor seasonal plant growth on a
yearly basis for North America, coarser spatial resolutions are favored. While this type of data
can be extremely useful there are data storage considerations and large file sizes to consider which may impact the accessibility of these images especially in web based applications.
1 meter
Where to obtain remote sensing data
There are a number of useful websites from which remote
sensing data (satellite images) can be downloaded. Here are some examples:
Free satellite image data @ University of Maryland:
It is also possible to BUY data! e.g. http://gs.mdacorporation.com/products/index.asp
Free Digital Terrain data:
Digitizing within a GIS
Sometimes digitizing is necessary if:
We need new to input new features.
Map features are incorrectly mapped.
Updates are needed for existing features.
And so, just what is digitizing?
Digitizing is the process of capturing map data in a GIS layer by tracing points, lines, or polygons from a map or image by using a mouse or a “puck” on on a digitizing tablet.
It can be done by either using a digitizing tablet, or by digitizing directly from a satellite image or a scanned map on screen.
On‐screen or “heads up” digitizing creates a spatial dataset by tracing over features displayed on a computer monitor with a mouse. The newly created dataset picks up the spatial reference of the source document and results in a string of points with (x, y) values. You will learn how to digitize in tutorial 6 of your analysis class but we will examine how we can do this now as well.
Digitizing table and PC workstation
A few years ago, the digitizing table was widely used in GIS. This method allows features to be input into a GIS
database as points, lines or polygons but often requires extensive data clean-up and editing. With the updates made to GIS software this method of data entry is essentially obsolete and has been replaced by on-screen digitizing.
Digitizing directly on screen
Digitizing directly on screen or “heads up” digitizing, is the approach I am most familiar with and it is what you will be doing in the analysis class.
In this case, you use your computer mouse to digitize paper maps, aerial photos, or other images displayed in the GIS. We can refer to this layer as a stable base map (that is geocoded) ‐ recall from slide 15 that the newly created dataset picks up
the spatial reference of the source document? Below is an example, where the outline of fluvial features (blue) and contour lines (brown) have been digitized. Only the fluvial features can have been digitized directly from the image on the left, however.
Selecting points to digitize
A vertex
Digitizing is not difficult to do. You can choose to digitize either a point, a line (snap to feature) or a polygon (snap to end). Snapping is a routine embedded within the GIS that makes sure the separate vertices in a line or an area are connected. That is, it ensures there are no gaps. This is important if calculations or correlations are going to be performed on the new layers. In these cases,
gaps would cause errors. The next slide shows how snapping can be achieved –it uses topology
Geocoding Address Data
What is it?
Geocoding address data is the process of relating an address to a geographic location (such as latitude/longitude coordinates) or geographic area (such as census tract, block group, block, or ZIP code). The address itself then is used to determine the geographical coordinates.
Geocoding can be affected by the quality of data, e.g., incorrect spelling and the use of different abbreviations (e.g., for Street and Avenue). Therefore, the use of standards is relevant to geocoding address data.
Address Matching
Address matching is the process of geocoding street addresses to a street network.
(Modified) ESRI’s definition of address matching:
A process that compares an address or a table of addresses to the address attributes of a reference dataset (e.g., locations may be determined based on address ranges stored for each street segment). This determines whether a particular address falls within an address range associated with a feature in the reference dataset.
If it does, it is considered a match and a location can be returned. The next slide provides an example of how an address can be converted to a point feature and drawn along a line segment based upon an address range.
Address Matching
In most instances a geocoding service or engine is built upon a street network in which line segments are assigned an address number range on the left and right side of the streets, a city or twin name, and a zip code. When a user enters an address the geocoding engine finds a match along a line segment and places the location of the address (a point) based upon where it falls along the street line segment.
GIS Services: Address Matching Resources
Creating a geocoding engine can be a time consuming process. However, there are some commercially available geocoding services that allow us to enter a table of addresses for geocoding
and then download the results and add them to our GIS database. Some geocoding services are free and some charge a fee.
https://geomap.ffiec.gov/FFIECGeocMap/GeocodeMap1.aspx https://www.census.gov/geo/maps-data/data/geocoder.html
Error, Accuracy, and Precision
Until quite recently, people involved in developing and using GIS paid little attention to the problems caused by error, inaccuracy, and imprecision in spatial datasets. Surely, a GIS is too powerful for error?
Not true! It is now recognized that error, inaccuracy, and imprecision, if left unchecked, can make the results of a GIS analysis almost worthless, it is like putting garbage in means getting garbage out.
But where do these errors come from?
Since a GIS can collate and cross‐reference many types of data by location and can integrate many discrete datasets (which is the heart of its power), it can also inherit error from the imported datasets.
17
Error, Accuracy, and Precision
We can discuss error in terms of data quality.
Data quality refers to the relative accuracy and precision of a GIS database. These facts are often documented in data quality reports.
Errors may exist both in map data (which can be reduced by maintaining topological integrity) and also attribute data.
Next we will discuss the differences between accuracy and precision in GIS data.
18
Data Accuracy
Accuracy is the degree to which information on a map, or in a database, matches true values. The level of accuracy required for particular applications and GIS analyses varies greatly.
Highly accurate data can be very difficult and costly to produce and maintain often requiring highly accurate GPS and/or survey data or the acquisition of LiDAR or remotely sends data (satellite imagery, aerial photographs, etc.).
In discussing a GIS database, it is possible to consider horizontal and vertical (spatial) accuracy with respect to geographic position (i.e. how close is the mapped feature to its actual location on the earths surface) as well as attribute accuracy (i.e. do the values in the Attribute table match the real world values).
19
Data Accuracy
1:1,200 ± 1 m
1:2,400 ± 2 m
1:4,800 ± 4 m
1:10,000 ± 8.5 m
1:12,000 ± 10 m
1:24,000 ± 12 m
1:63,360 ± 32 m
1:100,000 ± 50 m
This means that when we see a point on a map we have its "probable" location within a certain area. For example, a sewer manhole shown on a 1:1,200 scale map should be within 1 meter (3 ft) of it’s true location on the ground to be considered within the 1:1,200 accuracy standard. It’s important to understand the limitations of your data as well. While a layer of sewer manholes within the 1:1200 data standard may be suitable for broad planning or maintenance activities, using this data at scales it was not intended (engineering purposes for example) for can have negative and potentially costly consequences.
20
Defining Precision
Precision refers to the level of measurement and exactness of description in a GIS database.
The level of precision required for particular applications varies greatly. Engineering projects such as road and utility construction require very precise information measured
to the millimeter. Demographic analyses of marketing or electoral trends can often make do
with less. For example to the closest zip code or precinct boundary. Highly precise data can also be very difficult and costly to collect.
Precise data ‐ no matter how carefully measured ‐ may be inaccurate and vice-versa. Let’s try to understand this better by using Bolstad’s diagram on page 625, which I have included in the next slide.
21
Accuracy and Precision
Points (yellow circles) are digitized to represent the center of the cloverleaf intersection. Average accuracy is high when the average of the points falls near the true location, as in the panels on the left side of the figure. Precision is high when the points are all clustered near each other (top panels). A group of points may be accurate, but not precise (lower left), or precise, but not accurate (upper right). We typically strive for a process that provides both accuracy and precision (upper left), and avoid low accuracy and low precision (lower right).
22
Sources of Inaccuracy and Precision
Let us take a minute to discuss attribute accuracy and precision, that is the non‐spatial data linked to lo
cation.
Inaccuracies may result from mistakes of many sorts, including basic data entry. With respect to precision, precise attribute information describes phenomena in great detail. For example, a precise description of a person living at a particular address might include gender, age, income, occupation, level of education, and many other characteristics. A less precise description might include just income or gender. It is the application that will determine whether precise or less precise data are needed.
There are many sources of error that may affect the quality of a GIS dataset. Some are quite obvious, but others can be difficult to discern. Few of these will be automatically identified by the GIS itself.
It is the user's responsibility to prevent them.
Sources of error can be divided into 3 main categories:
Conceptual errors;
Errors arising through data processing; and
Errors arising from source data
23
Errors Arising from Source Data
Age of Data
Data sources may be too old to be useful or relevant to current GIS projects. Past collection standards may be unknown, non‐existent, or currently acceptable. For instance, John Wesley Powell's nineteenth century survey data of the Grand Canyon lacks the precision of data that can be developed and used today. Additionally, erosion, deposition, and other geomorphic processes will have modified the landscape. Therefore, reliance on old data could skew, bias, or negate results.
Density of Observations
The number of observations within an area is a guide to data reliability and they should be known by the map user. An insufficient number of observations may not provide the level of resolution required to adequately perform spatial analysis and determine the patterns GIS projects seek to resolve or define.
24
Errors Arising from Source Data
There are four ways we describe errors in spatial data
Positional accuracy – describes how close the locations of objects represented in a digital dataset correspond to the true locations of the real-world entities.
Attribute accuracy – summarizes how different the attributes are from their real world values.
Logical Consistency – reflects the presence, absence, or frequency of inconsistent data. Tests for logical consistency often require comparisons among themes (i.e. all buildings must be on dry land).
Completeness – Describes how well a layer reflects all of the real world features it is supposed to represent.
The next slide shows a figure from the text (Figure 14-3, Page 624) that depicts examples of these types of errors.
25
Errors Arising from Source Data
26
Data Standards
Regular checks and tests should be employed during a project to make sure that standards are being followed. This allows a designer to pinpoint difficulties at an early stage and correct them. Establishing data standards helps with data exchange –unfortunately, the history of GIS data exchange has been chaotic and has been wasteful in the past.
Examples of good data standards include:
USGS, National Mapping Program Standards, http://nationalmap.gov/standards//
Spatial Data Transfer Standard http://mcmcweb.er.usgs.gov/sdts/
27