Foundations of Geographic Information Systems Final Project

profileusaggwp
Week5-DataInput...FA19GIS5103.pptx

Foundations of Geographic Information Science

GIS5103 – Fall 2019

Week 5

Data Input, Throughput, & Quality

By the end of this week you should be able to:

Describe methods of data input

Describe Digitizing sources ‐ the cartographic base

Contrast Digitizing methods within a GIS (Tablet vs. “Heads Up”)

Describe Data editing (correcting digitizing errors, re‐projections

and generalizations)

Define accuracy and precision

Describe the importance of data standards and metadata

List sources of errors in a GIS project

Differentiate spatial and attribute errors

Describe techniques for coping with errors in Data

The Data Stream

The focus of this week’s lecture component is on creating GIS or spatial data from a variety of

sources. We will also be looking at how GIS data is updated.

Digitizing GIS Data

As we have discussed over the past several weeks one of the primary features of a GIS is its ability to integrate data from a variety of sources. Depending on how the data is to be used the need to convert the data or extract features from it may arise.

Fortunately there are a number of techniques and tools available to assist in acquiring data for

a new GIS is straightforward these days. Some of the sources for this data can include some of the following:

GPS (Global Positioning Systems) has become a major source of new GIS data. Data acquired through GPS can be highly spatially accurate (1cm horizontal) depending upon the type of receiver used.

Digital map images such as scanned maps, LiDAR , satellite data, and aerial photographs (Ortho photos) are now often used as a cartographic base for digitizing; this is another way to create new GIS data layers.

Remotely sensed satellite data are becoming an important source of GIS data as the cost of data falls, free data becomes available and new types of data emerge.

The following examples show the different spatial scales commonly found

with remote sensing data. Examples are also provided that show how they can be used to digitize temporal changes or changes to features over time (urbanization) or after an event (Natural Disaster).

Satellite data are raster data. The amount of spatial detail represented by one pixel of a raster

image determines its spatial resolution. So, an image with one meter spatial resolution means

that each pixel in the image represents one square meter on the ground. With improvements in

cameras and digital image acquisition the cost of acquiring aerial images has decreased. It is now possible to obtain imagery with a 6” pixel resolution for a few thousand dollars.

30 meters

10 meters

5 meters

The Different Spatial Resolutions

Image Resolutions Continued

2 meters

1 foot

In cartographic applications, higher spatial resolutions are often favored. However, for applications such as weather forecasting or to monitor seasonal plant growth on a

yearly basis for North America, coarser spatial resolutions are favored. While this type of data

can be extremely useful there are data storage considerations and large file sizes to consider which may impact the accessibility of these images especially in web based applications.

1 meter

Where to obtain remote sensing data

There are a number of useful websites from which remote

sensing data (satellite images) can be downloaded. Here are some examples:

Free satellite image data @ University of Maryland:

http://glcf.umd.edu/data/

It is also possible to BUY data! e.g. http://gs.mdacorporation.com/products/index.asp

Free Digital Terrain data:

http://www.webgis.com/terraindata.html

Digitizing within a GIS

Sometimes digitizing is necessary if:

We need new to input new features.

Map features are incorrectly mapped.

Updates are needed for existing features.

And so, just what is digitizing?

Digitizing is the process of capturing map data in a GIS layer by tracing points, lines, or polygons from a map or image by using a mouse or a “puck” on on a digitizing tablet.

It can be done by either using a digitizing tablet, or by digitizing directly from a satellite image or a scanned map on screen.

On‐screen or “heads up” digitizing creates a spatial dataset by tracing over features displayed on a computer monitor with a mouse. The newly created dataset picks up the spatial reference of the source document and results in a string of points with (x, y) values. You will learn how to digitize in tutorial 6 of your analysis class but we will examine how we can do this now as well.

Digitizing table and PC workstation

A few years ago, the digitizing table was widely used in GIS. This method allows features to be input into a GIS

database as points, lines or polygons but often requires extensive data clean-up and editing. With the updates made to GIS software this method of data entry is essentially obsolete and has been replaced by on-screen digitizing.

Digitizing directly on screen

Digitizing directly on screen or “heads up” digitizing, is the approach I am most familiar with and it is what you will be doing in the analysis class.

In this case, you use your computer mouse to digitize paper maps, aerial photos, or other images displayed in the GIS. We can refer to this layer as a stable base map (that is geocoded) ‐ recall from slide 15 that the newly created dataset picks up

the spatial reference of the source document? Below is an example, where the outline of fluvial features (blue) and contour lines (brown) have been digitized. Only the fluvial features can have been digitized directly from the image on the left, however.

Selecting points to digitize

A vertex

Digitizing is not difficult to do. You can choose to digitize either a point, a line (snap to feature) or a polygon (snap to end). Snapping is a routine embedded within the GIS that makes sure the separate vertices in a line or an area are connected. That is, it ensures there are no gaps. This is important if calculations or correlations are going to be performed on the new layers. In these cases,

gaps would cause errors. The next slide shows how snapping can be achieved –it uses topology

Geocoding Address Data

What is it?

Geocoding address data is the process of relating an address to a geographic location (such as latitude/longitude coordinates) or geographic area (such as census tract, block group, block, or ZIP code). The address itself then is used to determine the geographical coordinates.

Geocoding can be affected by the quality of data, e.g., incorrect spelling and the use of different abbreviations (e.g., for Street and Avenue). Therefore, the use of standards is relevant to geocoding address data.

Address Matching

Address matching is the process of geocoding street addresses to a street network.

(Modified) ESRI’s definition of address matching:

A process that compares an address or a table of addresses to the address attributes of a reference dataset (e.g., locations may be determined based on address ranges stored for each street segment). This determines whether a particular address falls within an address range associated with a feature in the reference dataset.

If it does, it is considered a match and a location can be returned. The next slide provides an example of how an address can be converted to a point feature and drawn along a line segment based upon an address range.

Address Matching

In most instances a geocoding service or engine is built upon a street network in which line segments are assigned an address number range on the left and right side of the streets, a city or twin name, and a zip code. When a user enters an address the geocoding engine finds a match along a line segment and places the location of the address (a point) based upon where it falls along the street line segment.

GIS Services: Address Matching Resources

Creating a geocoding engine can be a time consuming process. However, there are some commercially available geocoding services that allow us to enter a table of addresses for geocoding

and then download the results and add them to our GIS database. Some geocoding services are free and some charge a fee.

http://www.batchgeo.com

https://geomap.ffiec.gov/FFIECGeocMap/GeocodeMap1.aspx https://www.census.gov/geo/maps-data/data/geocoder.html

Error, Accuracy, and Precision

Until quite recently, people involved in developing and using GIS paid little attention to the problems caused by error, inaccuracy, and imprecision in spatial datasets. Surely, a GIS is too powerful for error?

Not true! It is now recognized that error, inaccuracy, and imprecision, if left unchecked, can make the results of a GIS analysis almost worthless, it is like putting garbage in means getting garbage out.

But where do these errors come from?

Since a GIS can collate and cross‐reference many types of data by location and can integrate many discrete datasets (which is the heart of its power), it can also inherit error from the imported datasets.

17

Error, Accuracy, and Precision

We can discuss error in terms of data quality.

Data quality refers to the relative accuracy and precision of a GIS database. These facts are often documented in data quality reports.

Errors may exist both in map data (which can be reduced by maintaining topological integrity) and also attribute data.

Next we will discuss the differences between accuracy and precision in GIS data.

18

Data Accuracy

Accuracy is the degree to which information on a map, or in a database, matches true values. The level of accuracy required for particular applications and GIS analyses varies greatly.

Highly accurate data can be very difficult and costly to produce and maintain often requiring highly accurate GPS and/or survey data or the acquisition of LiDAR or remotely sends data (satellite imagery, aerial photographs, etc.).

In discussing a GIS database, it is possible to consider horizontal and vertical (spatial) accuracy with respect to geographic position (i.e. how close is the mapped feature to its actual location on the earths surface) as well as attribute accuracy (i.e. do the values in the Attribute table match the real world values).

19

Data Accuracy

1:1,200 ± 1 m

1:2,400 ± 2 m

1:4,800 ± 4 m

1:10,000 ± 8.5 m

1:12,000 ± 10 m

1:24,000 ± 12 m

1:63,360 ± 32 m

1:100,000 ± 50 m

This means that when we see a point on a map we have its "probable" location within a certain area. For example, a sewer manhole shown on a 1:1,200 scale map should be within 1 meter (3 ft) of it’s true location on the ground to be considered within the 1:1,200 accuracy standard. It’s important to understand the limitations of your data as well. While a layer of sewer manholes within the 1:1200 data standard may be suitable for broad planning or maintenance activities, using this data at scales it was not intended (engineering purposes for example) for can have negative and potentially costly consequences.

20

Defining Precision

Precision refers to the level of measurement and exactness of description in a GIS database.

The level of precision required for particular applications varies greatly. Engineering projects such as road and utility construction require very precise information measured

to the millimeter. Demographic analyses of marketing or electoral trends can often make do

with less. For example to the closest zip code or precinct boundary. Highly precise data can also be very difficult and costly to collect.

Precise data ‐ no matter how carefully measured ‐ may be inaccurate and vice-versa. Let’s try to understand this better by using Bolstad’s diagram on page 625, which I have included in the next slide.

21

Accuracy and Precision

Points (yellow circles) are digitized to represent the center of the cloverleaf intersection. Average accuracy is high when the average of the points falls near the true location, as in the panels on the left side of the figure. Precision is high when the points are all clustered near each other (top panels). A group of points may be accurate, but not precise (lower left), or precise, but not accurate (upper right). We typically strive for a process that provides both accuracy and precision (upper left), and avoid low accuracy and low precision (lower right).

22

Sources of Inaccuracy and Precision

Let us take a minute to discuss attribute accuracy and precision, that is the non‐spatial data linked to lo

cation.

Inaccuracies may result from mistakes of many sorts, including basic data entry. With respect to precision, precise attribute information describes phenomena in great detail. For example, a precise description of a person living at a particular address might include gender, age, income, occupation, level of education, and many other characteristics. A less precise description might include just income or gender. It is the application that will determine whether precise or less precise data are needed.

There are many sources of error that may affect the quality of a GIS dataset. Some are quite obvious, but others can be difficult to discern. Few of these will be automatically identified by the GIS itself.

It is the user's responsibility to prevent them.

Sources of error can be divided into 3 main categories:

Conceptual errors;

Errors arising through data processing; and

Errors arising from source data

23

Errors Arising from Source Data

Age of Data

Data sources may be too old to be useful or relevant to current GIS projects. Past collection standards may be unknown, non‐existent, or currently acceptable. For instance, John Wesley Powell's nineteenth century survey data of the Grand Canyon lacks the precision of data that can be developed and used today. Additionally, erosion, deposition, and other geomorphic processes will have modified the landscape. Therefore, reliance on old data could skew, bias, or negate results.

Density of Observations

The number of observations within an area is a guide to data reliability and they should be known by the map user. An insufficient number of observations may not provide the level of resolution required to adequately perform spatial analysis and determine the patterns GIS projects seek to resolve or define.

24

Errors Arising from Source Data

There are four ways we describe errors in spatial data

Positional accuracy – describes how close the locations of objects represented in a digital dataset correspond to the true locations of the real-world entities.

Attribute accuracy – summarizes how different the attributes are from their real world values.

Logical Consistency – reflects the presence, absence, or frequency of inconsistent data. Tests for logical consistency often require comparisons among themes (i.e. all buildings must be on dry land).

Completeness – Describes how well a layer reflects all of the real world features it is supposed to represent.

The next slide shows a figure from the text (Figure 14-3, Page 624) that depicts examples of these types of errors.

25

Errors Arising from Source Data

26

Data Standards

Regular checks and tests should be employed during a project to make sure that standards are being followed. This allows a designer to pinpoint difficulties at an early stage and correct them. Establishing data standards helps with data exchange –unfortunately, the history of GIS data exchange has been chaotic and has been wasteful in the past.

Examples of good data standards include:

USGS, National Mapping Program Standards, http://nationalmap.gov/standards//

Spatial Data Transfer Standard http://mcmcweb.er.usgs.gov/sdts/

27