QGIS only need Drawing method and why i choisen this area, don't need drawing the map.

profilemrryanwang
example_report_.docx

Via encouraging increasingly accessible, reliable and user-friendly ways to interrogate and visualise geographic data, open source GIS methods are providing enormous scope for greater and more widespread understanding of spatial patterns across a multitude of fields and applications (Steiniger and Bocher, 2009). The visualisations created as part of this assessment aim to demonstrate how a range of these open source tools and software can be utilised in order to analyse and present spatial data in ways where it can have an informative and meaningful impact upon its audience. The patterns and messages revealed by the visualisations will consequently be discussed and comment will be made upon any limitations of the maps created and challenges encountered whilst putting the GIS methods into practise.

The first geovisualisation (Figure 1) presented consists of a series of heat maps showing the spatial pattern of road traffic collisions (RTC) in the borough of Manhattan, New York City (NYC) over a period of three years. Using a series of small, similar maps offers the benefit of visually enforcing the comparison of changes within them (Tufte, 1991). With this in mind, the small multiples approach was the ideal method for visualising and comparing the different spatial patterns of RTC causes. The idea behind this visualisation was to identify areas of high collision concentration both overall and in terms of specific causes. Highlighting where these “hotspots” occur may help authorities to understand why high numbers of collisions are concentrated in certain locations and therefore help to formulate measures to improve road safety.

The dataset used in this visualisation was sourced from the NYC Open Data website and is a product of all RTC recorded in the New York Police Department database between the 7th January 2012 and the 31st December 2014. The data was downloaded in the form of a CSV file with georeferenced fields for the majority of cases included. Once downloaded, the data was imported to QGIS using the “Add delimited text layer” function with geometry defined using the latitude and longitude fields. This could then be converted to a shapefile consisting of 103,398 individual georeferenced points. By filtering this dataset using different criteria, I was able to create new shapefiles for a number of different causal factors. These, along with the original unfiltered shapefile could then be used to create individual heat maps with the QGIS Heatmap Plugin. Ensuring the heat maps effectively displayed patterns in the data required significant experimentation, a large part of which was finding a suitable radius. Too large a radius, for example, would distort the pattern of collisions by suggesting they occurred over excessively wide an area. Given that the overall dataset being used contained over 100,000 points covering an area of around 90km2, a fairly small radius of 300m was chosen.

The basemap used in the visualisations consisted of vector lines data sourced from the Open Street Map (OSM) project. This was downloaded as a KML file, converted to a shapefile and clipped to the Manhattan boundary. Use of the basemap was important as the location and size of roads provided some useful context to the information on collisions. A grayscale basemap was chosen as it showed up most effectively beneath the coloured raster surfaces of the heatmaps. Labels for a number of major roads were added for further context using an additional filtered OSM layer.

Observation of the heatmap for total collisions shows us that the highest concentrations of RTCs are located around major road intersections close to 60th Street, the Lincoln Tunnel and two major bridges on the south of the island. Generally, the lower half of the island appears to experience a greater number of collisions than the upper half. Hotspots for collisions involving slippery roads can be identified along Henry Hudson Parkway and also in the Noho area near to Williamsburg Bridge. Speeding collisions appear to have a distribution less centred on the busy Midtown area of the borough. For example, hotspots can be found in the Financial District and along Broadway on the Upper West Side. Finally, road-rage related collisions show the greatest concentration in busy areas just south of Central Park.

The RTCs visualisation demonstrates the potential of the QGIS heat map function as a means of identifying spatial patterns. However, major challenges exist with regard to the scale at which the analysis is being carried out and the consequent definition of suitable parameters such as heatmap radius. These challenges should be viewed with particular caution in light of the wider non-expert user base which open source GIS is making map production accessible to. In terms of the visualisation presented here, the main limitations relate to the level of detail shown. The decision to highlight collision patterns on a boroughwide scale means that to some extent, the precision of individual collision locations had to be sacrificed. One way of overcoming this problem of scale would be to present the data as an interactive web map with the ability to zoom in on specific point locations.

Figure 1. Road traffic collisions in Manhattan, NYC.

The second visualisation (Figure 2) displays the distribution of English Heritage listed buildings across London’s boroughs. The data involved consisted of over 18,000 individual points situated within the area of Greater London. Due to its suitability for summarising densities of large numbers of points (Graser, 2013,) a hexagonal binning approach was chosen to visualise the distribution of London’s listed buildings.

The data was sourced from the English Heritage website, located via a search using the government open data portal. After downloading the data as a CSV, it was imported into QGIS and converted into a shapefile. Following this, the georeferenced point layer was clipped using a polygon representing the Greater London boundary. This polygon was constructed by filtering a national layer of unitary authority boundaries. Using the “hex grid from layer bounds” function in the QGIS processing toolbox, a 2km2 hexagonal grid was created covering the extent of the listed buildings points which could then be used in a “count points in polygon” operation. Choosing a suitable resolution for this grid was key to achieving an effective visualisation. A number of different resolutions were tested and following experimentation I decided that 2km was sufficient resolution to display boroughlevel variations in listed building density without being so fine-grained that overall clarity was sacrificed. The new hexagonal grid layer containing information on listed building numbers could then be clipped to the Greater London boundary and styled using no outline and a graduated colour scheme based on a natural breaks classification. Borough outlines and labels were consequently added to provide context.

Figure 2 shows that higher concentrations of listed buildings generally exist in central areas of London. In terms of boroughs, the highest concentration appears to occur in Westminster. Outer boroughs show much lower concentrations of listed buildings, particularly in the south and east of the city. Barking and Dagenham seems to contain the least listed buildings in the city, with only three areas where the concentration of listed buildings per 2km2 exceeds 3. The apparent negative relationship between distance from London’s centre and concentration of listed buildings can be explained by a number of factors, one of which may be the age of buildings. The English Heritage website states that “the older a building is, the more likely it is to be listed”. London’s historical growth from the centre outwards may hence explain why more, older buildings and hence more, listed buildings lie near to the centre of the city. The pattern of low listed building concentrations in eastern boroughs of the city may, to an extent, represent the effects of bombing in the Second World War. However, perhaps the most important factor underlying the patterns of listed building concentration demonstrated is the general density of the built environment. The fact that the number of all buildings per 2km2 in central London is significantly higher than in outer boroughs logically increases the probability that the number of listed buildings will also be higher.

The hex-bin approach used in this visualisation demonstrates an alternative method to the heatmap when visualising point density. Significant caution was required when selecting a radius distance for the heatmaps in Figure 1 and this was again necessary when defining the area of hexagons used in Figure 2. An issue with regard to the aggregation of points to any larger areas, in this case a 2km2 hexagonal grid, is that an assumption of uniformity is made. This assumption rarely holds. Parallels can be drawn with examples in raster data, where, despite sub-pixel variations, a pixel must take on a single value to represent, e.g. a landcover type (Fisher, 1997). With the hex-bin map presented here, some variation in listed building density may occur below the level of the 2km2 hexagon and hence remain undetected. For example, five listed buildings may occur in one street within a hexagon yet the remaining area may be completely devoid of them.

Figure 2. English Heritage listed buildings, London.

The final visualisation presented consists of an interactive web map created using Google Fusion Tables:

https://www.google.com/fusiontables/DataSource?docid=18rPxFU67BsWooVrr23X0NdLRc W - nim04FZS6ucDO .

The map displays European countries using a choropleth map coloured on the basis of the number of Summer Olympic medals won (1896-2008) per 1000 members of their respective populations (Figure 3). More detail, including total medal count, estimated population and most successful Olympic Games is provided via an info-window which appears on the clicking of individual countries.

All data on Olympic medals was sourced from the International Olympic Committee (IOC) Research and Reference Service via The Guardian Data Blog website. International boundary data was downloaded as part of the Natural Earth admin countries dataset. The first step in creating the visualisation was to identify the data that I wished to include and to organise this in a suitable tabular format in Excel. This involved applying a number of filters to the IOC data, adding a field for estimated population and calculating a field for medals per 1000 people. Next, a unique identifier field was added to the table which matched a corresponding field in the geographical boundary data. Following the import of the medals data to QGIS as a CSV, this unique identifier field could be used to create a tabular join between the medals and boundary data. The resulting joined data could then be saved as a shapefile and a KML which could consequently be uploaded for use in Google Fusion Tables.

Styling of the data on Google Fusion Tables involved using “buckets” to control fill colour on the basis of values within the “medals per 1000” field. I experimented with different classification schemes in QGIS before deciding on a quantile classification using 8 classes and consequently copying the break values into Google Fusion Tables. A sequential colour scheme was chosen in QGIS and hex-codes copied in. The info-window was reorganised by editing the window’s HTML code. This involved selecting the fields to be included and altering the order, style and format in which they would be displayed.

The Olympic medals visualisation shows us that the most successful countries in terms of medals per 1000 include the Scandinavian countries as well as Hungary. Through use of the info-window, a possible trend of countries with smaller populations having more medals per person can be identified. A pattern of low Olympic success in recently formed countries such as the Czech Republic, Bosnia and Herzegovina and Serbia is clearly evident in the visualisation. This is a product of the lower number of Olympic Games they have taken part in compared to their older neighbours.

The process of producing the Olympic medals web map was extremely straightforward, firstly due to the flexibility of QGIS, exemplified by the conversion of data from shapefile to KML form. Once the data was uploaded, the functionality of Google Fusion Tables, for example the option to view and edit fields in the attribute table, enabled me to experiment with a number of different map options before deciding on my preferred choice. One aspect of Google Fusion Tables that I felt was limited involved the creation of classes in the styling of the map and consequent colour scheme selection. Creating an option to classify data via a choice of algorithms and improving the range of pre-loaded colour schemes would prevent the need to copy in break values and hex codes from QGIS. Significant time could be spent improving the Olympic medals web map, much of which would involve adding more data and further experimenting with display options in the info-window. For example, adding a field for most successful sport, images of national flags or perhaps a second layer representing success at Winter Olympics.

The three visualisations presented here: a series of heatmaps, a hex-bin map and an interactive web map, demonstrate some of a wide range of geo-spatial techniques that open source GIS software makes freely available and accessible. Despite the great potential that open source GIS offers, the techniques involved have been shown to have their limitations, particularly if being used by a non-GIS expert. However, via thoughtful use of these techniques, I have been able to highlight important spatial trends which further our understanding of the phenomena being studied and have the potential to inform future decision making.

Figure 3. Olympic medals visualisation. Screenshot from Google Fusion Tables.

References

English Heritage, Listed buildings. Available from: http://www.english heritage.org.uk/caring/listing/listed - buildings/ [last accessed: 18/01/2015]

Fisher, P. (1997). The pixel: a snare and a delusion. International Journal of Remote Sensing, 18(3), 679-685.

Graser, A. (2013). Learning QGIS 2.0. Packt Publishing Ltd.

Guardian Data Blog, If Michael Phelps were a country, how would his medal haul compare? Available from: http://www.theguardian.com/sport/datablog/2012/aug/01/if michael - phelps - were - a - country [last accessed: 18/01/2015]

IOC Research and Reference Service. Available from: http://www.olympic.org/content/the olympic - studies - centre/service - pages - container/the - research - and - reference - service/ [last accessed: 18/01/2015]

New York City Open Data, NYPD motor vehicle collisions. Available from:

https://data.cityofnewyork.us/NYC - BigApps/NYPD - Motor - Vehicle - Collisions/h9gi nx95 [last accessed: 18/01/2015]

Open Street Map. Available from: https://www.openstreetmap.org/#map=13/52.1441/ -

0.4683 [last accessed: 18/01/2015]

Steiniger, S., & Bocher, E. (2009). An overview on current free and open source desktop GIS developments. International Journal of Geographical Information Science, 23(10), 1345-1370.

Tufte, E. R. (1991). Envisioning information. Optometry & Vision Science, 68(4), 322-324.