Reading Summary and Analysis Discussion Post: Exploring Information Systems
1
1 Introduction to Library Databases
As they said in the Sound of Music, “let’s start at the very beginning.” In this case, a definition and some history, leading up to the current state of the database industry and its major players. Last, we’ll go over the recent development known as “Discovery Services” and explain why those systems are outside the scope of this book.
Electronic access to information by means of the Web is so pervasive that we take it for granted. You have undoubtedly already used library data- bases somewhere in your academic life, and either heard or tossed the word “database” around yourself. But where did these “databases” come from? Why are they important? What is a database, anyway? Let’s address that last question first, and then find out where they came from.
What Is a Database?
The Oxford English Dictionary defines “database” as: “A structured set of data held in computer storage and typically accessed or manipulated by means of specialized software.” (So much for not using any part of a word in its definition.) For “data,” let us substitute “information.” A database is a way to structure, store, and rapidly access huge amounts of information electronically. That “information” can be numerical or textual, even visual. And as the Encyclopedia of Computer Science (2003) notes: “An important feature of a good database is that unnecessary redundancy of stored data is avoided.” The key concepts are structure (an organized way to store the information, accomplished by tables, records, and fields, which are discussed later in this chapter), efficiency (no redundancy), and rapid access (the abil- ity to search and retrieve material from the database as quickly as possible).
C o p y r i g h t 2 0 1 5 . L i b r a r i e s U n l i m i t e d .
A l l r i g h t s r e s e r v e d . M a y n o t b e r e p r o d u c e d i n a n y f o r m w i t h o u t p e r m i s s i o n f r o m t h e p u b l i s h e r , e x c e p t f a i r u s e s p e r m i t t e d u n d e r U . S . o r a p p l i c a b l e c o p y r i g h t l a w .
EBSCO Publishing : eBook Collection (EBSCOhost) - printed on 2/4/2022 6:34 PM via UNIVERSITY OF BRITISH COLUMBIA AN: 1197818 ; Bell, Suzanne S..; Librarian's Guide to Online Searching: Cultivating Database Skills for Research and Instruction, 4th Edition : Cultivating Database Skills for Research and Instruction Account: s5672194.main.ehost
2 Librarian’s Guide to Online Searching
As my husband the computer scientist puts it, “a database isn’t magic but it is pretty smart.”
Is Google a Database?
Absolutely, in the sense that Google is a vast collection of data (the con- tents of web pages and material linked to web pages) that is searchable and provides rapid access to results. But Google and similar web search engines are not the focus of this textbook. These search tools build their databases automatically, from material that is freely accessible on the web, and their structure, scope, size, and many other aspects are not obvious. As far as one can tell, there is no quality control, no human intervention involved in build- ing the database.1
The commercial and governmental databases considered in this text are products specifically crafted to achieve the goal of providing users access to formally published information (e.g. articles, conference papers, books, dis- sertations, reports), in a very organized and efficient fashion. (In the case of commercial databases, part of that crafting is a mechanism for limiting access to paid subscribers.) Let us refer to these as “library databases,” since that is where you usually encounter them. Library databases tend to be targeted to specific audiences, and to offer customized features accordingly. Their struc- ture, scope, size, date coverage, publication list and many other details are either obvious from their search interfaces, or explicitly provided. (I would like to say that library databases are more structured than Google, but I can’t. Because who knows how Google is structured or how its search algorithm re- ally works? The “black boxness” of Google is another thing that distinguishes it from the databases that are the focus of this text.) Library databases are much less well known and ubiquitous than Google, and usually not free, but there are good reasons for that. Finding out where these databases came from should help explain why (as the adage goes, “you get what you pay for”).
Historical Background
Indexing and Abstracting Services
In the Beginning . . .
There was hard copy. Writers wrote, and their works were published in (physical) magazines, journals, newspapers, or conference proceedings. Months or years afterward, other writers, researchers, and other alert read- ers wanted to know what was written on a topic. Wouldn’t it be useful if there were a way to find everything that had been published on a topic, with- out having to page through every likely journal, newspaper, and so forth? It certainly would, as various publishing interests demonstrated: as early as 1848, the Poole’s Index to Periodical Literature provided “An alphabeti- cal index to subjects, treated in the reviews, and other periodicals, to which no indexes have been published; prepared for the library of the Brothers in Unity, Yale college” (Figure 1.1).
EBSCOhost - printed on 2/4/2022 6:34 PM via UNIVERSITY OF BRITISH COLUMBIA. All use subject to https://www.ebsco.com/terms-of-use
Introduction to Library Databases 3
Not to be caught napping, the New York Times started publishing their Index in 1851, and in 1896, taking a page from Poole’s, the Cumulative In- dex to a Selected List of Periodicals appeared, which soon (1904) became the canonical Readers’ Guide to Periodical Literature. Thus, in the mid-19th century, the hard copy Index is born: an alphabetical list of words, represent- ing subjects, and under each word a list of articles deemed to be about that subject. The index is typeset, printed, bound, and sold, and all of this effort is done, slowly and laboriously, by humans.
Given the amount of work involved, and the costs of paper, printing, etc., how many subjects do you think an article would have been listed under? Every time the article entry is repeated under another subject, it costs the publisher just a little more. Suppose there was an article about a polar expedi- tion, which described the role the sled dogs played, the help provided by the native Inuit, the incompetence on the part of the provisions master, and the fund-raising efforts carried on by the leader’s wife back home in England. The
Figure 1.1. Index listing and title page (inset) from Poole’s Index to Periodical Literature. Courtesy of the Department of Rare Books and Special Collections, University of Rochester Li- braries, Rochester, NY.
EBSCOhost - printed on 2/4/2022 6:34 PM via UNIVERSITY OF BRITISH COLUMBIA. All use subject to https://www.ebsco.com/terms-of-use
4 Librarian’s Guide to Online Searching
publisher really can’t afford to list this article in more than one or two places. Under which subject(s) will people interested in this topic be most likely to look? The indexer’s career was a continuous series of such difficult choices.
An index, recording that an article exists and where it would be found, was a good start, but one could go a step further. The addition of a couple of sentences (e.g., to give the user an idea of what the article is about) increases the usefulness of the finding tool enormously—although the added informa- tion, of course, costs more in terms of space, paper, effort, etc. But some index publishers started adding abstracts, gambling that their customers would pay the higher price (which they did). Thus, we have the advent of Abstract- ing and Indexing services or “A & I,” terminology that you may still see in the library literature.
The abstracts were all laboriously written by humans. They needed to be skilled, literate humans, and skilled humans are very expensive (even when they’re underpaid, they are expensive in commercial terms). Humans are also slow, compared with technology. Paper and publishing are expensive, too. Given all this, how many times do you think an entry for an article would be duplicated (appear under multiple subjects) in this situation? The an- swers are obvious; the point is that the electronic situation we have today is all grounded in a physical reality. Once it was nothing but people and paper.
From Printed Volumes to Databases
Enter the Computer
The very first machines that can really be called digital computers were built in the period from 1939 to 1944, culminating in the construction of the ENIAC in 1946, “the first general-purpose, electronic computer” (Encyclopæ- dia Britannica Online 2014). These machines were all part of a long pro- gression of innovations to speed up the task of mathematical calculations. After that, inventions and improvements came thick and fast: the 1950s and 1960s were an incredibly innovative time in computing, although probably not in a way that the ordinary person would have noticed. The first ma- chine to be able to store a database, RCA’s Bizmac, was developed in 1952 (Lexikon’s History of Computing 2002). The first instance of an online trans- action-processing system, using telephone lines to connect remote users to central mainframe computers, was the airline reservation system known as SABRE, set up by IBM for American Airlines in 1964 (Computer His- tory Museum 2004). Meanwhile, at Lockheed Missile and Space Company, a man named Roger Summit was engaged in projects involving search and re- trieval, and management of massive data files. His group’s first interactive search-and-retrieval service was demonstrated to the company in 1965; by 1972, it had developed into a new, commercially viable product: Dialog—the “first publicly available online research service” (Dialog 2005).
Thus, in the 1960s and 1970s, when articles were still being produced on typewriters, indexes and abstracts were being produced in hard copy, and very disparate industries were developing information technologies for their own specialized purposes, Summit can be credited with having incredible vision. He asked the right questions:
EBSCOhost - printed on 2/4/2022 6:34 PM via UNIVERSITY OF BRITISH COLUMBIA. All use subject to https://www.ebsco.com/terms-of-use
Introduction to Library Databases 5
1. What do people want? Information.
2. Who produces information, and in what form? The government and commercial publishers, in the form of papers, articles, news- papers, etc.
3. What if you could put information about all that published mate- rial into a machine-readable file: a database—something you could search?
Summit also had the vision to see how the technological elements could be used. The database needed to be made only once, at his firm’s headquar- ters, and trained agents (librarians) could then access it over telephone lines with just some simple, basic equipment. The firm could track usage exactly and charge accordingly. Think of the advantages!
The advantages of an electronic version of an indexing/abstracting sys- tem are really revolutionary. In a system no longer bound by the confines of paper, space, and quite so many expensive skilled personnel:
• Articles could be associated with a greater number of terms describing their content, not just one or two (some skilled labor is still required).
• Although material has to be rekeyed (i.e., typed into the database), this doesn’t require subject specialists, simply typists (cheap labor).
• Turnaround time is faster: most of your labor force isn’t thinking and composing, just typing continuously—the process of adding to the information in the database goes on all the time, making the online product much more current.
• If you choose to provide your index “online only,” thus avoiding the time delays and costs of physical publishing, why, you might be able to redirect the funds to expanding your business: offering other in- dexes (databases) in new subject areas.
As time goes on, this process of “from article to index” gets even faster. When articles are created electronically (e.g., word processing), no rekeying is needed to get the information into your database, just software to convert and rearrange the material to fit your database fields. So, rather than typ- ists, you must pay programmers to write the software, and you still need some humans to analyze the content and assign the subject terms.
In the end, the electronic database is not necessarily cheaper to create; it very likely costs more! The costs have simply shifted. But customers buy it because . . . it is so much more powerful and efficient. It is irresistible, and printed indexes have vanished like the dodo. Online library databases are an integral part of the research process.
The Library Database Industry Today
For a line of business and a product you probably weren’t very aware of until you were in high school or college, the library database business is,
EBSCOhost - printed on 2/4/2022 6:34 PM via UNIVERSITY OF BRITISH COLUMBIA. All use subject to https://www.ebsco.com/terms-of-use
6 Librarian’s Guide to Online Searching
for the moment, surprisingly robust. The juggernaut of Google and espe- cially Google Scholar has not put the commercial database vendors out of business (yet—I’m sure there is a constant undercurrent of fear throughout the business). Probably the largest commercial vendors, and ones you might have heard of before reading this book, are EBSCO, ProQuest, and Gale (Gale Cengage). Other major names to add to your repertoire are Thomson- Reuters (creators of the Web of Science and many other databases), JSTOR, LexisNexis, OCLC FirstSearch, ABC-CLIO, Alexander Street Press, Project MUSE, and OVID. Most of the databases produced by these vendors have content drawn from many sources, many publishers: they aggregate content, bringing it together so you can search across all of it in one database. Thus the terminology aggregators that is often used to describe the multidisci- plinary article databases from the vendors listed above. In contrast, major publishers such as Elsevier, Oxford University Press, and Sage Publications are big enough to create databases just of the materials they publish, for ex- ample: Elsevier’s ScienceDirect database, Oxford Music Online, Sage Jour- nals and the CQ databases (an imprint of Sage).
In addition to the commercial entities mentioned above, some profes- sional associations create and manage the subscriptions to databases of their materials. Examples include the Association for Computing Machin- ery (ACM), the Institute of Electrical and Electronics Engineers (IEEE), the American Society of Mechanical Engineers (ASME), the American Math- ematical Society (AMS), and the American Chemical Society (ACS). Govern- ment and international agencies also produce databases. US government agencies such as the National Library of Medicine, the Department of Edu- cation, the Census Bureau, and the Bureau of Labor Statistics are the au- thors of key databases in their respective topical areas, which we will cover in subsequent chapters. At the international level, the World Bank,2 the In- ternational Monetary Fund, and the Organization for Economic Cooperation and Development (OECD) all offer databases of their information.
The names of library database vendors listed above represent only the largest and/or better-known entities. As in any line of business there are, of course, many more companies, either smaller or focused on a particular audience (the number of vendors that create databases specifically for the business community, both corporate and academic, is remarkably exten- sive). The database vendor industry is also a business like any other: it is subject to consolidation and occasionally to expansion. Companies come and go through mergers and acquisitions, start-ups, and occasional deaths. Changes may not happen as rapidly as in some industries, but when they do, they can be significant. Three of the notable changes in the current decade were EBSCO’s acquisition of the H.W. Wilson databases, and two major moves by ProQuest: the acquisition of the CSA databases and taking over publication of the Statistical Abstract of the United States from the US Census Bureau, including putting all the Statistical Abstract content into a new database.
At the beginning of this section, I made a reference to the database vendors’ (not to mention librarians’) fears about Google and Google Scholar: that these free, ubiquitous, embedded-in-daily-life resources might spell the end of the library database business. The vendors have been fighting
EBSCOhost - printed on 2/4/2022 6:34 PM via UNIVERSITY OF BRITISH COLUMBIA. All use subject to https://www.ebsco.com/terms-of-use
Introduction to Library Databases 7
back for many years, however, first with something called “federated search” (about which the less said the better; the title of Jody Fagan’s 2011 editorial on the topic says it all: “Federated Search Is Dead—and Good Riddance!”). The latest counter-attack by the database vendors, dubbed “Discovery Ser- vices” or “Web Scale Discovery Services,” is far superior. Reports on usage statistics from institutions that have adopted a discovery service indicate that these products may have a strong chance of winning ground back from the all-mighty Googleplex (Way 2010, Kemp 2012, Daniels, Robinson, and Wishnetsky 2013, Calvert 2014).
The following section will provide a brief overview of discovery services, concluding with why they will not be considered further in this text.
Discovery Services
Discovery Services are systems that harvest and pre-index a wide vari- ety of library content from separate sources (records from library databases, the online catalog, perhaps the local institutional repository or other locally developed databases), build one giant index of all that content, and provide near-instant, relevancy ranked results through one search box (Vaughan 2011, Adams et al. 2013). Sound familiar? It is exactly the Google model, but instead of web pages it draws on all the vetted and expensive resources for which the library has already paid, making them “discoverable.” These sys- tems are frequently referred to as Web-Scale Discovery Services, “meaning they search library collections the way Google searches the web: by search- ing the entire breadth of content available in the library’s collection” (Fry 2013). (The “entire breadth of content” is at least the goal if not the reality right now.) The essential key is the pre-indexing, getting the data from all the disparate resources ahead of time, as it were, to build that one giant index that can provide the speedy response time that users expect. Where the discovery systems start to part ways with Google is on the results page, which is loaded with options for refining and outputting results, and where library-owned full text is instantly accessible.
The vendors and products in the discovery service market at the time of this writing are EBSCO Discovery Service (EDS), Serials Solutions’ Summon (note that Serials Solutions is owned by ProQuest), Ex Libris’ PrimoCentral, OCLC’s WorldCat Discovery Services (WDS), and, though it works differ- ently from the others, Innovative Interface’s Encore Synergy. AquaBrowser is ProQuest’s discovery product aimed at the public library market.
The tricky part is that these companies are competitors both in the individual database and now in the discovery service market. Their major customers (large academic libraries) have resources from a wide variety of vendors. Achieving the goal of providing “one search” access to all that con- tent means that each discovery service company (A) must persuade the com- peting discovery service companies (B), and all the other database vendors (C), to give A access to their databases in order to harvest and pre-index the data therein. This is a delicate dance, as you might imagine, but again, the threat of Google is actually helping, and agreements are (carefully) be- ing negotiated. From a customer’s point of view, it’s obvious: “You’ve got to be in,” says Michael Kucsak, Director of Library Systems & Technology at
EBSCOhost - printed on 2/4/2022 6:34 PM via UNIVERSITY OF BRITISH COLUMBIA. All use subject to https://www.ebsco.com/terms-of-use
8 Librarian’s Guide to Online Searching
the University of North Florida, talking about inter-vendor discoverability. “You’re in—you win. You’re out—you’re not long for the world” (Fry 2013).
Discovery services hold immense promise for breaking down the silos in library content, especially the one between the library catalog (OPAC) and the article databases (each one of which is in its own silo). While students may eventually understand that the catalog and the databases are separate, and that the routes for accessing each one are different, for the casual user who needs 15 good articles for tomorrow’s paper—it is simply too much ef- fort. Discovery services meet that need fairly efficiently and painlessly. And as a librarian who has made many, many purchase decisions and is painfully aware of what quality database resources cost, a “tool [that] holds the poten- tial to significantly increase the discovery and use of such content” (Vaughan 2011) does indeed get my notice and my vote.
So how can a textbook on (individual) database searching still be justi- fied? Why master all sorts of esoteric knowledge and get comfortable with interfaces having three search boxes (with attendant options and settings) when there is a simple, one-box option that searches the same material? The discovery services tools are a wonderful way to woo undergraduates back to library resources. But you have this textbook in hand, presumably, because you are studying to become a librarian or an information professional or technologist. For you, a higher order of knowledge and familiarity with more sophisticated tools and approaches is one of the essential points—otherwise anyone could set up shop and call herself an expert searcher. Google and the discovery services will take care of the lower order questions. Someone still needs to be there to deal with the harder, higher order research queries. When the discovery service search isn’t providing the answer, someone needs to know how to go to next level: how to choose, access, and skillfully interact with highly crafted, subject-specific databases on an individual basis. Jody Fagan (2011) points out that “scholars working on more substantial research projects . . . have already found—or will need to find—the native interface to the subject-specific resources they need.” You need to be the person who can point those scholars to subject-specific resources, and help them get the most out of the “native interface” (which usually provides subject-specific fea- tures) of those resources.3 This book is designed to do precisely that. Let’s get started—because searching really can be just as rewarding as finding.
Notes
1. According to the Google Guide at http://www.googleguide.com/google_works.html, the Google database is built by the GoogleBot, indexed by the Google Indexer, and searches handled by the three parts of the Query Processor. The utterly massive scale simply precludes any kind of human involvement.
2. Worth noting, the World Bank databases, formerly subscription-based, are now avail- able to the world for free. Kudos to the World Bank for this daring and generous move!
3. Besides, it’s just ever so much more interesting. What fun is plunking words in a box? Trust me, database skills make research much more efficient and satisfying.
EBSCOhost - printed on 2/4/2022 6:34 PM via UNIVERSITY OF BRITISH COLUMBIA. All use subject to https://www.ebsco.com/terms-of-use