U10A1-68 - Qualitative Research Plan - ***TUTOR FOLLOW ALL INSTRUCTIONS AS OUTLINED TO COMPLETE THIS WORK. READ ATTACHMENTS.

profiledrcdopen82
Chaper9-EnhancingtheQualityandCredibilityofQualitativeStudiespt3.pdf

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 1/36

To generalize is to be an idiot. To particularize is the lone distinction of merit. General knowledges are those that idiots possess.

Stake (1978) continues,

Generalization may not be all that despicable, but particularization does deserve praise. To know particulars fleetingly, of course, is to know next to nothing. What becomes useful understanding is a full and thorough knowledge of the particular, recognizing it also in new and foreign contexts. That knowledge is a form of generalization too, not scientific induction but naturalistic generalization, arrived at by recognizing the similarities of objects and issues in and out of context and by sensing the natural covariations of happenings. To generalize this way is to be both intuitive and empirical, and not idiotic. (p. 6)

Stake (2000) extends naturalistic generalizations to include the kind of learning that readers take from their encounters with specific case studies. The “vicarious experience” that comes from reading a rich case account can contribute to the social construction of knowledge, which, in a cumulative sense, builds general, if not necessarily generalizable, knowledge.

Readers assimilate certain descriptions and assertions into memory. When researcher’s narrative provides opportunity for vicarious experience, readers extend their memories of happenings. Naturalistic, ethnographic case materials, to some extent, parallel actual experience, feeding into the most fundamental processes of awareness and understanding . . . [to permit] naturalistic generalizations. The reader comes to know some things told, as if he or she had experienced it. Enduring meanings come from encounter, and are modified and reinforced by repeated encounter.

In life itself, this occurs seldom to the individual alone but in the presence of others. In a social process, together they bend, spin, consolidate, and enrich their understandings. We come to know what has happened partly in terms of what others reveal as their experience. The case researcher emerges from one social experience, the observation, to choreograph another, the report. Knowledge is socially constructed, so we constructivists believe, and, in their experiential and contextual accounts, case study researchers assist readers in the construction of knowledge. (p. 442)

Guba (1978) considered three alternative positions that might be taken in regard to the generalizability of naturalistic inquiry findings:

1.Generalizability is a chimera; it is impossible to generalize in a scientific sense at all. . . .

2.Generalizability continues to be important, and efforts should be made to meet normal scientific criteria that pertain to it. . . .

3.Generalizability is a fragile concept whose meaning is ambiguous and whose power is variable. (pp. 68–70)

Having reviewed these three positions, Guba (1978) proposed a resolution that recognizes the diminished value and changed meaning of generalizations and echoes Cronbach’s emphasis, cited above, on treating conclusions as hypotheses for future applicability and testing rather than as definitive.

The evaluator should do what he can to establish the generalizability of his findings. . . . Often naturalistic inquiry can establish at least the “limiting cases” relevant to a given situation. But in the spirit of naturalistic inquiry he should regard each possible generalization only as a working hypothesis, to be tested again in the next encounter and again in the encounter after that. For the naturalistic inquiry evaluator, premature closure is a cardinal sin, and tolerance of ambiguity a virtue. (p. 70)

Guba and Lincoln (1981) emphasized appreciation of and attention to context as a natural limit to naturalistic generalizations. They ask, “What can a generalization be except an assertion that is context free? [Yet] it is virtually impossible to imagine any human behavior that is not heavily mediated by the context in which it occurs” (p. 62). They proposed substituting the concepts “transferability” and “fittingness” for generalization when dealing with qualitative findings:

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 2/36

The degree of transferability is a direct function of the similarity between the two contexts, what we shall call “fittingness.” Fittingness is defined as degree of congruence between sending and receiving contexts. If context A and context B are “sufficiently” congruent, then working hypotheses from the sending originating context may be applicable in the receiving context. (Lincoln & Guba, 1985, p. 124)

Cronbach (1980) offered a middle ground in the debate over generalizability. He found little value in experimental designs that are so focused on carefully controlling cause and effect (internal validity) that the findings are largely irrelevant beyond that highly controlled experimental situation (external validity). On the other hand, he was equally concerned about entirely idiosyncratic case studies that yield little of use beyond the case study setting. He was also skeptical that highly specific empirical findings would be meaningful under new conditions. He suggested instead that designs balance depth and breadth, realism and control so as to permit reasonable “extrapolation” (pp. 231–235).

SIDEBAR

TESTING THEORY FROM A PURPOSEFUL SAMPLE OF QUALITATIVE CASES TO GENERALIZE: A CLASSIC CASE EXAMPLE

Sociologist Alfred Lindesmith (1905–1991), Indiana University, wanted to test his theory about addiction to opiate drugs. The theory posited that people became addicted to opium, morphine, or heroin when they took the drug often enough and in sufficient quantity to develop physical withdrawal. But Lindesmith had observed that people become habituated to opiates in a hospital when medicated for pain and manifest junkie behavior of compulsively searching for drugs at almost any cost after hospitalization. He hypothesized that two other things had to happen: Having become habituated, the potential addict now had to (1) stop using drugs and experience the painful withdrawal symptoms that resulted and (2) consciously connect withdrawal distress with ceasing drug use, a connection not everyone made. Junkies, unlike former hospital patients, then had to act on that realization and take more drugs to relieve the symptoms. Those steps, taken together and taken repeatedly, create the compulsive activity that is addiction.

A well-known statistician criticized Lindesmith’s sample because he had generalized to a large population (all the addicts in the United States or in the world) from a small, purposefully selected sample rather than studying a random sample. Lindesmith replied that the purpose of random sampling was to ensure that every case had a known probability of being drawn for a sample and that researchers randomize to permit generalizations about distributions of some phenomenon in a population and in subgroups in a population. But, he argued, random sampling was irrelevant to his research on addicts because he was interested not in distributions but in a universal process—how one became and remained an addict. He didn’t want to know the probability that any particular case would be chosen for his sample. He wanted to maximize the probability of finding a negative case so as all the better to test the theory. Not finding disconfirming cases strengthened his confidence in generalizing his findings.

—Adapted from Becker (1998, pp. 86–87)

Extrapolation Unlike the usual meaning of the term generalization, an extrapolation clearly connotes that one has gone beyond the narrow confines of the data to think about other applications of the findings. Extrapolations are modest speculations on the likely applicability of findings other situations under similar, but not identical, conditions. Extrapolations are logical, thoughtful, case derived and problem oriented rather than statistical and probabilistic.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 3/36

Distinguished methodologist Thomas D. Cook (2014) has explained the nature and significance of extrapolation.

Informing future policy decisions also requires justified procedures for extrapolating past findings to future periods when the populations of treatment providers and recipients might be different, when adaptations of a previously studied treatment might be required, when a novel outcome is targeted, when the application might be to situations different from earlier, and when other factors affecting the outcome are novel too. We call this the extrapolation function since inferences are required about populations and categories that are now in some ways different from the sampled study particulars. Sampling theory cannot even pretend to deal with the framing of causal generalization as extrapolation since the emphasis is on taking observed causal findings and projecting them beyond the observed sampling specifics.

We argue here that both representation and extrapolation are part of a broad and useful understanding of external validity; that each has been quite neglected in the past relative to internal validity—namely, whether the link between manipulated treatments and observed effects is plausibly causal; that few practical methods exist for validly representing the populations and other constructs sampled in the existing literature; and that even fewer such methods exist for extrapolation. Yet, causal extrapolation is more important for the policy sciences, I argue, than is causal representation. (p. 527)

Extrapolations can be particularly useful when based on information-rich samples and designs—that is, studies that produce relevant information carefully targeted to specific concerns about both the present and the future. Users of evaluation, for example, will usually expect evaluators to thoughtfully extrapolate from their findings in the sense of pointing out lessons learned and potential applications to future efforts. Sampling strategies in qualitative evaluations can be planned with the stakeholders’ desire for extrapolation in mind.

High-Quality Lessons Learned The notion of identifying and articulating “lessons learned” has become popular as a way of extracting useful and actionable knowledge from cross-case analyses. Rather than being stated in the form of traditional scientific empirical generalizations, lessons learned take the form of principles of practice that must be adapted to particular settings in which the principle is to be applied. For example, a lesson learned from research on evaluation use is that evaluation use will likely be enhanced by designing an evaluation to answer the focused questions of specific primary intended users (Cousins & Bourgeois, 2014; Patton, 2008).

Ricardo Millett, former Director of Evaluation at the W. K. Kellogg Foundation, and I analyzed the lessons- learned sections of grantee evaluation reports. What we found was massive confusion and inconsistency. Listed under the heading “lessons” were findings, opinions, ideas, visions, and recommendations—but seldom lessons. Exhibit 9.12 provides examples of what we found.

EXHIBIT 9.12 Confusion About What Constitutes a Lesson Learned

A lesson, in the context of extracting useable knowledge from findings, takes the form of an if . . . then proposition that provides direction for future action in the real world.

Lesson about evaluation use. If you actively involve intended users in designing an evaluation to ensure its relevance, they are more likely to be interested in and actually use the findings.

This lesson meets two criteria: (1) it is based on evidence from studies of evaluation use (Cousins & Bourgeois, 2014; Patton, 2008) and (2) it provides guidance for future action (an extrapolation from past evidentiary patterns to future desired outcomes). A lesson provides guidance, but it is different from a law, a recipe, or a theoretical proposition.

A physical law. If you heat water to 100 degrees Celsius at sea level, it will boil.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 4/36

A recipe. Place a cup of oats in two cups of water, add a pinch of salt, and boil for five minutes. Remove from heat, and leave covered for two minutes. It is then ready to serve.

A theoretical proposition. It describes how the world works, as with natural selection: If a mutation provides a reproductive advantage that is heritable, over many generations that trait will become dominant in the population.

Using the definition of lesson and these distinctions, here is a sample of statements from evaluation reports illustrating confusion about what constitutes a lesson—and a lesson learned.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 5/36

High-Quality Lessons As we looked at examples of “lessons” listed in a variety of evaluation reports, it became clear that the label was being applied to any kind of insight, evidentially based or not. We began thinking about what would constitute “high-quality lessons” and decided that one’s confidence in the transferability or extrapolated relevance of a supposed lesson would increase to the extent towhich it was supported by multiple sources and types of learnings (triangulation). Exhibit 9.13 on the next page presents a list of kinds of evidence that could be accumulated to support a proposed lesson, making it more worthy of application and adaptation to new settings if it has triangulated support from a variety of perspectives and data sources. Questions for generating lessons learned are also listed. Thus, for example, the lesson that designing an evaluation to answer the focused questions of specific primary intended users enhances evaluation use is supported by research on use, theories about diffusion of innovation and change, practitioner wisdom, cross-case analyses of use, the profession’s articulation of standards, and expert testimony. High-quality lessons, then, constitute guidance extrapolated from multiple sources and independently triangulated to increase transferability as cumulative knowledge and working hypotheses that can be adapted and applied to new situations. This is a form of pragmatic utilitarian generalizability, if you will. The pragmatic bias in this approach reflects the wisdom of Samuel Johnson: “As gold which he cannot spend will make no man rich, so knowledge which he cannot apply will make no man wise.”

Principles Principles are lessons expressed more generically, taken to a higher level of generalizability, and stated in a more direct and less contingent manner.

Lesson about evaluation use: If you actively involve intended users in designing an evaluation to ensure its relevance, they are more likely to be interested in and actually use the findings.

Principle to enhance evaluation use: Form and nurture a relationship with primary intended users built around their information needs and intended uses of the evaluation.

Principles are built from lessons that are based on evidence about how to accomplish some desired result. Qualitative inquiry is an especially productive way to generate lessons and principles precisely because purposeful sampling of information-rich cases, systematically and diligently analyzed, yields rich, contextually sensitive findings. This combination of qualitative elements constitutes the intellectual farming system from which nutritious lessons and principles grow and thrive. I have discussed principles-focused qualitative inquiry throughout this book.

EXHIBIT 9.13 High-Quality Lessons Learned

High-quality lessons learned. Knowledge that can be applied to future action and derived from multiple sources of evidence (triangulation)

1. Evaluation findings—patterns across programs 2. Basic and applied research 3. Practice wisdom and experience of practitioners 4. Experiences reported by program participants/clients/intended beneficiaries 5. Expert opinion 6. Cross-disciplinary findings and patterns

The idea is that the greater the number and quality of supporting sources for a “lesson,” the more rigorous the supporting evidence, and the greater the triangulation of supporting sources, the more confidence one

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 6/36

has in the significance and meaningfulness of the lesson. Lessons promulgated with only one type of supporting evidence would be considered a “lessons” hypothesis. Nested within and cross-referenced to lessons should be the actual cases from which practice wisdom and evaluation findings have been drawn. A critical principle here is to maintain the contextual frame for lessons—that is, to keep lessons grounded in their context. For ongoing learning, the trick is to follow future-supposed applications of lessons to test their wisdom and relevance over time in action in new settings. If implemented and validated, they become high-quality lessons learned.

Questions for Generating High-Quality Lessons Learned

1. What is meant by a “lesson”? 2. What is meant by “learned”? 3. By whom was the lesson learned? 4. What’s the evidence supporting each lesson? 5. What’s the evidence the lesson was learned? 6. What are the contextual boundaries around the lesson (i.e., under what conditions does it apply)? 7. Is the lesson specific, substantive, and meaningful enough to guide practice in some concrete way? 8. Who else is likely to care about this lesson? 9. What evidence will they want to see? 10. How does this lesson connect with other “lessons”?

• Chapter 1: Examples of principles as both a focus of inquiry (Paris Declaration Principle for Development Aid, p. 10) and the result of comparative case study analysis (principles that distinguish great from good organizations, Collins, 2001a; adaptive from nonadaptive companies, Collins & Hansen, 2011)

• Chapter 2: Strategic principles for qualitative inquiry (Exhibit 2.1, pp. 46–47) • Chapter 3: Principles that undergird and guide various theoretical perspectives: constructivism,

hermeneutics, pragmatism • Chapter 4: Practical qualitative inquiry principles to get actionable answers (Exhibit 4.1, pp. 172–173);

principles of fully participatory and genuinely collaborative inquiry (p. 222); and principles-focused evaluation (p. 194)

• Chapter 5: Principles-focused purposeful sampling (p. 292) • Chapter 6: Principles for engaging in qualitative fieldwork (pp. 415–416) • Chapter 7: Ten interview principles and skills (Exhibit 7.2, p. 428) • Chapter 8: A principles-focused evaluation report (pp. 627–528) • Chapter 9: Rigor attribute analysis principles (pp. 675–676)

How to Extract Credible and Useful Principles: A Case Example

Scaling Up Excellence tackles a challenge that confronts every leader and organization—spreading constructive beliefs and behavior from the few to the many. This book shows what it takes to build and uncover pockets of exemplary performance, spread those splendid deeds, and as an organization grows bigger and older—rather than slipping toward mediocrity or worse—recharge it with better ways of doing the work at hand.

—Sutton and Rao (2014, p. 1)

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 7/36

This is how Robert Sutton and Huggy Rao (2014) open their influential book Scaling Up Excellence. Scaling is an applied version of the challenge of generalization. Scholars worry about generalizing findings. Philanthropic foundations, policymakers, and social innovators worry about spreading effective programs. Sutton and Rao identify five principles to guide scaling. How did they do it?

Sutton and Rao (2014) focused on two goals:

Uncovering the most rigorous evidence and theory we could find and generating observations and advice that were relevant to people who were determined to scale up excellence.

This meant bouncing back and forth between

the clean, careful, and orderly world of theory and research—that rigor we love so much as academics—and the messy problems, crazy constraints, and daily twists and turns that are relevant to real people as they strive and struggle to spread excellence to those who need it. (p. 298)

Seven Years of Inquiry

Sutton and Rao (2014) report that they began by gathering ideas and evidence, a process that took years.

We did case studies, reviewed theory and research, and huddled to develop insights about scaling challenges and how to overcome them. Little by little, this process changed from a private conversation between the two of us to ongoing conversations about scaling with an array of smart people. We were at the center of this process: making decisions about which leads, stories, and evidence to pursue; choosing which to keep, discard, or save for later; and weaving them together into (we hope) a coherent form. (p. 299)

Sutton and Rao (2014) then analyzed the evidence to reach preliminary conclusions. As conclusions emerged, they presented what they had found to people who had read their prior publications and/or attended their classes and speeches. They recruited knowledgeable and thoughtful people to review, question, and enhance their work.

This book is best described as the product of years of give-and-take between us and many thoughtful people, not as an integrated perspective that we constructed in private and are now unveiling for the first time. Hundreds of people played direct roles in helping us, and thousands more played indirect roles—even if they didn’t realize it. (p. 299)

To speak to issues of rigor and credibility, Sutton and Rao (2014) have distilled their inquiry process into seven core methods, each of which they elaborate in the methodological appendix of the book.

1. Combing through research from the behavioral sciences and beyond

2. Conducting and gathering detailed case studies

3. Brief examples from diverse media sources

4. Targeted interviews as unplanned conversations

5. Presenting emerging scaling ideas to diverse audiences

6. Teaching a “Scaling Up Excellence” class to Stanford graduate students

7. Participation in and observation of scaling at the Stanford school (an executive professional development program) (pp. 301–306)

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 8/36

What emerges from their description of their inquiry methods is a portrayal of an ongoing, generative, and iterative process of integrating theory, research, and practice around gathering and making sense of the evidence and, ultimately, distilling what they found into principles. The principles constitute a form of generalized guidance derived from and based on lessons. Remember, earlier I postulated that lessons lead to principles. The book opens with the four lessons they identified that became the basis for formulating their five scaling principles. Here’s how Sutton and Rao (2014) describe that connection and the first lesson, which is the basis for treating the principles as generalizations.

Our first big lesson is that, although the details and daily dramas vary wildly from place to place, the similarities among scaling challenges are more important than the differences. The key choices that leaders face and the principles that help organizations scale up without screwing up are strikingly consistent. (p. xi)

Why Principles? The seven years of inquiry described by Sutton and Rao (2014) generated five principles. Why principles? Because people engaged in a scaling initiative cannot simply look up some right answers and apply them. There is no recipe.

In the case of scaling, there are so many different aspects of the challenge, and the right answers vary so much across teams, organizations, and industries (and even across challenges faced by a single team or organization), that it is impossible to develop a useful “paint by numbers” approach. Regardless of how many cases, studies, and books (including this one) you read, success at scaling will always depend on making constantly shifting, complex, and not easily codified judgments. (p. 298)

Principles guide judgment. Context informs judgment. Qualitative inquiry generates principles, and then further qualitative inquiry, in a specific context, illuminates that context so that the principles can be interpreted and applied appropriately within that particular context. That process involves both extrapolation and assessing transferability, the qualitative approach to the challenge of generalizing.

Perspectives on Generalizability: A Review Four core epistemological issues are at the center of debates about the credibility and utility of qualitative

inquiry: (1) judging the quality of findings, (2) inferring causality (the challenge of attribution), (3) the validity of generalizations, and (4) determining what is true.

SIDEBAR

FROM LESSONS TO PRINCIPLES: SCALING THE TRANSFORMATIVE CHANGE INITIATIVE

Started in 2012, the Transformative Change Initiative (TCI) assists community colleges in scaling up innovation: “evidence-based strategies to improve student outcomes and program, organization, and system performance.” The TCI evaluation team reviewed case studies of effective innovations, extracted themes and lessons from those separate evaluations of diverse programs, and generated seven principles to guide the next stage of innovation.

The TCI Framework presents the rationale and guiding principles for scaling innovation in the community college context. It is important to link scaling to guiding principles because principles provide direction rather than prescription. They represent the intentionality of the innovation in ways that often allow for multiple actions (practices) to take place. Principles provide “guidance for action in the face of complexity” so that adaptation can occur in ways that achieve the intended outcome.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 9/36

The theory of change for TCI suggests scaling happens most successfully when practitioners apply guiding principles to their implementation and scaling efforts. In this view, scaling is not so much about replicating what others assert is good practice, which is a classic theory of scaling, but about practitioners and stakeholders becoming instrumental to the scaling process by igniting a chain of actions, reactions, and outcomes that reflect and ultimately reshape the context. To make this happen, practitioners need to

• be aware of the principles that guide the changes they are making to their practice, • reflect those principles in implementation over time, and • measure and assess whether the changes are producing the intended improved performance.

—Bragg et al. (2014, p. 6)

Transformative Change Initiative

The first part of this chapter dealt with the issue of quality by examining alternative criteria for judging quality (Modules 76 and 77). Chapter 8 included an extensive discussion of causal inference (pp. 582–595). This module has been examining perspectives on making generalizations. The next and final module will take up the issue of determining what is true. This module concludes with a summary of perspectives on and approaches to generalization in Exhibit 9.14.

EXHIBIT 9.14 Twelve Perspectives on and Approaches to Generalization of Qualitative Findings

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 10/36

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 11/36

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 12/36

SIDEBAR

“ALL GENERALIZATIONS ARE FALSE.”

“All generalizations are false, including this one”—doesn’t clarify much of anything.

“All generalizations are false, including this one” leads logically to “Some generalizations are true.” If you wish to trace this error back, consider “All Cretans lie,” uttered by a Cretan. It can’t be true, but it can be false. What’s interesting is that it not only leads to “Some Cretans tell the truth,” but it also leads to the conclusion that the Cretan speaking is not one of them.

—Errol Morris (2014) Documentary filmmaker and philosopher

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 13/36

MODULE

82 Enhancing the Credibility and Utility of Qualitative Inquiry byAddressing Philosophy of Science Issues

EXHIBIT 9.15 Criteria for Judging Quality

We come now to the fourth and final dimension of credibility. Let’s review. The first dimension is systematic, in-depth fieldwork that yields high-quality data. The second dimension that informs judgments of credibility is systematic and conscientious analysis. The third concerns judgments about the credibility of the researcher, which depends on training, experience, track record, status, and presentation of self. Now, to conclude, we take up the issue of philosophical belief in the value of qualitative inquiry, that is, a fundamental appreciation of naturalistic inquiry, qualitative methods, inductive analysis, purposeful sampling, and holistic thinking. Exhibit 9.15 graphically depicts these four dimensions of credibility. In the center of the graphic are the alternative criteria for judging quality that opened this chapter: traditional scientific research criteria,

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 14/36

constructivist and social construction criteria, artistic and evocative criteria, participatory and collaborative criteria, critical change criteria, systems and complexity criteria, and pragmatic criteria.

Philosophical belief in the value of qualitative inquiry is a prime determinant of credibility—and a matter of debate and controversy. Given the often-controversial nature of qualitative findings and the necessity, on occasion, to be able to explain and even defend the value and appropriateness of qualitative inquiry, this module will briefly discuss some of the most contentious issues. The selection of which philosophy of science issues to address in this closing section of the book is based on the workshops I regularly teach on qualitative evaluation methods. In those two- and three-day courses, which typically include participants from around the world, I reserve the final afternoon for open-ended exchanges about whatever matters of interest and concern participants want to raise. By then, we have covered types and applications of qualitative inquiry, design options, purposeful sampling approaches, fieldwork techniques, observational methods, interviewing skills, how to do systematic and rigorous analysis, and ethical standards. Inevitably, questions come pouring forth about the paradigms debate, political considerations, and fundamental doubts participants encounter about the legitimacy of qualitative inquiry. I’ll reproduce the questions that arise and offer my responses.

Paradigms question: Why are qualitative methods so controversial? I just want to interview people, see what they say, analyze the patterns, and report my findings? I don’t want to debate paradigms. Do we really have to deal with paradigms stuff?

You have to deal with what constitutes credible evidence. What constitutes credible evidence is a matter of debate among both scientists and nonscientists. While not always framed as a paradigms debate, and there are disagreements about what a paradigm is and whether it’s a useful concept, I think framing the controversy as a paradigms debate is both accurate and illuminating. In Chapter 3, I discussed the qualitative/quantitative paradigms debate at some length (see pp. 87–95), including an MQP Rumination against designating randomized controlled trials as the “gold standard.” In this module, I’m going to focus specifically on how that debate affects credibility and utility.

Paradigms are a way of distinguishing different perspectives in science about how best to study and understand the world. The debate sometimes takes the form of natural science versus social science, qualitative versus quantitative methods, behavioral psychology versus phenomenology, positivism versus constructivism, or realism versus interpretivism. How the debate is framed depends on the perspectives that people bring to it and the language available to them to talk about it. Whatever the terminology and labels for contrasting points of view, the debate is rooted in philosophical differences about the nature of reality and epistemological differences in what constitutes knowledge and how it is created. The paradigms debate, whatever form it takes, affects credibility and utility when particular worldviews are pitted against one another at the intersection of philosophy and methods to determine what kinds of evidence are acceptable, believable, and useful.

You may be able to carry out a qualitative study without ever addressing the issue of paradigms. But you ought to know enough about the debate and its implications, it seems to me, to address the issue if it comes up. I would alert those new to the debate that it has been and can be intense, divisive, emotional, and rancorous. And to those experienced in and tired of the debate, let me say that I’ve followed it, and been personally engaged in it, for more than 40 years. I’ve watched the debate ebb and flow, take on new forms, and attract new advocates and adversaries. But it doesn’t go away. The paradigms debate is an epistemological phoenix that emerges anew when fires of dissent mellow into become dying embers only to flame again on new winds of contention. I doubt that you can use qualitative methods without encountering and needing to deal with some aspects of the debate. As I have illustrated throughout this chapter, both scientists and nonscientists hold strong opinions about what constitutes credible evidence. Those opinions are paradigm derived and paradigm dependent because a paradigm constitutes a worldview built on epistemological assumptions, preferred definitions of key concepts, comfortable habits, entrenched values defended as truths, and beliefs offered up as evidence. As such, paradigms are deeply embedded in the socialization of adherents and practitioners, telling them what is important, legitimate, and reasonable.

So be prepared to address controversies and competing perspectives about what constitutes credible evidence even if it doesn’t come cloaked in the guise of a paradigms debate. Moreover, these are not simply

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 15/36

matters of academic debate. They have entered the public policy arena as matters of political debate.

Politics of evidence question: What makes research methods a matter of concern for politics and politicians?

In the public policy arena, advocates of randomized control trials are organized and funded to lobby the U.S. Congress to put their paradigm preferences into legislation (Coalition for Evidence-Based Policy, 2014). They have communications experts who supply reporters with positive news accounts (e.g., Keating, 2014; Kolata, 2013, 2014). On the other side, there are strong political advocacy statements for alternative paradigms: The Qualitative Manifesto (Denzin, 2010), Qualitative Inquiry and the Conservative Challenge (Denzin & Giardina, 2006), and Qualitative Inquiry and the Politics of Evidence (Denzin & Giardina, 2008). Ray Pawson (2013) has produced A Realist Manifesto. But there is no organized and funded lobbying effort on behalf of qualitative, mixed-methods, and/or realist approaches. So guess which group is successful in getting its paradigm legitimated and funded in legislation? Hint: It’s not the qualitative manifesto.

Objectivity question: Doesn’t the paradigms debate come down to objectivity versus subjectivity?

French philosopher Jean-Paul Sartre once observed that “words are loaded pistols.” The words “objectivity” and “subjectivity” are bullets people arguing fire at each other. It’s true that objectivity is held in high esteem. Science aspires to objectivity and a primary reason why decision makers commission an evaluation is to get objective data from an independent source external to the program being evaluated. The charge that qualitative methods are inevitably “subjective” casts an aspersion connoting the very antithesis of scientific inquiry. Objectivity is traditionally considered the sine qua non of the scientific method. To be subjective means to be biased, unreliable, and irrational. Subjective data imply opinion rather than fact, intuition rather than logic, impression rather than confirmation. Chapter 2 briefly discussed concerns about objectivity versus subjectivity, but I return to the issue here to address how these concerns affect the credibility and utility of qualitative analysis.

SIDEBAR

DIFFERENT MEANINGS AND USES OF OBJECTIVITY

1. Objective person. Unbiased, open-minded, and neutral 2. Objective process. Follow, document, and report procedures that do not predetermine results 3. Objective statement. Just the facts, unvarnished, put forward by an objective person following an

objective process 4. Objective reality. Belief that there is knowable, absolute reality 5. Objective scientific claim. Findings subjected to scientific peer review by members of a discipline

capable of judging the extent to which a claim has been produced by appropriate scientific methods and analysis

6. Objective methods. A design, data collection procedures, and analysis that follow accepted inquiry norms of a scientific discipline

7. Objective measure. The extent to which a given number can be interpreted as indicating the same amount of the thing measured, across persons or thing measured, using a validated and reliable instrument

8. Objective decisions. Fair and balanced judgment based on preponderance of evidence presented and explicit; transparent criteria for weighing the evidence

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 16/36

SOURCE: © Chris Lysy—freshspectrum.com

Let’s take a closer look at the objective/subjective distinction. The conventional means for controlling subjectivity and maintaining objectivity are the methods of quantitative social science: distance from the setting and people being studied, standardized quantitative measures, formal operational procedures, manipulation of isolated variables, and randomized controlled experimental designs. Yet the ways in which measures are constructed in psychological tests, questionnaires, cost–benefit indicators, and routine management information systems are no less open to the intrusion of biases than making observations in the field or asking questions in interviews. Numbers do not protect against bias; they merely disguise it. All statistical data are based on someone’s definition of what to measure and how to measure it. An “objective” statistic like the consumer price index is really made up of very subjective decisions about what consumer items to include in the index. Periodically, government economists change the basis and definition of such indices.

Philosopher of science Michael Scriven (1972a) has insisted that quantitative methods are no more synonymous with objectivity than qualitative methods are synonymous with subjectivity:

Errors like this are too simple to be explicit. They are inferred confusions in the ideological foundations of research, its interpretations, its application. . . . It is increasingly clear that the influence of ideology on methodology and of the latter on the training and behavior of researchers and on the identification and disbursement of support is staggeringly powerful. Ideology is to research what Marx suggested the economic factor was to politics and what Freud took sex to be for psychology. (p. 94)

Scriven’s (1972a) lengthy discussion of objectivity and subjectivity in educational research deserves careful reading by students and others concerned by this distinction. He skillfully detaches the notions of objectivity and subjectivity from their traditionally narrow associations with quantitative and qualitative methodology, respectively. He presents a clear explanation of how objectivity has been confused with consensual validation of something by multiple observers. Yet a little research will yield many instances of “scientific blunders” (Dyson, 2014; Livio, 2013; Youngson, 1998) where the majority of scientists were factually wrong while one dissenting observer described things as they really were (Kuhn, 1970).

Qualitative rigor has to do with the quality of the observations made by an inquirer. Scriven (1972a) emphasizes the importance of being factual about observations rather than being distant from the phenomenon being studied. Distance does not guarantee objectivity; it merely guarantees distance. Nevertheless, in the end, Scriven (1998) still finds the ideal of objectivity worth striving for as a counter to bias, and he continues to find the language of objectivity serviceable.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 17/36

In contrast, Lincoln and Guba (1986), as noted earlier, have suggested replacing the traditional mandate to be objective with an emphasis on trustworthiness and authenticity by being balanced, fair, and conscientious in taking account of multiple perspectives, multiple interests, multiple experiences, and diverse constructions of realities. Guba (1981) suggested that researchers and evaluators can learn something about these attributes from the stance of investigative journalists.

Journalism in general and investigative journalism in particular are moving away from the criterion of objectivity to an emergent criterion usually labeled “fairness” . . . Objectivity assumes a single reality to which the story or evaluation must be isomorphic; it is in this sense a one-perspective criterion. It assumes that an agent can deal with an objective (or another person) in a nonreactive and noninteractive way. It is an absolute criterion.

Journalists are coming to feel that objectivity in that sense is unattainable. . . .

Enter “fairness” as a substitute criterion. In contrast to objectivity, fairness has these features:

• It assumes multiple realities or truths—hence a test of fairness is whether or not “both” sides of the case are presented, and there may even be multiple sides.

• It is adversarial rather than one-perspective in nature. Rather than trying to hew the line with the truth, as the objective reporter does, the fair reporter seeks to present each side of the case in the manner of an advocate—as, for example, attorneys do in making a case in court. The presumption is that the public, like a jury, is more likely to reach an equitable decision after having heard each side presented with as much vigor and commitment as possible.

• It is assumed that the subject’s reaction to the reporter and interactions between them heavily determines what the reporter perceives. Hence one test of fairness is the length to which the reporter will go to test his own biases and rule them out.

• It is a relative criterion that is measured by balance rather than by isomorphism to enduring truth. (pp. 76–77)

But times change, and Guba would be unlikely to use the language of “fairness and balance” now that the most politically conservative and deliberately biased American television channel has adopted that phrase as its brand. Fairness and balance has become a euphemism for prejudiced and one-sided. Objectivity has also taken on unfortunate political and cultural connotations in some quarters, meaning uncaring, unfeeling, disengaged, and aloof. What about subjectivity, the constructivist badge of honor?

Subjectivity Deconstructed

In public discourse, it is not particularly helpful to know that philosophers of science now typically doubt the possibility of anyone or any method being totally “objective.” But subjectivity fares even worse. Even if acknowledged as inevitable (Peshkin, 1988), or valuable as a tool to understanding (Soldz & Andersen, 2012), subjectivity carries such negative connotations at such a deep level and for so many people that the very term can be an impediment to mutual understanding. For this and other reasons, as a way of elaborating with any insight the nature of the research process, the notion of subjectivity may have become as useless as the notion of objectivity.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 18/36

©2002 Michael Quinn Patton and Michael Cochran

The death of the notion that objective truth is attainable in projects of social inquiry has been generally recognized and widely accepted by scholars who spend time thinking about such matters. . . . I will take this recognition as a starting point in calling attention to a second corpse in our midst, an entity to which many refer as if it were still alive. Instead of exploring the meaning of subjectivity in qualitative educational research, I want to advance the notion that following the failure of the objectivists to maintain the viability of their epistemology, the concept of subjectivity has been likewise drained of its usefulness and therefore no longer has any meaning. Subjectivity, I feel obliged to report, is also dead. (Barone, 2000, p. 161)

But for other qualitative researchers, subjectivity is not so much about philosophy of science as it is about using one’s own experience to make sense of the world through reflexivity (Connolly & Reilly, 2007). That perspective, once entertained, can lead from a focus on the researcher’s subjectivity as a window into sense making to shared meaning making: intersubjectivity.

Intersubjectivity “Subjective” versus “objective” no longer makes sense, since everyone involved is a subject. . . . Human Social Research is intersubjective . . . built from encounters among subjects, including researchers who, like it or not, are also subjects. (Agar, 2013, pp. 108–109)

Eschewing both objectivity and subjectivity, intersubjectivity focuses on knowledge as socially constructed in human interactions. Human science research, what anthropologist Michael Agar (2013) calls the Lively Science, requires “human social relationships in order to happen at all. They are intersubjective sciences. They require social relationships with those who support the science, those who do it, those who serve as subjects of it, and those who consume it” (p. 215).

The difficult judgment call for the researcher is this: To some extent he or she should translate his or her own framework and jointly build a framework for communication with subjects of all those different types. . . . The bedrock of intersubjective research isn’t to preach or to lecture, but rather to learn and to communicate the results, though not at the price of abandoning the core principles of the science. The pressure always exists to achieve a balance, and a researcher always has to make the call of how much and in what way to handle it.

This fact has to be part of the science, not to mention a central part of training for human social researchers. How to navigate this ambiguous territory with professional integrity and product quality is a neglected topic, a neglect understandable in light of academic traditions where one could assume that whatever the dissertation committee or disciplinary peers would like was the right thing to do. That isolation is no longer possible. In my view, taking

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 19/36

human social research out into the world makes it more difficult, more interesting, more intellectually challenging, and of higher moral value than it has ever been. (pp. 215–216)

Empathic Neutrality

No consensus about substitute terminology has emerged. I prefer empathic neutrality, one of the 12 qualitative themes that I presented in Chapter 2.

While empathy describes a stance toward the people we encounter in fieldwork, calling on us to communicate interest, caring, and understanding, neutrality suggests a stance toward their thoughts, emotions, and behaviors, a stance of being nonjudgmental. Neutrality can actually facilitate rapport and help build a relationship that supports empathy by disciplining the researcher to be open to the other person and nonjudgmental in that openness.

(See pp. 57–62 for the full discussion of empathic neutrality.)

Open-Mindedness and Impartiality

I have evaluation colleagues who simply describe themselves as open-minded, which seems to satisfy most lay people. The political nature of evaluation means that individual evaluators must make their own peace with how they are going to describe what they do. The meaning and connotations of words like objectivity, subjectivity, neutrality, and impartiality will have to be worked out with particular stakeholders in specific evaluation settings. In her leadership role in evaluation in the U.S. federal government, former AEA president Eleanor Chelimsky emphasized her unit’s independence and impartiality. The perception of impartiality, she has explained, is at least as important as methodological rigor in highly political environments. Credibility, and therefore utility, are affected by “the steps we take to make and explain our evaluative decisions, [and] also intellectually, in the effort we put forth to look at all sides and all stakeholders of an evaluation” (Chelimsky, 1995, p. 219; see also Chelimsky, 2006).

I think it is worth noting that the official Program Evaluation Standards (Joint Committee on Standards, 2010) do not call for objectivity. The standards have been guiding evaluation practice for nearly four decades. They were originally formulated by social scientists and evaluators representing all the major disciplinary associations. They have twice gone through major review processes. The language used, therefore, has been thoroughly vetted. The standards call for evaluations to be credible, systematic, accurate, useful, accurate, and dependable, but not objective. The term objectivity has become a lightning rod attracting epistemological paradigms debate and therefore not useful as a standard for evaluation in the American context. In contrast, the international Quality Standards for Development Evaluation define evaluation as “objective assessment” (OECD-DAC, 2010, p. 5). Different context, different language.

Given the seven different sets of criteria for judging the quality of qualitative inquiry I identified at the beginning of this chapter, and the terms associated with each, it seems unlikely that a consensus about terminology is on the horizon. The methodological and scientific Tower of Babel stands tall and casts a long shadow. But the different perspectives on and uses of terms can be liberating because they opens up the possibility of getting beyond the meaningless abstractions and heavy-laden connotations of objectivity and subjectivity to move instead toward carefully selecting descriptive methodological language that best describes your own inquiry processes and procedures. That is, don’t label those processes as “objective,” “subjective,” “intersubjective,” “trustworthy,” or “authentic.” Instead, eschew overarching labels. Describe how you approach your inquiry, what you bring to your work, and how you’ve reflected on what you do, and then let the reader be persuaded, or not, by the intellectual and methodological rigor, meaningfulness, value, and utility of the result. In the meantime, be very careful how you use particular terms in specific contexts. Words are bullets. They are also landmines. I end this diatribe with a cautionary tale about being sensitive to the cultural context within which terms are used.

During a tour of America, former British prime minister Winston Churchill attended a buffet luncheon at which chicken was served. As he returned to the buffet for a second helping he asked, “May I have some more breast?”

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 20/36

His hostess, looking embarrassed, explained that “in this country we ask for white meat or dark meat.”

Churchill, taking the white meat he was offered, apologized and returned to his table.

The next morning the hostess received a beautiful orchid from Churchill with the following card: “I would be most obliged if you would wear this on your white meat.”

SIDEBAR

A REALIST PERSPECTIVE ON OBJECTIVITY

Evaluation cannot hope for perfect objectivity but neither does this mean it should slump into rampant subjectivity. We cannot hope for absolute cleanliness but this does not require us to enjoy a daily roll in the manure. The alternative to these two termini is for evaluation to embrace the goal of being “validity increasing” . . .

Skepticism . . . , in its English spelling, . . . constitutes the final desideratum of evaluation science.

Organised scepticism means that any scientific claim must be exposed to critical scrutiny before it becomes accepted. . . . What counts is the depth of critical scrutiny applied to the inferences drawn from any inquiry. And this level of attention depends, in turn, on the presence of a collegiate group of stakeholders and their willingness to put each other’s work under the microscope.

—Ray Pawson (2013, p. 107) The Science of Evaluation:

A Realist Manifesto

Truth and reality question: “I don’t understand this talk about multiple realities and different truths for different people. If research is anything, it ought to be about getting at true reality. I know you like quotes, so here’s one of my favorite quotes for you, from George Orwell: ‘In a time of universal deceit—telling the truth is a revolutionary act. ‘I think we ought to be research revolutionaries and speak the truth. In fact, the mantra of evaluation is: Speak truth to power. So, truth or not truth?”

It’s an important question. Certainly, there are a lot of quotes about truth. This is a thick book, and it could contain nothing but quotes about truth, which would serve to illustrate its evasiveness. Let me offer a quote from the great comedian Lily Tomlin, who, playing the character of a little girl accused by a scolding adult of making things up, responded thus:

Lady, I do not make up things. That is lies. Lies are not true. But the truth could be made up if you know how. And that’s the truth.

Or consider this observation by Thomas Schwandt, a philosopher of science and professional evaluator, who has spent much of a distinguished career grappling with this very issue. His conclusion:

TRUTH is one of the most difficult of all philosophical topics, and controversies surrounding the nature of truth lie at the heart of both apologies for and criticisms of varieties of qualitative work. Moreover, truth is intimately related to questions of meaning, and establishing the nature of that relationship is also complicated and contested.

There is general agreement that what is true or what carries truth are statements, propositions, beliefs, and assertions, but how the truth of same is established is widely debated. (Schwandt, 2007, p. 300)

Schwandt presents 10 different philosophical orientations to and theories about truth: (1) correspondence, (2) consensus, (3) coherence, (4) contextualist, (5) pragmatic, (6) hermeneutic, (7) critical theory (Foucault), (8) realist, (9) constructivist, and (10) objectivist theory. Pick your poison—or truth. We won’t resolve the

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 21/36

debate here. Not even close. Nor will others for, to add yet another quote to the collection, here’s cynic Ambrose Bierce’s (1999) assessment:

Discovery of truth is the sole purpose of philosophy, which is the most ancient occupation of the human mind and has a fair prospect of exiting with increasing activity to the end of time. (p. 201)

Since we can’t resolve the nature of truth, indulge me in a story that illustrates why it may be important to have figured out where you, yourself, stand on matters of truth. Following a presentation of evaluation findings at a public school board meeting, I was asked by the school district’s internal evaluator, “Do you, as a qualitative researcher, swear to tell the truth, the whole truth and nothing but the truth?” The question was meant to embarrass me. The researcher had an article I had written attacking overreliance on standardized tests for school evaluations and another advocating soliciting multiple perspectives from parents, teachers, students, and community members about their experiences with the school district to document diverse perspectives. In that article, and earlier editions of this book, I had expressed doubt about the utility of truth as a criterion of quality and I suspected that he hoped to lure me into an academic-sounding, arrogant, and philosophical discourse on the question “What is truth?” in the expectation that the public officials present would be alienated and dismiss my presentation. So when he asked, “Do you, as a qualitative researcher, swear to tell the truth, the whole truth and nothing but the truth?” I did not reply, “That depends on what truth means.” I said simply, “Certainly I promise to respond honestly.” Notice the shift from truth to honesty.

The researcher applying traditional social science criteria might respond, “I can show you truth insofar as it is revealed by the data.”

The constructivist might answer, “I can show you multiple truths.”

The artistically inclined might suggest that “beauty is truth.” And “fiction often reveals truth better than nonfiction.”

The critical theorist could explain that “truth depends on one’s consciousness.”

The participatory qualitative inquirer would say, “We create truth together.”

The critical change activist might say, “I offer you praxis. Here is where I take my stand. This is true for me.”

The pragmatic evaluator might reply, “I can show you what is useful. What is useful is true.”

Indeed, in this vein, Exhibit 9.7, in presenting the seven sets of criteria for judging quality, offers a political campaign button about TRUTH for each (pp. 680–681).

By the way, I noted earlier that the Program Evaluation Standards do not use the language of objectivity, but the “the Accuracy Standards are intended to increase the dependability and truthfulness [italics added] of evaluation representations” (Joint Committee on Standards, 2010). Note: Truthfulness is not TRUTH. You could do a little hermeneutic work on that distinction, should you be so inclined.

Ironically, it is sometimes easier to determine what is false than what is true. For insights into how the academic peer review process has been distorted and corrupted to generate invalid and untrustworthy results, see the widely cited and influential analysis by Professor of Health Research and Policy at Stanford School of Medicine, John P. A. Loannidis (2005) “Why Most Published Research Findings Are False.”

Truth Tests and Utility Tests

Previously I have cited the influential research by Weiss and Bucuvalis (1980) that decision makers apply both “truth” tests and “utility” tests to evaluation. “Truth,” in this case, however, means reasonably accurate and credible data (the focus of the program evaluation standards) rather than data that are true in some absolute sense. Savvy policymakers know better than most the context and perspective-laden nature of competing truths. Qualitative inquiry can present accurate data on various perspectives, including the evaluator’s perspective, without the burden of determining that only one perspective must be true.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 22/36

Evaluation theorist and methodologist Nick Smith (1978), pondering these questions, has noted that to act in the world we often accept either approximations to truth or even untruths.

For example, when one drives from city to city, one acts as if the earth is flat and does not try to calculate the earth’s curvature in planning the trip, even though acting as if the earth is flat means acting on an untruth. Therefore, in our study of evaluation methodology, two criteria replace exact truth as paramount: practical utility and level of certainty. The level of certainty required to make an adequate judgment under the law differs depending on whether one is considering an administrative hearing, an inquest, or a criminal case. Although it seems obvious that much greater certainty about the nature of things is required when legislators set national and educational policy than when a district superintendent decides whether to continue a local program, the rhetoric in evaluation implies that the same high level of certainty is required of both cases. If we were to first determine the level of certainty desired in a specific case, we could then more easily choose appropriate methods. Naturalistic descriptions give us greater certainty in our understanding of the nature of an educational process than randomized, controlled experiments do, but less certainty in our knowledge of the strength of a particular effect. . . . Our first concern should be the practical utility of our knowledge, not its ultimate truthfulness. (p. 17)

In studying evaluation use (Patton, 2008), I found that decision makers did not expect evaluation reports to produce “TRUTH” in any fundamental sense. Rather, they viewed evaluation findings as additional information that they could and did combine with other information (political, experiential, other research, colleague opinions, etc.), all of which fed into a slow, evolutionary process of incremental decision making. Kvale (1987) echoed this interactive and contextual approach to truth in emphasizing the “pragmatic validation” of findings in which the results of qualitative analysis are judged by their relevance to and use by those to whom findings are presented.

This criterion of utility can be applied not only to evaluation but also to qualitative analyses of all kinds, including textual analysis. Barone (2000), having rejected objectivity and subjectivity as meaningless criteria in the postmodern age, makes the case for pragmatic utility:

If all discourse is culturally contextual, how do we decide which deserves our attention and respect? The pragmatists offer the criterion of usefulness for this purpose. . . . An idea, like a tool, has no intrinsic value and is “true” only in its capacity to perform a desired service for its handler within a given situation. When the criterion of usefulness is applied to context-bound, historically situated transactions between itself and a text, it helps us to judge which textual experiences are to be valued. . . . The gates are opened for textual encounters, in any inquiry genre or tradition, that serve to fulfill an important human purpose. (pp. 169–170)

Focusing on the connection between truth tests and utility tests shifts attention back to credibility and quality, not as absolute generalizable judgments but as contextually dependent on the needs and interests of those receiving our analysis. This obliges researchers and evaluators to consider carefully how they present their work to others, with attention to the purpose to be fulfilled. That presentation should include reflections on how your perspective affected the questions you pursued in fieldwork, careful documentation of all procedures used so that others can review your methods for bias, and being open in describing the limitations of the perspective presented. Exhibit 9.16, at the end of this chapter (pp. 736–741), offers an in-depth description of how one qualitative inquirer dealt with these issues in a long-term participant–observer relationship. The exhibit, titled A Documenter’s Perspective, is based on her research journal and field notes. It moves the discussion from abstract philosophizing to day-to-day, in-the-trenches fieldwork encounters aimed at sorting out what is true (small t) and useful.

Finding TRUTH can be a heavy burden. I once had a student who was virtually paralyzed in writing an evaluation report because he wasn’t sure if the patterns he thought he had uncovered were really true. I suggested that he not try to convince himself or others that his findings were true in any absolute sense but, rather, that he had done the best job he could in describing the patterns that appeared to him to be present in the data and that he present those patterns as his perspective based on his analysis and interpretation of the data he had collected. Even if he believed that what he eventually produced was Truth, any sophisticated person reading the report would know that what he presented was no more than his perspective, and they would judge that perspective by their own commonsense understandings and use the information according to how it contributed to their own needs.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 23/36

SIDEBAR

TRUTH VERSUS RELATIVISM

Postmodern work is often accused of being relativistic (and evil) since it does not advocate a universal, independent standard of truth. In fact, relativism is only an issue for those who believe there is a foundation, a structure against which other positions can be objectively judged. In effect, this position implies that there is no alternative between objectivism and relativism. Postmodernists dispute the assumptions that produce the objectivism/relativism binary since they think of truth as multiple, historical, contextual, contingent, political, and bound up in power relations. Refusing the binary does not lead to the abandonment of truth, however, as Foucault emphasizes when he says, “I believe too much in truth not to suppose that there are different truths and different ways of speaking the truth.”

Furthermore, postmodernism does not imply that one does not discriminate among multiple truths, that “anything goes.”. . . If there is no absolute truth to which every instance can be compared for its truth-value, if truth is instead multiple and contextual, then the call for ethical practice shifts from grand, sweeping statements about truth and justice to engagements with specific, complex problems that do not have generalizable solutions. This different state of affairs is not irresponsible, irrational, or nihilistic. . . . As with truth, postmodern critiques argue for multiple and historically specific forms of reason. (St. Pierre 2000, p. 25)

As one additional source of reflection on these issues, perhaps the following Sufi story will provide some guidance about the difference between truth and perspective. Sagely, in this encounter, Nasrudin gathers data to support his proposition about the nature of truth. Here’s the story.

Mulla Nasrudin was on trial for his life. He was accused of no less a crime than treason by the king’s ministers, wise men charged with advising on matters of great import. Nasrudin was charged with going from village to village inciting the people by saying, “The king’s wise men do not speak truth. They do not even know what truth is. They are confused.” Nasrudin was brought before the king and the court. “How do you plead, guilty or not guilty?”

“I am both guilty and not guilty,” replied Nasrudin.

“What, then, is your defense?”

Nasrudin turned and pointed to the nine wise men who were assembled in the court. “Have each sage write an answer to the following question: ‘What is water?’”

The king commanded the sages to do as they were asked. The answers were handed to the king, who read to the court what each sage had written.

The first wrote, “Water is to remove thirst.”

The second, “It is the essence of life.”

The third, “Rain.”

The fourth, “A clear, liquid substance.”

The fifth, “A compound of hydrogen and oxygen.”

The sixth, “Water was given to us by God to use in cleansing and purifying ourselves before prayer.”

The seventh, “It is many different things—rivers, wells, ice, lakes, so it depends.”

The eighth, “A marvelous mystery that defies definition.”

The ninth, “The poor man’s wine.”

Nasrudin turned to the court and the king: “I am guilty of saying that the wise men are confused. I am not, however, guilty of treason because, as you see, the wise men are confused. How can they know if I have

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 24/36

committed treason if they cannot even decide what water is? If the sages cannot agree on the truth about water, something which they consume every day, how can one expect that they can know the truth about other things?”

The king ordered that Nasrudin be set free.

SIDEBAR

TRUE FACTS VERSUS TRUE THEORIES

Facts and theories are born in different ways and are judged by different standards. Facts are supposed to be true or false. They are discovered by observers or experimenters. A scientist who claims to have discovered a fact that turns out to be wrong is judged harshly. One wrong fact is enough to ruin a career.

Theories have an entirely different status. They are free creations of the human mind, intended to describe our understanding of nature. Since our understanding is incomplete, theories are provisional. Theories are tools of understanding; and a tool does not need to be precisely true in order to be useful. Theories are supposed to be more-or-less true, with plenty of room for disagreement. A scientist who invents a theory that turns out to be wrong is judged leniently. Mistakes are tolerated, so long as the culprit is willing to correct them when nature proves them wrong.

—Physicist Freeman Dyson (2014, p. 4) Institute for Advanced Studies, Princeton

Enhanced Credibility and Increased Legitimacy for Qualitative Methods: Looking Back and Looking Ahead

The distinction between the past, present, and future is only a stubbornly persistent illusion. —Physicist Albert Einstein

Theory of Relativity

Chapter Summary This chapter has reviewed ways of enhancing the quality, credibility, and utility of qualitative analysis by dealing with four distinct but related inquiry concerns:

• Rigorous methods for doing fieldwork that yield high-quality data • Systematic and conscientious analysis with attention to issues of credibility • The credibility of the researcher, which depends on training, experience, track record, status, and

presentation of self • Philosophical belief in the value of qualitative inquiry—that is, a fundamental appreciation of naturalistic

inquiry, qualitative methods, inductive analysis, purposeful sampling, and holistic thinking.

Exhibit 9.15 presented a graphic depicting these four dimensions, with criteria for judging quality in the center.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 25/36

Conclusion: Beyond the Qualitative/Quantitative Debate Question: What’s the status of the qualitative/quantitative debate today? From your perspective, what does the future look like for qualitative inquiry?

The debate between qualitative and quantitative methodologists was often strident historically, but in recent years the debate has mellowed. A consensus has gradually emerged that the important challenge is to appropriately match methods to purposes and inquiry questions, not to universally and unconditionally advocate any single methodological approach for all inquiry situations. Indeed, eminent methodologist Thomas Cook, one of evaluation’s luminaries, pronounced in his keynote address to the 1995 International Evaluation Conference in Vancouver that “qualitative researchers have won the qualitative/quantitative debate.”

Won in what sense?

Won acceptance.

The validity of experimental methods and quantitative measurement, appropriately used, was never in doubt. Now, qualitative methods have ascended to a level of parallel respectability. I have found increased interest in and acceptance of qualitative methods in particular and multiple methods in general. Especially in evaluation, a consensus has emerged that researchers and evaluators need to know and use a variety of methods in order to be responsive to the nuances of particular empirical questions and the idiosyncrasies of specific stakeholder needs. The debate has shifted from quantitative versus qualitative to strong differences of opinion about how to establish causality (the attribution and so-called gold standard debate discussed in Chapters 3 and 8). While related, that’s a narrower issue.

The credibility and respectability of qualitative methods varies across disciplines, university departments, professions, time periods, and countries. In the field I know best, program evaluation, the increased legitimacy of qualitative methods is a function of more examples of useful, high-quality evaluations employing qualitative methods and an increased commitment to providing useful and understandable information based on stakeholders’ concerns. Other factors that contribute to increased credibility include more and higher- quality training in qualitative methods and the publication of a substantial qualitative literature.

The history of the paradigms debate parallels the history of evaluation. The earliest evaluations focused largely on quantitative measurement of clear, specific goals and objectives. With the widespread social and educational experimentation of the 1960s and early 1970s, evaluation designs were aimed at comparing the effectiveness of different programs and treatments through rigorous controls and experiments. This was the period when the quantitative/experimental paradigm dominated. By the middle 1970s, the paradigms debate had become a major focus of evaluation discussions and writings. By the late 1970s, the alternative qualitative/naturalistic paradigm had been fully articulated (Guba, 1978; Patton, 1978; Stake, 1975, 1978). During this period, concern about finding ways to increase use became predominant in evaluation, and evaluators began discussing standards. A period of pragmatism and dialogue followed, during which calls for and experiences with multiple methods and a synthesis of paradigms became more common. The advice of Cronbach (1980), in his important book on reform of program evaluation, was widely taken to heart: “The evaluator will be wise not to declare allegiance to either a quantitative–scientific–summative methodology or a qualitative–naturalistic–descriptive methodology” (p. 7).

Signs of detente and pragmatism now abound. Methodological tolerance, flexibility, eclecticism, and concern for appropriateness rather than orthodoxy now characterize the practice, literature, and discussions of evaluation. Several developments seem to me to explain the withering of the methodological paradigms debate.

1. The articulation of professional standards has emphasized methodological appropriateness rather than paradigm orthodoxy (Joint Committee, 2010; OECD-DAC, 2010). Within the standards as context, the focus on conducting evaluations that are useful, practical, ethical, accurate, and accountable have reduced paradigms polarization.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 26/36

2. The strengths and weaknesses of both quantitative/experimental methods and qualitative/naturalistic methods are now better understood. In the original debate, quantitative methodologists tended to attack some of the worst examples of qualitative evaluations while the qualitative evaluators tended to hold up for critique the worst examples of quantitative/experimental approaches. With the accumulation of experience and confidence, exemplars of both qualitative and quantitative approaches have emerged with corresponding analyses of the strengths and weaknesses of each. This has permitted more balance and a better understanding of the situations for which various methods are most appropriate as well as grounded experience in how to combine methods.

3. A broader conceptualization of evaluation, and of evaluator training, has directed attention to the relation of methods to other aspects of evaluation, like use, and has therefore reduced the intensity of the methods debate as a topic unto itself.

4. Advances in methodological sophistication and diversity within both paradigms have strengthened diverse applications to evaluation problems. The proliferation of books and journals in evaluation, including but not limited to methods contributions, has converted the field into a rich mosaic that cannot be reduced to quantitative versus qualitative in primary orientation. Moreover, the upshot of all the developmental work in qualitative methods is that, as documented in Chapter 3, today there is as much variation among qualitative researchers as there is between qualitatively and quantitatively oriented scholars.

5. Support for methodological eclecticism from major figures and institutions in evaluation increased methodological tolerance. When eminent measurement and methods scholars like Donald Campbell and Lee J. Cronbach began publicly recognizing the contributions that qualitative methods could make, the acceptability of qualitative/naturalistic approaches was greatly enhanced. Another important endorsement of multiple methods came from the Program Evaluation and Methodology Division of the U.S. General Accounting Office (GAO), which arguably did the most important and influential evaluation work at the national level. Under the leadership of Assistant Comptroller General and former AEA president (1995) Eleanor Chelimsky, GAO published a series of methods manuals, including Case Study Evaluations (GAO, 1987), Prospective Evaluation Methods (GAO, 1989), and The Evaluation Synthesis (GAO, 1992). The GAO manual Designing Evaluations put the paradigms debate to rest as it described what constituted a “strong evaluation.”

Strength is not judged by adherence to a particular paradigm. It is determined by use and technical adequacy, whatever the method, within the context of purpose, time, and resources.

Strong evaluations employ methods of analysis that are appropriate to the question; support the answer with evidence; document the assumptions, procedures, and modes of analysis; and rule out competing evidence. Strong studies pose questions clearly, address them appropriately, and draw inferences commensurate with the power of the design and the availability, validity, and reliability of the data. Strength should not be equated with complexity. Nor should strength be equated with the degree of statistical manipulation of data. Neither infatuation with complexity nor statistical incantation makes an evaluation stronger.

The strength of an evaluation is not defined by a particular method. Longitudinal, experimental, quasi-experimental, before-and-after, and case study evaluations can be either strong or weak. . . . That is, the strength of an evaluation has to be judged within the context of the question, the time and cost constraints, the design, the technical adequacy of the data collection and analysis, and the presentation of the findings. A strong study is technically adequate and useful—in short, it is high in quality. (GAO, 1991, pp. 15–16)

6. Evaluation professional societies have supported exchanges of views and high-quality professional practice in an environment of tolerance and eclecticism. The evaluation professional societies and journals serve a variety of people from different disciplines who operate in different kinds of organizations at different levels, in and out of the public sector, and in and out of universities. This diversity, and opportunities to exchange views and perspectives, has contributed to the emergent pragmatism, eclecticism, and tolerance in the field. A good example was the appearance two decades ago of a volume of New Directions for Program Evaluation on The Qualitative–Quantitative Debate: New Perspectives (Reichardt & Rallis, 1994). The tone of the eight distinguished contributions in that volume is captured by

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 27/36

phrases such as “peaceful coexistence,” “each tradition can learn from the other,” “compromise solution,” “important shared characteristics,” and “a call for a new partnership.”

7. There is increased advocacy of and experience in combining qualitative and quantitative approaches. The Reichardt and Rallis (1994) volume just cited also included these themes: “blended approaches,” “integrating the qualitative and quantitative,” “possibilities for integration,” “qualitative plus quantitative,” and “working together.” Exhibit 9.2 presented 10 developments enhancing mixed-methods triangulation (p. 666).

Matching Claims and Criteria The withering of the methodological paradigms debate holds out the hope that studies of all kinds can be judged on their merits according to the claims they make and the evidence marshaled in support of those claims. The thing that distinguishes the seven sets of criteria for judging quality introduced in this chapter (Exhibit 9.7) is that they support different kinds of claims. Traditional scientific claims, constructivist claims, artistic claims, participatory inquiry claims, critical change claims, systems claims, and pragmatic claims will tend to emphasize different kinds of conclusions with varying implications. In judging claims and conclusions, the validity of the claims made is only partly related to the methods used in the process.

Validity is a property of knowledge, not methods. No matter whether knowledge comes from an ethnography or an experiment, we may still ask the same kind of questions about the ways in which that knowledge is valid. To use an overly simplistic example, if someone claims to have nailed together two boards, we do not ask if their hammer is valid, but rather whether the two boards are now nailed together, and whether the claimant was, in fact, responsible for that result. In fact, this particular claim may be valid whether the nail was set in place by a hammer, an airgun, or the butt of a screwdriver. A hammer does not guarantee successful nailing, successful nailing does not require a hammer, and the validity of the claim is in principle separate from which tool was used. The same is true of methods in the social behavioral sciences. (Shadish, 1995a, p. 421)

This brings us back to a pragmatic focus on the utility of findings as a point of entry for determining what’s at stake in the claims made in a study and therefore what criteria to use in assessing those claims. As I noted in opening this chapter, judgments about credibility and quality depend on criteria. And though this chapter has been devoted to ways of enhancing quality and credibility, all such efforts ultimately depend on the willingness of the inquirer to weigh the evidence carefully and be open to the possibility that what has been learned most from a particular inquiry is how to do it better next time.

Canadian-born bacteriologist Oswald Avery, discoverer of DNA as the basic genetic material of the cell, worked for years in a small laboratory at the hospital of the Rockefeller Institute in New York City. Many of his initial hypotheses and research conclusions turned out, on further investigation, to be wrong. His colleagues marveled that he never turned argumentative when findings countered his predictions and never became discouraged. He was committed to learning and was often heard telling his students, “Whenever you fall, pick up something.”

A final Halcolm story on the nature of journeys ends this chapter—and this book.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 28/36

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 29/36

EXHIBIT 9.16 A Documenter’s Perspective

by Beth Alberty

Introduction

This exhibit provides a reflective case study of the struggle experienced by one internal, formative program evaluator of an innovative school art program as she tried to figure out how to provide useful information to program staff from the voluminous qualitative data she collected. Beth begins by describing what she means by “documentation” and then shares her experiences as a novice in analyzing the data, a process of moving from a mass of documentary material to a unified, holistic document.

Documentation

Documentation, as the word is commonly used, may refer to “slice of life” recordings in various media or to the marshalling of evidence in support of a position or point of view. We are familiar with “documentary” films; we require lawyers or journalists to “document” their cases. Both meanings contribute to my view of what documentation is, but they are far from describing it fully. Documentation,

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 30/36

to my mind, is the interpretive reconstitution of a focal event, setting, project, or other phenomenon, based on observation and on descriptive records set in the context of guiding purposes and commitments.

I have always been a staff member of the situations I have documented, rather than a consultant or an employee of an evaluation organization. At first this was by accident, but now it is by conviction: My experience urges that the most meaningful evaluation of a program’s goals and commitments is one that is planned and carried out by the staff and that such an evaluation contributes to the program as well as to external needs for information. As a staff member, I participate in staff meetings and contribute to decisions. My relationships with other staff members are close and reciprocal. Sometimes I provide services or perform functions that directly fulfill the purposes of the program—for example, working with children or adults, answering visitor’s questions, and writing proposals and reports. Most of my time, however, is spent planning, collecting, reporting, and analyzing documentation.

First Perceptions

With this context in mind, let me turn to the beginning plunge. Observing is the heart of documenting, and it was into observing that I plunged, coming up delighted at the apparent ease and swiftness with which I could fish insight and ideas from the ceaseless ocean of activity around me. Indeed, the fact that observing (and record keeping) does generate questions, insight, and matters for discussion is one of many reasons why records for any documentation should be gathered by those who actually work in the setting.

My observing took many forms, each offering a different way of releasing questions and ideas— interactive and noninteractive observations were transcribed or discussed with other staff members and thereby rethought; children’s writing was typed out, the attention to every detail involving me in what the child was saying; notes of meetings and other events were rewritten for the record; and so on. Handling such detail with attention, I found, enabled me to see into the incident or piece of work in a way I hadn’t on first look. Connections with other things I knew, with other observations I made, or questions I was puzzling over seemed to proliferate during these processes; new perceptions and new questions began to form.

I have heard others describe similarly their delighted discovery of the provocativeness of record-keeping processes. The teacher who begins to collect children’s art, without perhaps even having a particular reason for the collecting, will, just by gathering the work together, begin to notice things about them that he or she had not seen before—how one child’s work influences another’s, how really different (or similar) are the trees they make, and so on. The in-school advisor or resource teacher who reviews all his or her contacts with teachers—as they are recorded or in a special meeting with his or her colleagues— may begin, for example, to see patterns of similar interest in the requests he or she is getting and thus become aware of new possibilities for relationships within the school.

My own delight in this apparently easy access to a first level of insight made me eager to collect more and more, and I also found the sheer bulk of what I could collect satisfying. As I collected more records, however, my enthusiasm gradually changed to alarm and frustration. There were so many things that could be observed and recorded, so many perspectives, such a complicated history! My feelings of wanting more changed to a feeling of needing to get everything. It wasn’t enough for me to know how the program worked now—I felt I needed to know how it got started and how the present workings had evolved. It wasn’t enough to know how the central part of the program worked—I felt I had to know about all its spinoff activities and from all points of view. I was quickly drawn into a fear of losing something significant, something I might need later on. Likewise, in my early observations of class sessions, I sought to write down everything I saw. I have had this experience of wanting to get everything in every setting in which I have documented, and I think it is not unique.

I was fortunate enough to be able to indulge these feelings and to learn from where they led me. It did become clear to me after a while that my early ambitions for documenting everything far exceeded my time and, indeed, the needs of the program. Nevertheless, there was a sense to them. Collecting so much was a way of getting to know a new setting, of orienting myself. And, not knowing the setting, I couldn’t

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 31/36

know what would turn out to be important in “reconstituting” it; also, the purpose of “reconstituting” it was sufficiently broad to include any number of possibilities from which I had not yet selected. In fact, I found that the first insights, the first connections that came from gathering the records were a significant part of the process of determining what would be important and what were the possibilities most suited to the purposes of the documentation. The process of gathering everything at first turned out to be important and, I think, needs to be allowed for at the beginning of any documenting effort. Even though much of the material so gathered may remain apparently unused, as it was in my documenting, in fact it has served its purpose just in being collected. A similar process may be required even when the documenter is already familiar with the setting, since the new role entails a new perspective.

The first connections, the first patterns emerging from the accumulating records were thus a valuable aspect of the documenting process. There came a moment, however, when the data I had collected seemed more massive than was justified by any thought I’d had as a result of the collecting. I was ill at ease because the first patterns were still fairly unformed and were not automatically turning into a documentation in the full sense I gave earlier, even though I recognized them as part of the documentary data. Particularly, they did not function as “evaluation.” Some further development was needed, but what? “What do I do with them now?” is a cry I have heard regularly since then from teachers and others who have been collecting records for a while.

I began with the relatively simple procedure of rereading everything I had gathered. Then, I returned to rethink what my purposes were and sought out my original resources on documentation. Rereading qualitative references, talking with the staff of the school and with my staff colleagues, I began to imagine a shape I could give to my records that would make a coherent representation of the program to an outside audience.

At the same time, I began to rethink how I could make what I had collected more useful to the staff. Conceiving an audience was very important at this stage. I will be returning to this moment of transition from initial collecting to rethinking later, to analyze the entry into interpretation that it entails. Descriptively, however, what occurred was that I began to see my observations and records as a body with its own configurations, interrelationships, and possibilities, rather than simply as excerpts of the larger program that related only to the program. Obviously, the observations and records continued to have meaning through their primary relationship to the setting in which they were made; but they also began to have meaning through their secondary relationships to each other.

These secondary relationships also emerge from observation as a process of reflecting. Here, however, the focus of observation is the setting as it appears in and through the observations and records that have accumulated, with all their representation of multiple perspectives and longitudinal dimensions. These observations in and through records —“thickened observations”—are of course confirmed and added to by continuing direct observation of the setting.

Beginning to see the records as a body and the setting through thickened observation is a process of integrating data. The process occurs gradually and requires a broad base of observation about many aspects of the program over some period of time. It then requires concentrated and systematic efforts to find connections within the data and weave them into patterns, to notice changes in what is reported, and find the relationship of changes to what remains constant. This process is supported by juxtaposing the observations and records in various ways as well as by continual return to reobserve the original phenomenon. There is, in my opinion, no way to speed up the process of documenting. Reflectiveness takes time.

In retrospect, I can identify my own approach to an integration of the data as the time when I began to give my opinions on long-range decisions and interpretations of daily events with the ease of any other staff member. Up to the moment of transition, I shared specific observations from the records and talked them over as a way of gathering yet more perspectives on what was happening. I was aware, however, that my opinions or interpretations were still personal. They did not yet represent the material I was collecting.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 32/36

Thus, it may be that integration of the documentary material becomes apparent when the documenter begins to evince a broad perspective about what is being documented, a perspective that makes what has been gathered available to others without precluding their own perceptions. This perspective is not a fixed-point view of a finished picture, both the view and the picture constructed somehow by the documenter in private and then unveiled with a flourish. It is also not a personal opinion; nor does it arise from placing a predetermined interpretive structure or standard on the observations. The perspective results from the documenter’s own current best integration of the many aspects of the phenomenon, of the teachers’ or staff’s aims, ideas, and current struggles, and of their historical development as these have been conveyed in the actions that have been observed and the records that have been collected.

As documenter, my perspective of a program or a classroom is like my perspective of a landscape. The longer I am in it, the sharper defined become its features, its hills and valleys, forests and fields, and the folds of distance; the more colorful and yet deeply shaded and nuanced in tone it appears; the more my memory of how it looks in other weather, under other skies, and in other seasons, and my knowledge of its living parts, its minute detail, and its history deepen my viewing and valuing of it at any moment. This landscape has constancy in its basic configurations, but is also always changing as circumstances move it and as my perceptions gather. The perspective the documenter offers to others must evoke the constancy, coherence, and integrity of the landscape, and its possibilities for changing its appearance. Without such a perspective, an organization or integration that is both personal and informed by all that has been gathered by myself and by others in the setting—others could not share what I have seen—could not locate familiar landmarks and reflect on them as they exhibit new relationships to one another and to less familiar aspects. All that material, all those observations and records, would be a lifeless and undoubtedly dusty pile.

The process of forming a perspective in which the data gathered are integrated into an organic configuration is obviously a process of interpretation. I had begun documenting, however, without an articulated framework for interpretation or a format for representation of the body of records, like the theoretical framework researchers bring to their data. Of course, there was a framework. Conceptions of artistic process, of learning and development, were inherent in the program; but these were not explicit in its goals as a program to provide certain kinds of service. The plan of the documentation had called for certain results, but there was no specified format for presentation of results. Therefore, my entry into interpretation became a struggle with myself over what I was supposed to be doing. It was a long internal debate about my responsibilities and commitments.

When I began documenting this particular school’s art program, for example, I had priorities based on my experience and personal commitments. It seemed to me self-evidently important to provide art activities for children and to try and connect these to other areas of their learning. I knew that art was not something that could be “learned” or even experienced on a once-a-week basis, so I thought it was important to help teachers find various ways of integrating art and other activities into their classrooms. I had already made a personal estimate that what I was documenting was worthwhile and honest. I had found points of congruence between my priorities and the program. I could see how the various structures of the program specified ways of approaching the goals that seemed possible and that also enabled the elaboration of the goals.

This initial commitment was diffuse; I felt a kind of general enthusiasm and interest for the efforts I observed and a desire to explore and be helpful to the teachers. In retrospect, however, the commitment was sufficiently energizing to sustain me through the early phases of collecting observations and records, when I was not sure what these would lead to. Rather than restricting me, the commitment freed me to look openly at everything (as reflected in the early enthusiasm for collecting everything). Obviously, it is possible to begin documenting from many other positions of relative interest and investment, but I suspect that even if there is no particular involvement in program content on the part of the documenter, there must be at least some idea of being helpful to its staff. (Remember, this was a formative evaluation.) Otherwise, for example, the process of gathering data may be circumscribed.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 33/36

At the point of beginning to “do something” with the observations and records, I was forced to specify the original commitment, to rethink my purposes and goals. Rereading the observations and records as a preliminary step in reworking to address different audiences, I found myself at first reading with an idea of “balancing” success and failure, an idea that constricted and trivialized the work I had observed and recorded. Thankfully, it was immediately evident from the data itself that such balance was not possible. If, during 10 days of observation, a child’s experience was intense 1 day and characterized by rowdy socializing the other 9, a simple weigh-off would not establish the success or failure of the child’s experience. The idea was ludicrous. Similarly, the staff might be thorough in its planning and follow- through on one day and disorganized on another day, but organization and planning were clearly not the totality of the experience for children.

Such trade-offs implied an external, stereotyped audience awaiting some kind of quantitative proof, which I was supposed to provide in a disinterested way, like an external, summative evaluator. The “balanced view” phase was also like my early record gathering of everything. What I was documenting was still in fragments for me, and my approach was to the particulars, to every detail.

A second approach to interpreting, also brief, took a slightly broader view of the data, a view that acknowledged my original estimate of program value and attempted to specify it. Perceiving through the data the landscape-like configurations of program strengths, I made assessments that included statements of past mistakes or inadequacies like minor “flaws” in the landscape (e.g., a few odd billboards and a garbage dump in one of Poussin’s dreams of classical Italy) rather than debits on a balance sheet. Here again, the implication was of an external audience, expecting some absolute of accomplishment. The “flaws” could be “minor” only by reference to an implied major flaw—that of failing to carry out the program goals altogether.

The formulation of strength subsuming weakness could not withstand the vitality of the records I was reading. The reality the data portrayed became clearer as the inadequacy of my first formulations of how to interpret the documentary material was revealed. Similarly, the implications of external audience expectations were not justified by the actuality of my relationship to the program and staff. My stated goal as documenter had been originally to set up record-keeping procedures that would preserve and make available to staff and to other interested persons aspects of the beginnings and workings of the program, and to collect and analyze some of the material as an assessment of what further possibilities for development actually existed. My goals had not been to evaluate in the sense of an external judgment of success or failure.

Thinking over what other approaches to interpretation were possible, I recalled that I had gathered documentary materials quite straightforwardly as a participant, whose engagement was initially through recognition of shared convictions and points of congruence with the program. Perhaps, I decided, I could share my viewpoint of the observations just as straightforwardly, as a participant with a particular point of view. In examining this possibility, I came to a view of interpreting observational data as a process of “rendering,” much as a performer renders a piece of classical music. The interpretation follows a text closely—as a scientist might say, it sticks closely to the facts. But it also reflects the performer, specifically the performer’s particular manner of engagement in the enterprise shared by text and performer, the enterprise of music. The same relationship could exist, it seemed to me, between a body of observations and records gathered participatively and as documenter. The relationship would allow my personal experience and viewpoint to enhance rather than distort the data. Indeed, I would become their voice.

Through this relationship I could make the observations available to staff and to other audiences in a way that was flexible and responsive to their needs, purposes, and standards. In so doing, of course, the framework of inherent conceptions underlying the work of the program would be incorporated. Thus, to interpret the observational data I had gathered, I had to reaffirm and clarify my relationship, my attachment to and participation in the program.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 34/36

My initial engagement, with its strong coloring of prior interests and ideas, had never meant that I understood or was sympathetic with every goal or practice of every participant of the program all the time. In any joint enterprise, such as a school or program, there are diverse and multiple goals and practices. Part of the task of documenting is to describe and make these various understandings, points of view, and practices visible so that participants can reflectively consider them as the basis for planning. No participant agrees on all issues and points of practice. Part of being a participant is exploring differences and how these illuminate issues or contribute to practice. My participation allowed me to examine and extend the interests and ideas I came with as well as observing and recording those other people brought. In this process, my engagement was deepened, enabling me to make assessments closer to the data than my first readings brought. These assessments are evaluation in its original sense of “drawing-value-from,” an interactive process of valuing, of giving weight and meaning.

In the context of renewed engagement and deepened participation, assessments of mistakes or inadequacies are construed as discrepancies between a particular practice and the intent behind it, between immediate and long-range purposes. The discrepancy is not a flaw in an otherwise perfect surface, but— like the discrepancy in a child’s understanding that stimulates new learning—is the occasion for growth. It is a sign of life and possibility. The burden of the discrepancy can lie either with the practice or with the intent, and that is the point for further examination. Assessment can also occur through the observation of and search for underlying themes of continuity between present and past intent and practice, and the point of change or transformation in continuity. Whereas discrepancy will usually be a more immediate trigger to evaluation, occasions for the consideration of continuity may tend to be longer-range planning for the coming year, contemplating changes in staff and function, or commemorating an anniversary.

I have located the documenter as participant, internal to the program or setting, gathering and shaping data in ways that make them available to participants and potentially to an external audience. Returning to the image of a landscape, let me comment on the different forms availability assumes for these different audiences.

Participant access to the landscape through the documenter’s perspective cannot be achieved through ponderous written descriptions and reports on what has been observed but must be concentrated in interaction. Sometimes this may require the development of special or regular structures—a series of short-term meetings on a particular issue or problem; an occasional event that sums up and looks ahead; a regular meeting for another kind of planning. But many times the need is addressed in very slight forms, such as a comment in passing about something a child or adult user is doing, about the appearance of a display, or the recounting of another staff member’s observation. I do not mean that injecting documentation into the self-assessment process is a juggling act or some feat of manipulation; merely that the documenter must be aware that his or her role is to keep things open and that, while the observations and records are a resource for doing this, a sense of the whole they create is also essential. The landscape is, of course, changed by the new observations offered by fellow viewers.

The external audience places different requirements on the documenter who seeks to represent to it the documentary perspective. By external audience I refer to funding agencies, supervisors, school boards, institutional hierarchies, and researchers. Proposals, accounts, and reports to these audiences are generally required. They can be burdensome because they may not be organically related to the process of internal self-reflection and because the external audience has its own standards, purposes, and questions; it is unfamiliar with the setting and with the documenter, and it needs the time offered by written accounts to return and review the material. The external audience will need more history and formal description of the broad aspects than the internal audience, with commentary that indicates the significance of recent developments. This need can be met in the overall organization, arrangement, and introduction of documents, which also convey the detail and vividness of daily activity.

To limit the report to conventional format and expectations would probably misrepresent the quality of thought, of relating, of self-assessment that goes into developing the work. If there is intent to use the occasion of a report for reflection—for example, by including staff in the development of the report—the reporting process can become meaningful internally while fulfilling the legitimate external demands for

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 35/36

accounting. Naturally, such a comment engages the external audience in its own evaluative reflections by evoking the phenomenon rather than reducing it.

In closing, I return to what I see as the necessary engaged participation of the documenter in the setting being documented, not only for data gathering but for interpretation. Whatever authenticity and power my perspective as documenter has had has come, I believe, from my commitment to the development of the setting I was documenting and from the opportunities in it for me to pursue my own understanding, to assess and reassess my role, and to come to terms with issues as they arose.

We come to new settings with prior knowledge, experience, and ways of understanding, and our new perceptions and understandings build on these. We do not simply look at things as if we had never seen anything like them before. When we look at a cluster of light and dark greens with interstices of blue and some of deeper browns and purples, what we identify is a tree against the sky. Similarly, in a classroom we do not think twice when we see, for example, a child scratching his head, yet the same phenomenon might be more strictly described as a particular combination of forms and movements. Our daily functioning depends on this kind of apparently obvious and mundane interpretation of the world. These interpretations are not simply personal opinions—though they certainly may be unique—nor are they made up. They are instead organizations of our perceptions as “tree” or “child scratching” and they correspond at many points with the phenomena so described.

It is these organizations of perception that convey to someone else what we have seen and that make objects available for discussion and reflection. Such organizations need not exclude our awareness that the tree is also a cluster of colors or that the child scratching his head is also a small human form raising its hand in a particular way. Indeed, we know that there could be many other ways to describe the same phenomena, including some that would be completely numerical—but not necessarily more accurate, more truthful, or more useful! After all, we organize our perceptions in the context of immediate purposes and relationships. The organizations must correspond to the context as well as to the phenomenon.

Facts do not organize themselves into concepts and theories just by being looked at; indeed, except within the framework of concepts and theories, there are no scientific facts but only chaos. There is an inescapable a priori element in all scientific work. Questions must be asked before answers can be given. The questions are all expressions of our interest in the world; they are at bottom valuations. Valuations are thus necesssarily involved already at the stage when we observe facts and carry on theoretical analysis and not only at the stage when we draw political inferences from facts and valuations (Myrdal, 1969, p. 9).

My experience suggests that the situation in documenting is essentially the same as what I have been describing with the tree and the child scratching and what Myrdal describes as the process of scientific research. Documentation is based on observation, which is always an individual response both to the phenomena observed and to the broad purposes of observation. In documentation, observation occurs both at the primary level of seeing and recording phenomena and at secondary levels of re-observing the phenomena through a volume of records and directly, at later moments. Since documentation has as its purpose to offer these observations for reflections and evaluation in such a way as to keep alive and open the potential of the setting, it is essential that observations at both primary and secondary levels be interpreted by those who have made them. The usefulness of the observations to others depends on the documenter’s rendering them as finely as he or she is able, with as many points of correspondence to both the phenomena and the context of interpretation as possible. Such a rendering will be an interpretation that preserves the phenomena and so does not exclude but rather invites other perspective.

Of course, there is a role for the experienced observer from outside who can see phenomenon freshly; who can suggest ways of obtaining new kinds of information about it, or, perhaps more important, point to the significance of already existing procedures or data; who can advise on technical problems that have arisen within a documentation; and who can even guide efforts to interpret and integrate documentary

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 36/36

information. I am stressing, however, that the outside observer in these instances provides support, not judgment or the criteria for judgment.

The documenter’s obligation to interpret his or her observations and those reflected in the records being collected becomes increasingly urgent, and the interpretations become increasingly significant, as all the observers in the setting become more knowledgeable about it and thus more capable of bringing range and depth to the interpretation. Speaking of the weight of her observations of the Manus over a period of some 40 years to great change, Margaret Mead clarifies the responsibility of the participant–observer to contribute to both people studied and to a wider audience the rich individual interpretation of his or her own observations:

Uniqueness, now, in a study like this (of people who have come under the continuing influence of contemporary world culture), lies in the relationships between the fieldworker and the material. I still have the responsibility and incentives that come from the fact that because of my long acquaintance with this village I can perceive and record aspects of this people’s life that no one else can. But even so, this knowledge has a new edge. This material will be valuable only if I myself can organize it. In traditional fieldwork, another anthropologist familiar with the area can take over one’s notes and make them meaningful. But here it is my individual consciousness that provides the ground on which the lives of these people are figures. (Mead, 1977, pp. 282–283)

In documenting, it seems to me the contribution is all the greater, and all the more demanded, because what is studied is one’s own setting and commitment.