U10A1-68 - Qualitative Research Plan - ***TUTOR FOLLOW ALL INSTRUCTIONS AS OUTLINED TO COMPLETE THIS WORK. READ ATTACHMENTS.

profiledrcdopen82
Chaper9-EnhancingtheQualityandCredibilityofQualitativeStudiespt2.pdf

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 1/37

From the General to the Particular: Seven Sets of Criteria for Judging the Quality of Different Approaches to Qualitative Inquiry

Once we move beyond general criteria for scientific inquiry (Exhibit 9.6) to address specific quality criteria for qualitative inquiry, we must move from the general to the particular and contextual. Exhibit 9.7 lists criteria that are embedded in and flow from distinct qualitative inquiry frameworks. The traditional scientific research criteria are embedded in and derived from what I discussed in Chapter 3 as reality-testing inquiry frameworks that include positivist, postpositivist, empiricist, and foundationalist epistemologies pp. 105–108. The social construction criteria are derived from the discussion of “constructivism” in Chapter 3 pp. 121– 126. The artistic and evocative criteria are derived from the discussion of autoethnography and evocative forms of inquiry in Chapter 3, especially the criteria suggested by Richardson (2000b) for “creative analytic practice of ethnography.” The fourth set of criteria, participatory and collaborative approaches, are based on traditions and approaches reviewed in Chapter 4 pp. 213–222. The fifth set of criteria, critical change criteria, flow from critical theory, feminist inquiry, activist research, and participatory research processes aimed at empowerment. The sixth set of criteria, systems and complexity criteria, are derived from the discussion in Chapter 3 pp. 139–151.The seventh and final set of criteria, pragmatic and utilization-focused criteria, are based on discussions in Chapters 3 and 4 pp. 152–157 as well as program evaluation standards and principles (Joint Committee and Standards, 2010) and “Guiding Principles for Evaluators” (AEA Task Force on Guiding Principles for Evaluators, 1995).

To some extent, all of the theoretical, philosophical, and applied orientations reviewed in Chapters 3 and 4 provide somewhat distinct criteria, or at least priorities and emphases, for what constitutes a quality contribution within those particular perspectives and concerns. I’ve chosen these seven broader sets of criteria to capture the primary debates that differentiate qualitative approaches and, more specifically, to highlight what seem to me to differentiate reactions to qualitative inquiry. In this chapter, we are primarily concerned with how others respond to our work. With what perspectives and by what criteria will our work be judged by those who encounter and engage it?

SIDEBAR

DIFFERENT AUDIENCES INTERESTED IN AND INVOLVED IN ASSESSING THE QUALITY OF QUALITATIVE RESEARCH AND EVALUATION

Criteria of quality can and often do vary by audience (Flick, 2007b, pp. 3–8). Here are some questions to consider in thinking about the intersection of quality criteria and audience.

1. Your criteria. You, the inquirer, presumably have an interest in doing quality work. How do you decide what standards and criteria of quality you will adhere to?

2. Primary users of your findings. Others will read and potentially use your findings. Who are the intended users of what you generate, and what criteria will they apply in judging the quality and credibility of your work?

3. Funders of your inquiry. If your inquiry has been funded by a grant, an agency, an evaluation contract, or some other funding mechanism, funders will be judging whether what you produced was worth what it cost. How will they make that judgment?

4. Publication reviewers. You may want to publish your findings. How will journals, book editors, and peer reviewers judge your work?

You need not be passive about others’ criteria and judgments. Indeed, you ought not to be passive. You should make explicit the quality criteria you have applied in designing and implementing your inquiry and invite readers, funders, and peer reviewers to join you in using your criteria. You may also add the caveat

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 2/37

that if they apply different criteria, their judgments of quality may well differ from yours and from those who follow the criteria you’re operating under. In all of this keep in mind that

the question of how to ascertain the quality of qualitative research has been asked since the beginning of qualitative research and attracts continuous and repeated attention. However, answers to this question have not been found—at least not in a way that is generally agreed upon. (Flick, 2007b, p. 11)

Some of the confusion that people have in assessing qualitative research stems from thinking it represents a uniform perspective, especially in contrast to quantitative research. This makes it hard for them to make sense of the competing approaches within qualitative inquiry. By understanding the criteria that others bring to bear on our work, we can anticipate their reactions and help them position our intentions and criteria in relation to their own expectations and criteria. In terms of the Reflexive Triangulated Inquiry model presented in Chapter 2 as Exhibit 2.5 (see p. 72), we’re dealing here with the intersection between the inquirer’s perspective and the perspective of those receiving the study (the audiences).

Criteria Determine What We See: The Umpires’ Perspectives

Different perspectives about things such as truth and the nature of reality constitute paradigms or worldviews based on alternative epistemologies and ontologies. People viewing qualitative findings through different paradigmatic lenses will react differently just as we, as researchers and evaluators, vary in how we think about what we do when we study the world. These differences are nicely illustrated by the classic story of three baseball umpires who, having retired after a game to a local establishment for the dispensing of reality- distorting but truth-enhancing libations, are discussing how they call balls and strikes.

“I call them as I see them,” says the first.

“I call them as they are,” says the second.

“They ain’t nothing until I call them,” says the third.

That’s the classic version of the story. Now, thanks to high-speed camera technology, we can update the story.

As chance would have it, two management researchers, Brayden King of Northwestern University and Jerry Kim of Columbia Business School, happened to be in the same bar going over their research on the accuracy of umpires’ calls. Overhearing the three umpires, they went up to them and said, “Fourteen percent of the time you call them wrong.” Before the umpires could argue, they explained,

We analyzed more than 700,000 pitches thrown during the 2008 and 2009 seasons. In addition to an average error rate of 14%, we found that umpires tended to favor the home team and that umpires were more likely to make mistakes when the game was on the line. (Based on King & Kim, 2014, p. SR12)

The two researchers went on like this for some time, breaking down the error rates by innings, situation, pitcher and batter ethnicity and race, pitcher reputation, and so forth and so on, until finally the umpires together put up their hands and told them to stop.

The first umpire said, “Your criteria are based on a high-speed camera. We get that. You love your numbers. We get that. And your analysis is interesting, even fascinating. We get that. But during a game we don’t use a camera. So I still call them as I see them,” he reiterated.

“And I call them as they are,” repeated the second.

“And they ain’t nothing until I call them,” concluded the third.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 3/37

We turn now to discussion and elaboration of the seven alternative sets of criteria for judging the quality of qualitative work summarized in Exhibit 9.7.

1. Traditional Scientific Research Criteria

The saddest aspect of life right now is that science gathers knowledge faster than society gathers wisdom.

—Isaac Asimov (1920–1992) Science author and science fiction writer

One way to increase the credibility and legitimacy of qualitative inquiry among those who place priority on traditional scientific research criteria is to emphasize those criteria that have priority within that tradition. Science has traditionally emphasized objectivity, so qualitative inquiry within this tradition emphasizes procedures for minimizing investigator bias. Those working within this tradition will emphasize rigorous and systematic data collection procedures, for example, cross-checking and cross-validating sources during fieldwork. In analysis it means, whenever possible, using multiple coders and calculating intercoder consistency to establish the validity and reliability of pattern and theme analysis. Qualitative researchers working in this tradition are comfortable using the language of “variables” and “hypothesis testing” and striving for causal explanations and generalizability, especially in combination with quantitative data (e.g., Hammersley, 2008b). Qualitative approaches that manifest some or all of these characteristics include grounded theory (Glaser, 2000), qualitative comparative analysis (Ragin, 1987, 2000), and realism (Miles et al., 2014). Their common aim is to use qualitative methods to describe and explain phenomena as accurately and completely as possible so that their descriptions and explanations correspond as closely as possible to the way the world is and actually operates (Reynolds et al., 2011). Government agencies supporting qualitative research (e.g., the U.S. Government Accounting Office, the National Science Foundation, or the National Institutes of Health) usually operate within this traditional scientific framework.

SIDEBAR

THE ROOTS OF TRADITIONAL SOCIAL SCIENCE CRITERIA APPLIED TO QUALITATIVE INQUIRY

An emphasis on valid and reliable knowledge, as generated by neutral researchers utilizing the scientific method to discover universal Truth, reflects an epistemology commonly referred to as positivism. Historically, social scientists understood positivism as reflected in a “realist ontology, objective epistemology, and value-free axiology.” Few, if any, qualitative researchers currently subscribe to an absolute faith in positivism, however. Many postpositivists, or researchers who believe that achievement of objectivity and value-free inquiry are not possible, nonetheless embrace the goal of production of generalizable knowledge through realist methods and minimization of researcher bias, with objectivity as a “regulatory ideal” rather than an attainable goal. In short, postpositivism does not embrace naive belief in pure scientific truth; rather, qualitative research conducted in a strict postpositivist tradition utilizes precise, prescribed processes and produces social scientific reports that enable researchers to make generalizable claims about the social phenomenon within particular populations under examination.

Postpositivists commonly utilize qualitative methods that bridge quantitative methods, in which researchers conduct an inductive analysis of textual data, form a typology grounded in the data (as contrasted with a preexisting, validated typology applied to new data), use the derived typology to sort data into categories, and then count the frequencies of each theme or category across data. Such research typically emphasizes validity of the coding schema, inter-coder reliability, and careful delineation of procedures, including random or

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 4/37

otherwise systematic sampling of texts. Content analyses of media typify this approach. (Ellingson, 2011, pp. 596, 598; within-quote references omitted)

2. Social Construction and Constructivist Criteria

What is perceived as real is real in its consequences. —The Thomas theorem

Social construction, constructivist, and “interpretivist” perspectives have generated new language and concepts to distinguish quality in qualitative research (e.g., Glesne, 1999, pp. 5–6). Lincoln and Guba (1986) proposed that constructivist inquiry demanded different criteria from those inherited from traditional social science. They suggested “credibility as an analog to internal validity, transferability as an analog to external validity, dependability as an analog to reliability, and confirmability as an analog to objectivity.” In combination, they viewed these criteria as addressing “trustworthiness (itself a parallel to the term rigor)” (pp. 76–77). They went on to emphasize that naturalistic inquiry should be judged by dependability (a systematic process systematically followed) and authenticity (reflexive consciousness about one’s own perspective, appreciation for the perspectives of others, and fairness in depicting constructions in the values that undergird them). They viewed the social world (as opposed to the physical world) as socially, politically, and psychologically constructed, as are human understandings and explanations of the physical world. They advocated triangulation to capture and report multiple perspectives rather than seek a singular truth. The team of researchers who reviewed approaches to assessing quality in qualitative research found that “the post- positivist criteria developed by Lincoln and Guba, based around the construct of ‘trustworthiness,’ were referenced frequently and appeared to be the basis upon which a number of authors made their recommendations for improving quality of qualitative research” (Reynolds et al., 2011).

Constructivists embrace subjectivity as a pathway deeper into understanding the human dimensions of the world in general as well as whatever specific phenomena they are examining (Peshkin, 1985, 1988, 2000a,b). They’re more interested in deeply understanding specific cases within a particular context than in hypothesizing about generalizations and causes across time and space. Indeed, they are suspicious of causal explanations and empirical generalizations applied to complex human interactions and cultural systems. They offer perspective and encourage dialogue among perspectives rather than aiming at singular truths and linear predictions. Social constructivists’ case studies, findings, and reports are explicitly informed by attention to praxis and reflexivity—that is, understanding how one’s own experiences and background affect what one understands and how one acts in the world, including acts of inquiry. For an in-depth discussion of this perspective and its implications, see the Handbook of Constructionist Research (Holstein & Gubrium, 2008). Also see Chapter 3 pp. 121–126 for a much lengthier discussion of constructionism and constructivism.

Here are three examples of social construction as a framework for program evaluation.

1. The evaluation of a community development project in an ethnically and racially diverse neighborhood collected and reported stories from residents purposefully sampled to present a range of experiences and perspectives. The evaluation did not render judgments but was called a “multivocal evaluation” in which the diverse stories were used for dialogue and to enhance mutual understanding.

2. The evaluation of the international Paris Declaration on Development Aid included case studies that revealed the different perspectives and contexts within which aid is given and received. Donors (wealthier countries) and beneficiaries (poorer countries) experience different “realities.” One purpose of the evaluation was to capture those different realities, including diverse experiences with the Paris Declaration principles, to facilitate dialogue on future international development policies and practices.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 5/37

3. Social constructivism was the foundation of Nora Murphy’s (2014) study of homeless youth. The 14 case studies showed diverse experiences and perspectives on homelessness. The study concluded that the unique situation of each homeless youth meant that program responses needed to be socially constructed together with the youth to be meaningful to them and to build trusting adult–youth relationships. (For more on this evaluation, see pp. 194, 626–628.)

SIDEBAR

CONSTRUCTIVIST TRUSTWORTHINESS

The credibility of your findings and interpretations depends on your careful attention to establishing trustworthiness. Lincoln and Guba (1985) describe prolonged engagement (spending sufficient time at your research site) and persistent observation (focusing in detail on those elements that are most relevant to your study) as critical in attending to credibility. “If prolonged engagement provides scope, persistent observation provides depth” (p. 304). With each, time is a major factor in the acquisition of trustworthy data. Time at your research site, time spent interviewing, and time building sound relationships with respondents all contribute to trustworthy data. When a large amount of time is spent with your research participants, they less readily feign behavior or feel the need to do so; moreover, they are more likely to be frank and comprehensive about what they tell you. Lincoln and Guba posited four constructivist criteria as parallel to but distinct from traditional research criteria:

First, credibility (parallel to internal validity) addressed the issue of the inquirer providing assurances of the fit between respondents’ views of their life ways and the inquirer’s reconstruction and representation of same. Second, transferability (parallel to external validity) dealt with the issue of generalization in terms of case-to- case transfer. It concerned the inquirer’s responsibility for providing readers with sufficient information on the case studied such that readers could establish the degree of similarity between the case studied and the case to which findings might be transferred. Third, dependability (parallel to reliability) focused on the process of the inquiry and the inquirer’s responsibility for ensuring that the process was logical, traceable, and documented. Fourth, confirmability (parallel to objectivity) was concerned with establishing the fact that the data and interpretations of an inquiry were not merely figments of the inquirer’s imagination. It called for linking assertions, findings, interpretations, and so on to the data themselves in readily discernible ways. For each of these criteria, Lincoln and Guba also specified a set of procedures that could be used to meet the criteria. For example, auditing was highlighted as a procedure useful for establishing both dependability and confirmability, and member check and peer debriefing, among other procedures, were defined as most appropriate for credibility.

In Fourth Generation Evaluation (1989), Guba and Lincoln reevaluated this initial set of criteria. They explained that trustworthiness criteria were parallel, quasi-foundational, and clearly intended to be analogs to conventional criteria. Furthermore, they held that trustworthiness criteria were principally methodological criteria and thereby largely ignored aspects of the inquiry concerned with the quality of outcome, product, and negotiation. Hence, they advanced a second set of criteria called authenticity criteria, arguing that this second set was better aligned with the constructivist epistemology that informed their definition of qualitative inquiry. (Schwandt, 2007, pp. 299–300)

Continual alertness to your own biases and subjectivity (reflexivity) also assists in producing more trustworthy interpretations. Consider your subjectivity within the context of the trustworthiness of your findings. Ask yourself a series of questions: Whom do I not see? Whom have I seen less often? Where do I not go? Where have I gone less often? With whom do I have special relationships, and in what light would they interpret phenomena? What data-collecting means have I not used that could provide additional insight? Triangulated findings contribute to credibility. Triangulation may involve the use of multiple data collection methods, sources, investigators, or theoretical perspectives. To improve trustworthiness, you can also consciously search for negative cases.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 6/37

Alternative Criteria Review

Exhibit 9.7 (pp. 680–681) presented seven different sets of criteria for judging the quality of qualitative studies. This module reviewed the first two sets of criteria: (1) traditional scientific research criteria versus (2) constructivist criteria. Constructivist criteria emerged from the critique that traditional scientific research criteria were based on quantitative and experimental design thinking that, by the very nature of using those criteria for defining quality, led to qualitative studies being judged inferior. The next module makes the issue of judging quality even more complicated by adding five more sets of alternative and competing criteria.

© 2002 Michael Quinn Patton and Michael Cochran

A Realist Views a Constructivist Proposal

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 7/37

MODULE

79 Alternative and Competing Criteria, Part 2

Artistic, Participatory, Critical Change, Systems, Pragmatic, and Mixed Criteria

The moral and social yearnings of fully realized human beings are not reducible to universal laws and cannot be studied like physics.

—Brooks (2010, p. A27)

This module continues the presentation and discussion of seven alternative sets of criteria for judging the quality of qualitative studies. The previous module covered (1) traditional social science research criteria and (2) social construction and constructivist criteria. This module covers (3) artistic and evocative criteria, (4) participatory and collaborative criteria, (5) critical change criteria, (6) systems and complexity criteria, and (7) pragmatic and utilization-focused criteria. We’ll then examine mixing criteria.

3. Artistic and Evocative Criteria

TRUTH is visceral, palpable, sensuous, wrenching, hormonal, cognitive, cathartic, lyrical, contextual, awakening, fleeting, universal, and debatable. In other words, truth is art.

—From Halcolm’s Ruminations

Researchers and audiences operating from the perspective of traditional scientific research criteria emphasize the scientific nature of qualitative inquiry. Researchers and audiences who view the world through the lens of social construction emphasize qualitative inquiry as a particularly human form of understanding centered on the capacity of people and groups to construct meaning. That brings us to this third alternative, which emphasizes that human beings both think and feel. Traditional social science and constructivist inquiries focus on cognitive, logical, sense-making analyses. The artistic and evocative approaches to qualitative inquiry want to bring forth our emotional selves and do so by integrating art and science. Science makes us think. Great art makes us feel. From the perspective of artistic and evocation qualitative inquirers, great qualitative studies should evoke both understandings (cognition) and feelings (emotions).

Persons are moved by emotion. . . . People are their emotions. To understand who a person is, it is necessary to understand emotion. . . . Emotions cut to the core of people. Within and through emotion people come to define the surface and essential, or core, meanings of who they are. Emotions and moods are ways of disclosing the world for the person. (Denzin, 2009, pp. 1–2)

Artistic and evocative criteria focus on aesthetics, creativity, interpretive vitality, and expressive voice. Case studies become literary works. Poetry or performance art may be used to enhance the audience’s direct experience of the essence that emerges from analysis. Artistically oriented qualitative analysts seek to engage those receiving the work, to connect with them, move them, provoke them, and stimulate them. Creative nonfiction and fictional forms of representation blur the boundaries between what is “real” and what has been

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 8/37

created to represent the essence of a reality. A literal presentation of reality, real (scientific) or perceived (constructivism), yields to artistically created reality. The results may be called creative syntheses, ideal- typical case constructions, scientific poetics, or any number of phrases that suggest the artistic emphasis. Artistic expressions of qualitative analysis strive to provide an experience with the findings where “truth” or “reality” is understood to have a feeling dimension that is every bit as important as the cognitive dimension. Such qualitative inquiry is explicitly sensuous (Stoller, 2004) and emotional (Denzin, 2009).

The performance art of The Vagina Monologues (Ensler, 2001), based on interviews with women about their experiences of coming of age sexually, but presented as theater, offers a prominent example. The audience feels as much as knows the truth of the presentation because of the essence it reveals. In the artistic tradition, the analyst’s interpretive and expressive voice, experience, and perspective may become as central to the work as depictions of others or the phenomenon of interest. Here are some examples of artistic and evocative approaches used in program evaluation.

• A program for low-income, pregnant, drug-addicted teenagers asked the young women to draw pictures of their hearts and tell what the pictures meant. The initial hearts were portrayed as wounded, knifed, torn, mangled, and tortured. Over a period of four months (four drawings and accompanying stories), the pictures showed some sunshine, flowers, rainbows, and, most striking, connections to other hearts. No perfect valentines. Not even close. But to look at those raw drawings was to see hearts healing.

• Theater for development is being used in Nigeria to engage community members in discussing preliminary results from an evaluation, with actors replaying scenarios relating to a program uncovered through fieldwork. By involving community members in role-play, the early findings can be verified or corrected based on the participants’ experiences of the program, and new perspectives can be recorded, which might otherwise have remained dormant (Folorunsho, 2014).

• A team of educational evaluators concerted interviews with students and teachers about an international cross-cultural summer experience into a play dramatizing critical events and key learnings. The play was performed for the school board as the project’s evaluation report.

• Photographs taken before and after Vietnam implemented a motorbike helmet law showed dramatic changes in compliance. (See Exhibit 8.20, pp. 609 ; for other examples of using visuals created to present evaluation findings, see Exhibits 8.21 through 8.27, pp. 610–619.)

Exhibit 9.8 shows how an interview transcript was converted into a poem when presenting the findings, all the better to give the reader a feel for what was said and the affect it carried.

EXHIBIT 9.8 From Interview Transcript to Poem: An Artistic and Evocative Presentation

In May 1994, Corrine Glesne (1997) interviewed Dona Juana, an 86-year-old professor in the College of Education at the University of Puerto Rico.

That she chose a bird to represent her was no surprise. Standing 5 feet tall, very thin (“a problem all my life”), and with bright dark eyes, she was birdlike in appearance. Her office was a nest of books, papers, and folders in organized piles on her large desk, on the beige metal filing cabinets next to the door opposite her desk, in the wooden cabinet along the wall to the right of her desk, on the shelves below the window to her left, and on the two chairs before her desk. There was no sense of disorder, but rather an impression of an archive that would illuminate Dona Juana’s 50 years in research and higher education. (p. 203)

Below is the poem Glesne (1997) created from the interview transcript, followed by a table showing the conversion from transcript to poem.

The Poem

That Rare Feeling

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 9/37

I am a flying bird moving fast seeing quickly looking with the eyes of God, from the tops of trees. How hard for country people picking green worms from fields of tobacco, sending their children to school, not wanting them to suffer as they suffer. In the urban zone, students worked at night and so they slept in school. Teaching was the real university. So I came to study to find out how I could help. I am busy here at the university, there is so much to do. But the University is not the Island. I am a flying bird moving fast, seeing quickly so I can give strength, so I can have that rare feeling of being useful. (pp. 202–203)

Composing the Poem From the Interview Transcript

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 10/37

Crystallization

Sociologist Laurel Richardson (2000b) introduced crystallization as a criterion of quality in artistic and evocative qualitative inquiry, a replacement for triangulation as a criterion.

The scholar draws freely on his or her productions from literary, artistic, and scientific genres, often breaking the boundaries of each of those as well. In these productions, the scholar might have different “takes” on the same topic, what I think of as a postmodernist deconstruction of triangulation. . . . In postmodernist mixed-genre texts, we do not triangulate, we crystallize. . . . I propose that the central image for “validity” for postmodern texts is not the triangle—a rigid, fixed, two-dimensional object. Rather, the central imaginary is the crystal, which combines symmetry and substance with an infinite variety of shapes, substances, transmutations, multidimensionalities, and angles of approach. . . . Crystallization provides us with a deepened, complex, thoroughly partial, understanding of the topic. Paradoxically, we know more and doubt what we know. Ingeniously, we know there is always more to know. (p. 934)

Crystallization’s roots can be traced to

the creative and courageous work of feminist methodologists who blasphemed the boundaries of art and science. . . .

Art and science do not oppose one another; they anchor ends of a continuum of methodology, and most of us situate ourselves somewhere in the vast middle ground. When scholars argue that we cannot include narratives alongside

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 11/37

analysis or poems within grounded theory, they operate under the assumption that art and science negate one another and hence are incompatible, rather than merely differ in some dimensions. . . . My explanation of crystallization assumes a basic understanding of the complexities involved in combining methods and genres from across regions of the continuum. (Ellingson, 2009, pp. 3, 5)

4. Participatory and Collaborative Criteria

To be human is to engage in interpersonal dynamics. Inter: between. Personal: people. Dynamics: forces that produce activity and change. Combining these definitions, interpersonal dynamics are the forces between people that lead to activity and change. Whenever and wherever people interact, these dynamics are at work.

—King and Stevahn (2013, p. 2) Interactive Evaluation Practice

Participatory and collaborative qualitative inquiries have four purposes and justifications:

1. Values premise: The right way to inquire into a phenomenon of interest is to do it with the people involved and affected. This means doing research and evaluation with as opposed to people. It means engaging them as fellow inquirers and coresearchers rather than as research subjects.

SIDEBAR

RIGOR IN ARTISTIC AND EVOCATIVE CRYSTALLIZATION

One of the most helpful (albeit not foolproof) ways to enhance your account and ward off editorial defensiveness toward creative analytic work, in general, and crystallization, in particular, is to be absolutely clear about what you did (and did not do) in producing your manuscript. This includes data collection, analysis, and especially choices made about representation. . . . By explaining my process, I help alleviate suspicions that I took an “anything goes” sloppy attitude toward constructing my representation.

While some colleagues may not like or approve of what you did no matter how you explain it, concise, explicit details of your process make it more difficult for them to dismiss it as careless or random. Accounting for your process (even in an appendix or endnote) constitutes an important nod toward methodological rigor. As many have posited, engaging in creative analytic work should be no less rigorous, exacting, and subject to strict standards of peer evaluation. . . . Moreover, such a roadmap assists others who may seek to follow your lead. . . .

Some suggestions on issues to own: • Explain choices you made in composing narratives, poems, or other artistic work; in other words, how

did you get from data to text? • Describe your standpoint vis-à-vis your topic, not just what it is, but (at least some of) how it shapes

your interactions with your data (e.g., I am a cancer survivor studying clinics so I tend to be more empathetic with patients than health care providers; I am a feminist so I pay a lot of attention to power dynamics).

• Indicate your awareness of and response to ethical considerations about voice, privacy, and responsibility to others. What steps did you take to ensure participant confidentiality? To privilege participants’ voices? Consider how your work might be read in ways that do not reflect your

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 12/37

intentions—for example, what quotes from participants could be taken out of context and used as justification for blaming the victim?—and surround vulnerable voices with preemptory statements that make it more difficult for oppositional forces to excerpt and reinterpret their meaning in regressive ways.

• Detail your analytic procedures. . . . Even if you construct a unique, outside-the-box artistic creation, you should explain your methodology and cite some sources to contextualize your work. Again, this need not interfere with your aesthetic goals; details should be concise and can be placed in an appendix, footnote, or even a separate piece altogether. The goal is to reveal crystallized projects as embodied, imperfect, insightful constructions rather than as immaculate end products. (Ellingson, 2009, pp. 199–120)

2. Quality premise: Data will be better when people who are the focus of the inquiry willingly participate, understand the nature of the inquiry, and agree with the importance of the study. Interviews will be richer and more detailed. Observations will be open and unguarded. Documents will be readily available. Data are better.

3. Reciprocity premise: Researchers get data, publications, knowledge, and career advancement from research and evaluation studies. Those who are the focus of inquiry should benefit as well. As coresearchers, through participation in the inquiry, they learn research skills, learn to think more systematically, and gain knowledge that they can use for their own purposes.

4. Utility premise: In program evaluation and action research inquiries, the findings are more likely to be useful—and actually used—when those who must act on the findings collaborate in generating and interpreting them.

From the classic articulation and justification of Participatory Action Research by William Foote Whyte (1989, 1991) to methods and facilitation guides on how to actually do it (Caister et al., 2011; Hacker, 2013; King & Stevahn, 2013; Pyrch, 2012; Taylor, Suarez-Balcazar, Forsyth, & Kielhofner, 2006), participatory and collaborative engagement has been a major approach to qualitative inquiry. When conducting research in a collaborative mode, professionals and nonprofessionals become coresearchers. Participatory action research encourages collaboration within a mutually acceptable inquiry framework to understand and/or solve organizational or community problems. Chapter 4 includes an in-depth discussion of participatory and collaborative approaches (pp. 213–222), including Exhibit 4.13, Principles of Fully Participatory and Genuinely Collaborative Inquiry (p. 222).

Here’s an example of a participatory and collaborative qualitative inquiry. Robin Boylorn (2008) studied the experiences of black Southern women. She recruited a group of participants from the community in which she had grown up and invited them to share stories about their experiences and lives growing up and raising families in the rural South. She facilitated their interactions together as co-investigators so that they felt “equally invested and equally involved in the process of collecting, writing, interpreting, and editing the stories they wrote” (p. 600). She shared her experiences with the participants, and together they compared and contrasted their ideas and experiences.

Their involvement began during the early stages of recommending other participants and retelling stories in individual and group settings to ensure adequate information was available. As co-investigators their stories were instrumental in establishing and representing a corporate set of themes and experiences. Though the co-researchers in this project were not involved in the writing stages, they did have the opportunity to respond to the stories the author wrote, offering their unique perspectives and feedback as participants in the research and characters in the stories. The resulting research project is a collaboration between the researcher and the researched, including participants as co-researchers. (p. 600)

SIDEBAR

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 13/37

INTERPERSONAL VALIDITY

Educational evaluator Karen Kirkhart (1995) coined the term interpersonal validity: the extent to which an evaluator is able to relate meaningfully and effectively to individuals in the evaluation setting. The interpersonal factor that undergirds interpersonal validity highlights the competence of a participatory evaluator or researcher do two things: (1) interact with people constructively throughout the framing and implementation of an inquiry and (2) create activities and conditions conducive to positive interactions among participants. The interpersonal factor is concerned with creating, managing, and ultimately mastering the interpersonal dynamics that make a collaborative inquiry possible and inform its findings. One is concerned with eventual use, the other with establishing buy-in among participants and a valid inquiry process. (King & Stevahn, 2013, p. 6)

5. Critical Change Criteria

We are distressed by underprivilege. We see gaps among privileged patrons and managers and staff and underprivileged participants and communities. . . . We are advocates of a democratic society.

—Robert Stake (2004, pp. 103–107) Qualitative evaluation pioneer

How Far Dare an Evaluator Go Toward Saving the World?

Those engaged in qualitative inquiry as a form of critical analysis aimed at social and political change eschew any pretense of open-mindedness or objectivity; they take an activist stance. Critical change inquiry aims to critique existing conditions and through that critique bring about change. Critical change criterion is derived from critical theory, which frames and engages in qualitative inquiry with an explicit agenda of elucidating power, economic, and social inequalities. The “critical” nature of critical theory flows from a commitment to go beyond just studying society for the sake of increased understanding. Critical change researchers set out to use inquiry to critique society, raise consciousness, and change the balance of power in favor of those less powerful. Influenced by Marxism, informed by the presumption of the centrality of class conflict in understanding community and societal structures, and updated in the radical struggles of the 1960s, critical theory provides both philosophy and methods for approaching research and evaluation as fundamental and explicit manifestations of political praxis (connecting theory and action), and as change-oriented forms of engagement.

Critical social science and critical social theory attempt to understand, analyze, criticize, and alter social, economic, cultural, technological, and psychological structures and phenomena that have features of oppression, domination, exploitation, injustice, and misery. They do so with a view to changing or eliminating these structures and phenomena and expanding the scope of freedom, justice, and happiness. The assumption is that this knowledge will be used in processes of social change by people to whom understanding their situation is crucial in changing it. (Bentz & Shapiro, 1998, p. 146; Kincheloe & McLaren, 2000)

Critical change has three interconnected elements: (1) inquiry into situations of social injustice, (2) interpretation of the findings as a critique of the existing situation, and (3) using the findings and critique to mobilize and inform change.

Critical theory looks at, exposes, and questions hegemony—traditional power assumptions held about relationships, groups, communities, societies, and organizations—to promote social change. Combined with action research, critical theory questions the assumed power that researchers typically hold over the people they typically research. Thus, critical action research is based on the assumption that society is essentially discriminatory but is capable of becoming less so through purposeful human action.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 14/37

Critical action research also assumes that the dominant forms of professional research are discriminatory and must be challenged. Critical action research takes the concept of knowledge as power and equalizes the generation of, access to, and use of that knowledge. Critical action research is an ethical choice that gives voice to, and shares power with, previously marginalized and muted people. (Davis, 2008, p. 140)

Critical change criteria apply to a number of specialized areas of qualitative inquiry (Given, 2008, pp. 139–179; Schwandt, 2007, pp. 50–55):

In addition, feminist inquiry often includes an explicit agenda of bringing about social change (e.g., Benmayor, 1991; Brisolara, Seigart, & SenGupta, 2104; Hesse-Biber, 2013; Podems, 2014b). Liberation research and empowerment evaluation derive, in part, from Paulo Freire’s philosophy of praxis and liberation education, articulated in his classics Pedagogy of the Oppressed (1970) and Education for Critical Consciousness (1973), still sources of influence and debate (e.g., Glass, 2001). Barone (2000) aspires to “emancipatory educational storysharing” (p. 247). Qualitative studies informed by critical change criteria range from largely intellectual and research-oriented approaches that aim to expose injustices to more activist forms of inquiry that actually engage in bringing about social change. Stephen Brookfield (2004) uses critical theory to illuminate adult education issues, trends, and inequities. Plummer (2011) integrates critical their and queer theory. Caruthers and Friend (2014) bring critical inquiry to online learning and engagement. Crave, Zaleski, and Trent (2014) emphasize the role of critical change in building a more equitable future through participatory program evaluation.

Here are two examples of critical change studies that would expect to be evaluated for quality by critical change criteria (Davis, 2008, p. 141):

1. Martin Diskin worked with policymakers and development agencies in Latin American studies to conduct what they called “power structure research,” in which they exposed injustice as a strategy for building coalitions and motivating movements.

2. Christine Davis’s ethnography of a children’s mental health treatment team was an interdisciplinary research project involving the fields of communication studies, social work, and mental health. Conducted in partnership with community agencies, this research examined issues of power, marginalization, and control within these teams. It suggested a stance toward children and families that rejects the traditional hierarchical medical model of care and instead treats them as unique valuable humans and as equal partners in treatment.

Consequential validity as a critical change criterion for judging a research design or instrument makes the social consequences of its use a value basis for assessing its credibility and utility. Thus, standardized achievement tests are criticized because of the discriminatory consequences for minority groups of educational decisions made with “culturally biased” tests. Consequential validity asks for assessments of who benefits and who is harmed by an inquiry, measurement, or method (Brandon, Lindberg, & Wang, 1993; Messick, 1989; Shepard, 1993).

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 15/37

SIDEBAR

A QUALITATIVE MANIFESTO: A CALL TO ARMS —Norman K. Denzin (2010)

The social sciences . . . should be used to improve quality of life. . . . For the oppressed, marginalized, stigmatized and ignored . . . and to bring about healing, reconciliation and restoration between the researcher and the researched.

—Stanfield (2006, p. 725)

Mills wanted his sociology to make a difference in the lives that people lead. He challenged persons to take history into their own hands. He wanted to bend the structures of capitalism to the ideologies of radical democracy. . . .

I want a critical methodology that enacts its own version of the sociological imagination. Like Mills, my version of the imagination is moral and methodological. And like Mills, I want a discourse that troubles the world, understanding that all inquiry is moral and political.

This book is an invitation and a call to arms. It is directed to all scholars who believe in the connection between critical inquiry and social justice (Denzin, 2010, p. 10).

Qualitative inquiry can contribute to social justice in the following ways:

1. It can help identify different definitions of a problem and/or a situation that is being evaluated with some agreement that change is required. It can show, for example, how battered wives interpret the shelters, hotlines, and public services that are made available to them by social welfare agencies. Through the use of personal experience narratives, the perspectives of women and workers can be compared and contrasted.

2. The assumptions, often belied by the facts of experience, that are held by various interested parties— policy makers, clients, welfare workers, online professionals—can be located and shown to be correct, or incorrect (Becker, 1967, p. 239).

3. Strategic points of intervention into social situations can be identified. Thus, the services of an agency and a program can be improved and evaluated.

4. It is possible to suggest “alternative moral points of view from which the problem, the policy and the program can be interpreted and assessed” (see Becker, 1967, pp. 239–240). Because of its emphasis on experience and its meanings, the interpretive method suggests that programs must always be judged by and from the point of view of the persons most directly affected.

5. The limits of statistics and statistical evaluations can be exposed with the more qualitative, interpretive materials furnished by this approach. Its emphasis on the uniqueness of each life holds up the individual case as the measure of the effectiveness of all applied programs. (Denzin, 2010, pp. 24–25)

6. Systems Thinking and Complexity Criteria In a finite game, it is easy to make sense. Everyone agrees on the goal; the rules are known; and the field of play has clear boundaries. Baseball, football, and bridge are examples of finite games. At one time in the not-so-distant past we expected careers, marriages, parenthood, education, and citizenship to be finite games. When everyone agrees on the rules, and the consequences of our actions are undeniable, responsible people plan for what they want, take steps to achieve it, and enjoy the fruits of their labor. We know what it takes to make sense in a finite game.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 16/37

Most of us realize we are playing in a very different game. We are playing in an infinite game. In which the boundaries are unclear or nonexistent, the scorecard is hidden, and the goal is not to win but to keep the game in play. There are still rules, but the rules can change without notice. There are still plans and playbooks, but many games are going on at the same time, and the winning plans can seem contradictory. There are still partners and opponents, but it is hard to know who is who, and besides that, the “who is who” changes unexpectedly.

—Glenda Eoyang and Royce Holladay (2013, p. 4)

Adaptive Action: Leveraging Uncertainty in Your Organization

Studying “infinite games” in highly dynamic situations characterized by uncertainty and rapid change creates special challenges for qualitative inquiry. Systems thinking and complexity concepts offer a framework for studying such situations, tools for both inquiry and “coping with chaos” (Eoyang, 1997), and criteria for deciding whether such studies are of high quality. To be credible to systems thinkers and complexity scientists, the qualitative inquiry must capture, describe, map, and analyze, and map systems of interests; must attend to interrelationships, capture diverse perspectives, attend to emergence, and be sensitive to and explicit about boundary implications; and must document nonlinearities, adapt the inquiry in the face of uncertainties, and describe systems changes and their implications. In so doing, the explanatory approach moves from attribution to contribution analysis (see pp. 596–597 in Chapter 8).

Chapter 3 discussed systems theory and complexity theory as distinct, though intersecting, theoretical frameworks (see pp. 139–151). Exhibit 3.14 presents complexity theory concepts and qualitative inquiry implications (pp. 147–148). Exhibit 3.16 presents the relationship of systems theory to complexity theory (p. 150). • Systems theory inquiry questions: How and why does this system function as it does? What are the

system’s boundaries and interrelationships, and how do these affect perspectives about how and why the system functions as it does?

• Complexity theory inquiry question: How can the emergent and nonlinear dynamics of complex adaptive systems be captured, illuminated, and understood?

For my purpose here, namely, differentiating distinct sets of criteria by which to judge the quality and credibility of various approaches to qualitative inquiry, the core systems and complexity dimensions can be integrated, as they are in Exhibit 9.7. That said, the systems field exemplifies the challenge of settling on some definitive set of quality criteria for judging qualitative inquiry, especially using a systems and complexity framing, because there are multiple approaches within the systems field (e.g., Hieronymi, 2013), each of which would assert and favor particular criteria unique to that perspective.

Systemic inquiry covers a wide range methodologies, methods, and techniques with a strong focus on the behaviors of complex situations and the meanings we draw from those situations. It spans both the qualitative and quantitative research method domains but also includes approaches that fit neither category nor both categories. . . .

Any attempt to summarize a trans discipline like systemic inquiry is fraught with difficulties. Despite relatively simple origins, the field has sprawled many directions so that no single, universally accepted theory has emerged, and neither are there universally agreed definitions of basic concepts such as what is and what is not a system. Although we will find many definitions in the systems literature, many authors argue that single fixed definitions promote the kind of reductionist thinking that runs counter to systemic principles. Instead, they argue, the field should promote debates around methodological principles to create learning rather than fixed definitions—what Kurt Richardson calls “critical pluralism.” (Williams, 2008, p. 858)

Thinking Systemically So as not to get lost in or overwhelmed by different approaches to systems, let me close this section with an example grounded in the basics of attending to interrelationships, boundaries, perspectives, and emergence.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 17/37

An exemplar of applying systems thinking to understand an issue is the analysis done by Christopher Wells (2012) of the role and impact of the automobile in the United States. His analysis begins before there were automobiles (what complexity theorists call initial conditions). He examines emergent land use patterns in the nineteenth century, sanitation problems in cities, the development of agricultural markets, the role of horses in transportation, the influence of train routes, the challenges of riding bicycles on rutted and muddy roads, the function of farmers in maintaining roads along their farms, population growth, and many other factors that established the initial conditions that automobiles emerged into. To understand the automobile in American society, culture, politics, and economics, you must look at the systems before the automobile existed (transportation, commerce, public health, political jurisdictions, land use, and community values as starting places) and continue to examine those systems and their interactions through to the present day. The irony is that engaging and thinking through those complex interactions yields extraordinary clarity.

7. Pragmatic, Utilization-Focused Criteria

Usefulness! It is not a fascinating word, and the quality is not one of which the aspiring spirit can dream o’ nights, yet on the stage it is the first thing to aim at.

—Dame Ellen Terry (1847–1928) Leading Shakespearean actress in Britain

How use doth breed a habit in a man. —William Shakespeare

It is intriguing to find a great Shakespearean actress lauding usefulness as a matter of prime concern in her performances. Based on her musings about what she aspired to, usefulness concerned using anything and everything at her disposal to bring the play to life and connect with the audience. This is, perhaps, an artistic and evocative view of usefulness, but it also connotes a practical twist that makes for a provocative introduction to our final set of quality criteria: pragmatic, utilization-focused criteria.

• Observations of a high school cafeteria revealed substantial food waste. Interviews showed why the students were so dissatisfied with the food offered. The school had recently experienced an influx of immigrants from Asian countries, where people preferred rice rather than potatoes and bread. The results were used by school officials and the student council to advocate for more culturally appropriate food. Their efforts were successful.

• An early-childhood parent education program was experiencing a high dropout rate. Fewer than half the parents who started the program completed it. Interviews with the dropouts revealed that the program materials being used were academic and difficult for poorly educated parents to understand. Materials were revised to be more accessible and appropriate for parents with lower reading skills.

• The agricultural extension service serving a remote rural area in West Africa had very poor attendance at field trips aimed at helping farmers improve their basic growing practices for the subsistence crops sorghum and millet. Interviews with farmers revealed that they received no advance notice when the field training would occur. A radio news program agreed to announce extension field visits. Attendance increased significantly.

SIDEBAR

DIVERSE METHODS BASED ON SYSTEMS AND QUALITY CONCEPTS

Systems and complexity concepts manifest nuances of difference under varying application frameworks:

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 18/37

• System dynamics: Focuses on the interrelationships between components of a situation, especially the consequences of feedback and delay

• Viable systems: Explores relationships that support an organization’s viability within its environment • Soft systems methodology: Looks at a situation from multiple viewpoints to understand and anticipate

both interactions and unanticipated consequences • Critical systems heuristics: Focuses on ethical issues, marginalization of people, and ideas of power

and coercion • Activity systems: Draws on cultural-historical activity theory to identify and track roles, tools, past

features and dynamics, contradictions, tensions, conflicts, disturbances, innovations, processes, and learning opportunities

• Complex adaptive systems: Independent and interdependent elements or agents adapting to each other, self-organizing and emergent patterns, and nonlinear dynamics

• Network analysis: Examines dynamic interactions, connectivity, processes, and outcomes among a group or system of interconnected people or things. (Network Impact and Center for Evaluation Innovation, 2014)

SOURCES: Williams (2005) and Williams and Hummelbrunner (2011).

These are examples of simple inquiries aimed at providing practical and useful information to solve immediate problems. The pragmatic, utilization-focused criteria emphasize qualitative data generated to solve problems and inform decisions. This means focusing the inquiry on informing action and decisions. To be useful, specific intended users must be identified and their information needs met. Interactive engagement with intended users enhances relevance and use. Findings and feedback are timed to support use. Findings must be actionable and results understandable. The methods used need to be credible to those who will use the findings. Epistemologically, the orientation of pragmatic qualitative inquiry is that what is useful is true.

Pragmatic, utilization-focused inquiry begins with the premise that studies should be judged by their utility and actual use; therefore, evaluators and researchers should facilitate the inquiry process and design any study with careful consideration of how everything that is done, from beginning to end, will affect use. Use concerns how real people in the real world apply findings and experience the inquiry process. Therefore, the focus is on intended use by intended users. Since no study can be value-free, utilization-focused inquiry answers the question of whose values will frame the study by working with clearly identified, primary intended users who have the responsibility to apply findings and take action (Patton, 2008, 2012a).

SIDEBAR

PRAGMATIC EVALUATION STANDARDS

The evaluation profession has adopted standards that call for evaluations to be useful, practical, ethical, accurate, and accountable (Joint Committee on Standards, 2010). In the 1970s, as evaluation was just emerging as a field of professional practice, many evaluators took the position of traditional researchers that their responsibility was merely to design studies, collect data, and publish findings; what decision makers did with those findings was not their problem. This stance removed from the evaluator any responsibility for fostering use and placed all the “blame” for nonuse or underutilization on decision makers. Moreover, before the field of evaluation identified and adopted its own standards, criteria for judging evaluations could scarcely be differentiated from criteria for judging research in the traditional social and behavioral sciences, namely, technical quality and methodological rigor. Utility was largely ignored. Methods decisions dominated the evaluation design process. Validity, reliability, measurability, and generalizability were the dimensions that received the greatest attention in judging evaluation research proposals and reports. Indeed, evaluators concerned about increasing a study’s usefulness often

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 19/37

called for ever more methodologically rigorous evaluations to increase the validity of findings, thereby supposedly compelling decision makers to take findings seriously.

By the late 1970s, however, program staff and funders were becoming openly skeptical about spending scarce funds on evaluations that they couldn’t understand and/or found irrelevant. Evaluators were being asked to be “accountable,” just as program staff were supposed to be accountable. The questions emerged with uncomfortable directness: Who will evaluate the evaluators? How will evaluation be evaluated? It was in this context that professional evaluators began discussing standards.

The most comprehensive effort at developing standards was hammered out over five years by a 17- member committee appointed by 12 professional organizations with input from hundreds of practicing evaluation professionals. Just prior to publication, Dan Stufflebeam, chair of the committee, summarized the results as follows:

The standards that will be published essentially call for evaluations that have four features. These are utility, feasibility, propriety and accuracy. And I think it is interesting that the Joint Committee decided on that particular order. Their rationale is that an evaluation should not be done at all if there is no prospect for its being useful to some audience. Second, it should not be done if it is not feasible to conduct it in political terms, or practicality terms, or cost effectiveness terms. Third, they do not think it should be done if we cannot demonstrate that it will be conducted fairly and ethically. Finally, if we can demonstrate that an evaluation will have utility, will be feasible and will be proper in its conduct, then they said we could turn to the difficult matters of the technical adequacy of the evaluation. (Stufflebeam, 1980, p. 90)

High-Stakes Debate: What Counts as Credible Evidence, and by What Criteria Shall Credibility Be Judged? The seven frameworks just reviewed show the range of criteria that can be brought to bear in judging a qualitative study. They can also be viewed as “angles of vision” or “alternative lenses” for expanding the possibilities available, not only for critiquing inquiry but also for undertaking it. What is most important to understand is that researchers and evaluators attending to and operating with any one of the seven different sets of quality criteria will (a) ask different questions, (b) use different methods, (c) follow different analytical processes, (d) report their findings in different ways, and (e) aim their claims of credibility to different audiences. These are not just academic distinctions. The differences are far from trivial. Quite the contrary, the different orientations have far-reaching implications for every aspect of inquiry. These different quality criteria constitute the underpinnings of significantly different ways of engaging in qualitative inquiry. At the heart of all scientific debate throughout history has been this burning question: What counts as credible evidence and by what criteria shall credibility be judged?

Nor is the debate about what counts as credible evidence just a matter of contention among scientists. Policymakers and politicians have gotten involved. It is down-and-dirty politics with millions of dollars in government-funded and philanthropic-sponsored research at stake (Denzin & Giardina, 2006, 2008; Donaldson, Christie, & Mark, 2008; Scriven, 2008). This means that advocates of qualitative inquiry must understand and be prepared to enter the debate about the politics of evidence (e.g., Eyben, 2013; Nutley et al., 2013; Schorr, 2012). In so doing, understanding the variety of approaches to qualitative inquiry, and which approaches are legitimate by what criteria, will become part of the debate.

So make no mistake about it, advocates of one particular set of criteria are likely to be vociferous critics of alternative criteria. Using the research process as an intervention to correct injustices and foment change is anathema to those who advocate traditional scientific research criteria as the only acceptable standards for judging quality. Those traditional criteria insist on a clear line of demarcation between studying a phenomenon (basic research and independent, external evaluation) versus engaging in change through the research process (advocacy). On the other hand, attempts to make traditional scientific research criteria the only legitimate approach to government-funded research are criticized as narrow-minded, self-serving

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 20/37

political advocacy that constitutes “a conservative challenge to qualitative inquiry” (Denzin & Giardina, 2006, p. x). Constructivists generated their criteria of quality as a direct reaction to what they considered the gross inadequacies and methodological distortions of traditional scientific research criteria, which are essentially derived from the experimental/quantitative paradigm (see pp. 87–95). Thus, they systematically set out to replace traditional research criteria like validity and reliability with trustworthiness and authenticity (Lincoln & Guba, 1985, 1986). Advocates of artistic and evocative approaches attack both traditional research and constructivism as emotionally void. Traditional researchers have been disinclined to use participatory and collaborative approaches, sometimes believing that involving nonresearchers in research inevitably leads to poorer quality; in other cases, it’s a matter of lacking incentives, capacity, or interest. Pragmatic, utilization- focused inquiries are attacked for being theoretically useless, unscholarly, and so practical as to be worthless for generating explanations or generalizations. Many traditional researchers don’t even consider action research worthy of the name “research.”

Choosing a Framework Within Which to Work

Which criteria you choose to emphasize in your work will depend on the purpose of your inquiry, the values and perspectives of the audiences for your work, and your own philosophical and methodological orientation. Operating within any particular framework and using any specific set of criteria will invite criticism from those who judge your work from a different framework and with different criteria. (For examples of the vehemence of such criticisms between those using traditional social science criteria and those using artistic narrative criteria, see Bochner, 2001; English, 2000.) Understanding that criticisms (or praise) flow from criteria can help you anticipate how to position your inquiry and make explicit what criteria to apply to your own work as well as what criteria to offer others given the purpose and orientation of your work.

The profession of program evaluation is a microcosm of these larger divisions. Program evaluation is a diverse, multifaceted profession manifesting many different models and approaches (Christie & Alkin, 2013; Fitzpatrick, Sanders, & Worthen, 2010; Funnell & Rogers, 2011; Patton, 2008; Stufflebeam, Madeus, & Kellaghan, 2000). All seven alternative quality criteria are advocated by various evaluation theorists, methodologists, and practitioners.

Any particular evaluation study has tended to be dominated by one set of criteria, with a second set as possibly secondary. For example, a primarily constructivist approach might add some artistic techniques as supporting methods. An evaluation dominated by the traditional scientific research approach might have a section dedicated to dealing with pragmatic issues. Exhibit 9.9 shows how the seven frameworks can be found in the approaches of various evaluation theorists, methodologists, and practitioners.

Clouds and Cotton: Mixing and Changing Perspectives

While each set of criteria manifest a certain coherence, many researchers mix and match approaches. The work of Tom Barone (2000), for example, combines aesthetic, political (critical change), and constructivist elements. Denzin’s Performance Ethnography (2003) uses artistic and evocative approaches to foment and contribute to “radical social change, to economic justice, to a culture of politics that extends critical race theory and the principles of a radical democracy to all aspects of society” (p. 3). A team of evaluators collaborated to integrate constructivism, participatory evaluation, critical change, and a utilization focus (evaluations for improvement):

EXHIBIT 9.9 Alternative Quality Criteria Applied to Program Evaluation

Program evaluation is a diverse, multifaceted profession manifesting many different models and approaches (Alkin & Christie, 2013; Fitzpatrick, Sanders, & Worthen, 2010; Funnell & Rogers, 2011; Patton, 2008; Stufflebeam, Madeus, & Kellaghan, 2000). All seven alternative quality criteria are advocated by various evaluation theorists, methodologists, and practitioners.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 21/37

Evaluations for improvement, understanding lived experience, or advancing social justice are fundamentally participatory, involving key stakeholders in critical decisions about the evaluation’s agenda, direction, and use. Such a principle is rooted epistemologically in the importance of understanding multiple perspectives and experiences in evaluation, and also politically in the importance of democratic inclusion.

—Whitmore et al. (2006, p. 341)

As an evaluator, I have worked with mixed criteria from all seven frameworks to match particular designs to the needs and interests of specific stakeholders and clients (Patton, 2008). But mixing and combining criteria means dealing with the tensions between them. After reviewing the tensions between traditional social science criteria and postmodern constructivist criteria, narrative researchers Lieblich, Tuval-Mashiach, and Zilber (1998) attempted “a middle course,” but that middle course reveals the very tensions they were trying to supersede as they worked with one leg in each camp.

We do not advocate total relativism that treats all narratives as texts of fiction. On the other hand, we do not take narratives at face value, as complete and accurate representations of reality. We believe that stories are usually

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 22/37

constructed around a core of facts or life events, yet allow a wide periphery for freedom of individuality and creativity in selection, addition to, emphasis on, and interpretation of these “remembered facts.” . . .

Life stories are subjective, as is one’s self or identity. They contain “narrative truth” which may be closely linked, loosely similar, or far removed from “historical truth.” Consequently, our stand is that life stories, when properly used, may provide researchers with a key to discovering identity and understanding it—both in its “real” or “historical” core, and as narrative construction. (p. 8)

Traditional scientific research criteria and critical change criteria are polar opposites. The same study cannot aspire to independence, objectivity, and a primary focus on contributing to theory while also being deeply engaged in using the inquiry process to foment change and ameliorate oppression. Mixing methods (qualitative and quantitative) is one thing. Mixing criteria of quality is a bit more challenging, one might even say daunting. Certainly, constructivist, artistic, and participatory criteria can be intermingled. But traditional scientific research criteria are less amenable to comingling.

The remainder of this chapter will elaborate some of the most prominent of these competing criteria that affect judgments about the quality and credibility of qualitative inquiry and analysis. But it’s not always easy to tell whether someone is operating from a realist, constructionist, artistic, activist, or evaluative framework. Indeed, the criteria can shift quickly. Consider this example. My six-year-old son, Brandon, was explaining a geography science project he had done for school. He had created an ecological display out of egg cartons, ribbons, cotton, bottle caps, and styrofoam beads. “These are three mountains and these are four valleys,” he said, pointing to the egg cup arrangement. “And is that a cloud?” I asked, pointing to the big hunk of cotton. He looked at me, disgusted, as though I’ve just said about the dumbest thing he’s ever heard. “That’s a piece of cotton, Dad.”

Foreshadowing research rigor mortis, MQP Rumination # 9 in the next module.

SOURCE: © Chris Lysy—freshspectrum.com

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 23/37

MODULE

80 Credibility of the Inquirer

The previous modules in this chapter have reviewed strategies for enhancing the quality and credibility of qualitative analysis: selecting appropriate criteria for judging quality, searching for rival explanations, explaining negative cases, triangulation, and keeping data in context. Technical rigor in analysis is a major factor in the credibility of qualitative findings. This section now takes up the issue of how the credibility of the inquirer affects the way findings are received.

One barrier to credible qualitative findings stems from the suspicion that the analyst has shaped findings according to her or his predispositions and biases. Whether this may have happened unconsciously, inadvertently, or intentionally (with malice and forethought) is not the issue. The issue is how to counter such a suspicion before it takes root. One strategy involves discussing your predispositions and making biases explicit, to the extent possible. This involves systematic and studious reflexivity (see pp. 70–74).Another approach is engaging in mental cleansing processes (e.g., epoche in phenomenological analysis, p. 575). Or one may simply acknowledge one’s orientation as a feminist researcher (Podems, 2014b) or critical theorist and move on from there. The point is that you have to address the issue of your credibility.

The Researcher as the Instrument in Qualitative Inquiry

Because the researcher is the instrument in qualitative inquiry, a qualitative report should include some information about you, the researcher. What experience, training, and perspective do you bring to the study? Who funded the study and under what arrangements with you? How did you gain access to the study site and the people observed and interviewed? What prior knowledge did you bring to the research topic and study site? What personal connections do you have to the people, program, or topic studied? For example, suppose the observer of an Alcoholics Anonymous program is a recovering alcoholic. This can either enhance or reduce credibility depending on how it has enhanced or detracted from data gathering and analysis. Either way, the analyst needs to deal with it in reporting findings. In a similar vein, it is only honest to report that the evaluator of a family counseling program was going through a difficult divorce at the time of fieldwork.

No definitive list of questions exists that must be addressed to establish investigator credibility. The principle is to report any personal and professional information that may have affected data collection, analysis, and interpretation—either negatively or positively—in the minds of users of the findings. For example, health status should be reported if it affected one’s stamina in the field. Were you sick part of the time? Let’s say that the fieldwork for evaluation of an African health project was conducted over three weeks, during which time the evaluator had severe diarrhea. Did that affect the highly negative tone of the report? The evaluator said it didn’t, but I’d want to have the issue out in the open to make my own judgment. Background characteristics of the researcher (e.g., gender, age, race, and/or ethnicity) may be relevant to report in that such characteristics can affect how the researcher was received in the setting under study and what sensitivities the inquirer brings to the issues under study.

In preparing to interview farm families in Minnesota, I began building up my tolerance for strong coffee a month before the fieldwork. Being ordinarily not a coffee drinker, I knew my body would be jolted by 10 to 12 cups of coffee a day doing interviews in farm kitchens. In the Caribbean, I had to increase my tolerance for rum because some farmer interviews took place in rum shops. These are matters of personal preparation— both mental and physical—that affect perceptions about the quality of the study. Preparation and training for fieldwork, discussed at the beginning of Chapter 6, should be reported as part of the study’s methodology.

Reflexivity and Intellectual Rigor

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 24/37

(Othello to Iago, interpreting what it means for someone to mutter something while sleeping)

But this denoted a foregone conclusion. —William Shakespeare

(Othello to Iago, interpreting what it means for someone to mutter something while sleeping)

The credibility of qualitative inquiry is so closely connected to the credibility of the person or team conducting the inquiry that the quality of reflexivity and reflectivity offered in a report is a window into the thinking processes that are the bedrock of qualitative analysis. Essentially, reflexivity involves turning qualitative analysis on yourself. Who are you, and how has who you are affected what you’ve found and reported in the study? This puts your intellectual rigor on display. The very notion of intellectual rigor connotes that as important as it is to employ systematic analytical strategies and techniques, the effectiveness and quality of those strategies and techniques depend on the quality of thinking that directs them. Which brings me to this chapter’s rumination: Avoiding Research Rigor Mortis.

MQP Rumination # 9

Avoiding Research Rigor Mortis

I am offering one personal rumination per chapter. These are issues that have persistently engaged, sometimes annoyed, occasionally haunted, and often amused me over more than 40 years of research and evaluation practice. Here’s where I state my case on the issue and make my peace.

Look for a pattern in what follows. See if you detect a theme.

Rigor (definition). Unyielding or inflexible; the quality of being extremely thorough, exhaustive, or accurate; being strict in conduct, judgment, and decision (Oxford Dictionary); scrupulous or inflexible accuracy or adherence (Random House Dictionary)

Measurement rigor. The underlying psychometric properties of a measure and its ability to fully and meaningfully capture the relevant construct; the fact that data have been collected in essentially the same manner, across time, the program, and jurisdictions, adds methodological rigor; the reliability and validity of instruments

(Weitzman & Silver, 2012)

Research design rigor. The true experiment (randomized controlled trials) as the optimal (gold standard) design for developing evidence-based practice (Ross, Barkaoui, & Scott, 2007)

Methodological rigor. Design elements that support strong causal attributions and analytical generalization (Chatterji, 2007; Coryn, Schröter, & Hanssen, 2009)

Evaluation research rigor. Evidence testing the extent to which valid and reliable measures of program outcomes can be directly and confidently attributed to a standardized, high-fidelity, consistently implemented program intervention; the most rigorous evaluation is the randomized controlled trial (Chatterji, 2007; Henry, 2009; Ross et al., 2007; Rossi et al., 2004); “methodological rigor can be assessed from the evaluation plan and the quality of the evaluation’s implementation” (Braverman, 2013, p. 101)

Analytical rigor. “Meticulous adherence to standard process . . .; scrupulous adherence to established standards for the conduct of work” (Zelik, Patterson, & Woods, 2007, p. 1)

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 25/37

Rigor mortis. Latin: rigor “stiffness,” mortis “of death”—one of the recognizable signs of death, caused by chemical changes in the muscles after death, causing the limbs of the corpse to become stiff and difficult to move or manipulate

Research rigor mortis. Rigid designs, rigidly implemented, then rigidly analyzed through standardized, rigidly prescribed operating procedures and judged hierarchically by standardized, rigid criteria, thereby manifesting rigorism at every stage

Rigorism. Extreme strictness; no course may be followed that is contrary to doctrine (Random House Dictionary)

Research rigorism. Technicism—reducing research to “the application of techniques or the following of rules” (Hammersley, 2008b, p. 31)

Did you find the pattern? Did you detect a theme? Read on for the countertheme. (A countertheme is like a counterfactual: a theme that might be dominant, even should be dominant, in an alternate universe where the dominant theme is not so dominant.)

The Problem

“The Problem of Rigor in Qualitative Research”—that’s the title of a classic article (Sandelowski, 1986) and a common refrain in textbooks about research methods. The “problem,” it turns out, is that by traditional and dominant definitions of rigor, qualitative methods are inferior. But different criteria for what constitutes methodological quality lead to different judgments about rigor, the central point of this chapter. “The ‘problem of multiple standards’ describes the inherent difficulties in selecting which, among many viable candidates, is the standard process to which performance should be compared” (Zelik, Patterson, & Woods, 2007, p. 2). Rigor begets credibility. Different criteria for what constitutes methodological quality and rigor will yield different judgments about credibility. That much is straightforward.

The larger problem, it seems to me, is the focus on methods and procedures as the basis for determining quality and rigor. The notion that methods are more or less rigorous decouples methods from context and the thinking process that determined what questions to ask, what methods to use, what analytical procedures to follow, and what inferences to draw from the findings. Avoiding research rigor mortis requires rigorous thinking.

Rigorous Thinking No problem can withstand the assault of sustained thinking.

—Voltaire (1694–1778) French philosopher

Rigorous thinking combines (a) critical thinking, (b) creative thinking, (c) evaluative thinking, (d) inferential thinking, and (e) practical thinking. Critical thinking demands questioning assumptions; acknowledging and dealing with preconceptions, predilections, and biases; diligently looking for negative and disconfirming cases that don’t fit the dominant pattern; conscientiously examining rival explanations; relentlessly seeking diverse perspectives; and analyzing what and how you think, why you think that way, and the implications for your inquiry (Kahneman, 2011; Klein, 2011; Loseke, 2013).

Creative thinking invites putting the data together in new ways to see the interactions among separate findings more holistically; synthesizing diverse themes in a search for coherence and essence while simultaneously developing comfort with ambiguity and uncertainty in the messy, complex, and dynamic real work; distinguishing signal from noise while also learning from the noise; asking wicked questions that enter into the intersections and tensions between the search for coherent meaning and persistent uncertainties and ambiguities; bringing artistic, evocative, and visualization techniques to data analysis and presentations; and inviting outside-the-box, off-the-wall, and beyond-the-ken perspectives and interpretations.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 26/37

Evaluative thinking forces clarity about the inquiry purpose, who it is for, with what intended uses, to be judged by what quality criteria; it involves being explicit about what criteria are being applied in framing inquiry questions, making design decisions, determining what constitutes appropriate methods, and selecting and following analytical processes and being aware of and articulating values, ethical considerations, contextual implications, strengths and weaknesses of the inquiry, and potential (or actual) misinterpretations, misuses, and misapplications. In contrast with the perspective of rigor as strict adherence to a standardized process, evaluative thinking emphasizes the importance of understanding the sufficiency of rigor relative to context and situational factors (Clarke, 2005; Patton, 2012a).

Inferential thinking involves examining the extent to which the evidence supports the conclusions reached. Inferential thinking can be deductive, inductive, or abductive—and often draws on and creatively integrates all three analytical processes—but at the core, it is a fierce examination of and allegiance to where the evidence leads.

A rigorously conducted evaluation will be convincing as a presentation of evidence in support of an evaluation’s conclusions and will presumably be more successful in withstanding scrutiny from critics. Rigor is multifaceted and relates to multiple dimensions of the evaluation. . . . The concept of rigor is understood and interpreted within the larger context of validity, which concerns the “soundness or trustworthiness of the inferences that are made from the results of the information gathering process” (Joint Committee on Standards for Educational Evaluation, 1994, p. 145). . . . There is relatively broad consensus that validity is a property of an inference, knowledge claim, or intended use, rather than a property either of a research or evaluation study from the study’s findings. (Braverman, 2013, p. 101)

In reflecting on and writing about “what counts as credible evidence in applied research and evaluation practice,” Sharon Rallis (2009), former president of the AEA and experienced qualitative researcher, emphasized rigorous reasoning: “I have come to see a true scientist [italics added], then, as one who puts forward her findings and the reasoning that led her to those findings for others to contest, modify, accept, or reject” (p. 171).

Practical thinking calls for assiduously integrating theory and practice, examining real-world implications of findings, inviting interpretations and applications from nonresearchers (e.g., community members, program staff, and participants) who can and will apply to the data what ordinary people refer to as “common sense”; and applying real-world criteria to interpreting the findings, criteria like understandability, meaningfulness, cost implications, and implications in addressing societal issues and problems.

What’s at Stake? My words fly up, my thoughts remain below:

Words without thoughts, never to heaven go.

—William Shakespeare (1564–1616) The king in Hamlet

As I noted in Chapter 4, and is worth repeating here, philosopher Hannah Arendt (1968) concluded that to resist efforts by the powerful to deceive and control thinking, people need to practice thinking: “Experience in thinking . . . can be won, like all experience in doing something, only through practice, through exercises” (p. 4).

Regardless of what one thinks of the U.S. invasion of Iraq to depose Saddam Hussein in 2003, both those who supported the war and those who opposed it ultimately agreed that the intelligence used to justify the invasion was deeply flawed and systematically distorted (U.S. Senate Select Committee on Intelligence, 2004). Under intense political pressure to show sufficient grounds for military action, those charged with analyzing and evaluating intelligence data began doing what is sometimes called cherry- picking or stove-piping—selecting and passing on only those data that support preconceived positions and ignoring or repressing all contrary evidence (Hersh, 2003; Tan, 2014; Zelik et al., 2007). The failure

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 27/37

of the intelligence community to appropriately and accurately assess whether Iraq had weapons of mass destruction was not a function of poor data but of weak analysis, political manipulation of the analysis process, and a fundamental failure to think critically, creatively, evaluatively, and practically. The generation of the Rigor Attribute Model to support more rigorous intelligence analysis and restore credibility to the intelligence community focuses on rigorous thinking (Zelik et al., 2007; see Exhibit 9.5, pp. 675–677).

Despite the etymological implication that to be rigorous is to “be stiff,” expert information analysis processes often are not rigid in their application of a standard process, but rather, flexible and adaptive to highly dynamic environments. In information analysis, judgment of rigor reflects a relationship in the appropriateness of fit between analytic processes and contextual requirements. Thus, as supported by this and other research, rigor is more meaningfully viewed as an assessment of degree of sufficiency, rather than degree of adherence to an established analytic procedure. (Zelik, Patterson, & Woods, 2007, p. 1)

The phrase “degree of sufficiency” as a criterion for assessing rigor refers to an evaluation of the extent to which a multidimensional, multiperspectival, and critical thinking process was followed determinedly to yield conclusions that best fit the data, and therefore findings that are credible to and inspire confidence among those who must use the findings.

Bottom-Line Conclusion

Methods do not ensure rigor. A research design does not ensure rigor. Analytical techniques and procedures do not ensure rigor. Rigor resides in, depends on, and is manifest in rigorous thinking—about everything, including methods and analysis.

The thread that runs through this rumination is the importance of intellectual rigor. There are no simple formulas or clear-cut rules about how to do a credible, high-quality analysis. The task is to do one’s best to make sense of things. A qualitative analyst returns to the data over and over again to see if the constructs, categories, interpretations, and explanations make sense—if they sufficiently reflect the nature of the phenomena studied. Creativity, intellectual rigor, perseverance, insight—these are the intangibles that go beyond the routine application of scientific procedures. It is worth quoting again Nobel prize–winning physicist Percy Bridgman: “There is no scientific method as such, but the vital feature of a scientist’s procedure has been merely to do his utmost with his mind, no holds barred” (quoted in Waller, 2004, p. 106).

Varieties of and Concerns About Reactivity: How What We See and Do Affects What Is Seen and Done

Nasrudin denied that he was a fisherman. From a passing tourist he had heard of something called philanthropy and, feeling transformed by what he had learned, he instantly adopted the moniker for himself. He explained to his fellow villagers: “When we see a problem that needs solving, it is wrong to just stand by and observe as scholars are wont to do. We must react. It is wrong to remain passive and detached in the face of need and noble to render help.”

“I am a philanthropist. Each day I strive to help fish that are drowning in the lake. I save them. I throw out my net and the fish rush in. I quickly put the many fish I’ve rescued on the dry ground, where they dance about in joy. But the dancing soon exhausts them and before long they cease to move. Alas, they dance themselves to death.”

“It is sad, but it is also wrong not to honor their struggle. So I take the dead fish to market where people contribute money to my effort to save more fish in exchange for my gifts to them of those fish who have lost the struggle. With the financial tokens of appreciation I receive for my charitable work, I purchase more nets so I can rescue more fish.”

—From Halcolm’s Chronicles of Lessons Learned: Teach a Man to Fish

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 28/37

SIDEBAR

IN-DEPTH REFLEXIVITY: GUIDELINES FOR QUALITY IN AUTOBIOGRAPHICAL FORMS OF SELF-STUDY RESEARCH

• Autobiographical self-studies should ring true and enable connection. • Self-studies should promote insight and interpretation. • Autobiographical self-study research must engage history forthrightly, and the author must take an

honest stand. • Authentic voice is a necessary but not sufficient condition for the scholarly standing of a biographical

self-study. • The autobiographical self-study researcher has an ineluctable obligation to seek to improve the

learning situation not only for the self but also for the other. • Powerful autobiographical self-studies portray character development and include dramatic action:

Something genuine is at stake in the story. • Quality autobiographical self-studies attend carefully to persons in context or setting. • Quality autobiographical self-studies offer fresh perspectives on established truths. • To be scholarship, edited conversation or correspondence must not only have coherence and structure

but that coherence and structure should also provide argumentation and convincing evidence. • Interpretations made of self-study data should not only reveal but also interrogate the relationships,

contradictions, and limits of the views presented (adapted from Bullough & Pinnegar, 2001, pp. 13– 21).

Considering and Reporting Investigator Effects: Varieties of Reactivity Reflectivity includes considering and reporting how your presence as an observer or evaluator may have affected what you observed. There are four primary ways in which the presence of an outside observer, or the fact that an evaluation is taking place, can affect, and possibly distort, the findings of a study, namely,

1. reactions of those in the setting (e.g., program participants and staff) to the presence of the qualitative fieldworker;

2. changes in you, the fieldworker (the measuring instrument), during the course of the data collection or analysis—that is, what has traditionally been called instrumentation effects in quantitative measurement;

3. the predispositions, selective perceptions, and/or biases you might bring to the inquiry that become evident to others during data collection; and

4. researcher incompetence (including lack of sufficient training or preparation).

Reactivity

All accounts produced by researchers must be interpreted within the context in which they were generated. Interpretations must examine, as carefully as possible, how the presence of the researcher, the context in which data were obtained, and so on shaped the data.

—Schwandt (2007, p. 256)

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 29/37

Problems of reactivity are well documented in the anthropological literature, which is one of the prime reasons why qualitative methodologists advocate long-term observations that permit an initial period during which observers and the people in the setting being observed get a chance to get used to each other. This increases trustworthiness, which supports credibility both within and outside the study setting.

The credibility of your findings and interpretations depend upon your careful attention to establishing trustworthiness. . . . Time is a major factor in the acquisition of trustworthy data. Time at your research site, time spent interviewing, and time building sound relationships with respondents all contribute to trustworthy data. When a large amount of time is spent with your research participants, they less readily feign behavior or feel the need to do so; moreover, they are more likely to be frank and comprehensive about what they tell you. (Glesne, 1999, p. 151)

On the other hand, prolonged engagement may actually increase reactivity as the researcher becomes more a part of the setting and begins to affect what goes on through prolonged engagement. Thus, whatever the length of inquiry or method of data collection, researchers have an obligation to examine how their presence affects what goes on and what is observed.

It is axiomatic that observers must record what they perceive to be their own reactive effects. They may treat this reactivity as bad and attempt to avoid it (which is impossible), or they may accept the fact that they will have a reactive effect and attempt to use it to advantage. . . . The reactive effect will be measured by daily field notes, perhaps by interviews in which the problem is pointedly inquired about, and also in daily observations. (Denzin, 1978b, p. 200)

Anxieties that surround an evaluation can exacerbate reactivity. The presence of an evaluator can affect how a program operates as well as its outcomes. The evaluator’s presence may, for example, create a halo effect so that staff perform in an exemplary fashion and participants are motivated to “show off.” On the other hand, the presence of the evaluator may create so much tension and anxiety that performances are below par. Some forms of program evaluation, especially “empowerment evaluation” and “intervention-oriented evaluation,” (Patton, 2008, chap. 5) turn this traditional threat to validity into an asset by designing data collection to enhance achievement of the desired program outcomes. For example, at the simplest level, the observation that “what gets measured gets done” suggests the power of data collection to affect outcomes attainment. A leadership program, for example, that includes in-depth interviewing and participant journal writing as ongoing forms of evaluation data collection may find that participating in the interviewing and writing reflectively have effects on participants’ learning and program outcomes. Likewise, a community- based AIDS awareness intervention can be enhanced by having community participants actively engaged in identifying and doing case studies of critical community incidents. In short, a variety of reactive responses are possible, some that support program processes, some that interfere, and many that have implications for interpreting findings. Thus, the evaluator has a responsibility to think about the problem, make a decision about how to handle it in the field, attempt to monitor evaluator/observer effects, and reflect on how reactivities may have affected the findings.

Evaluator effects can be overrated, particularly by evaluators. There is more than a slight touch of self- importance in some concerns about reactivity. Lillian Weber, director of the Workshop Center for Open Education, City College School of Education, New York, once set me straight on this issue, and I pass her wisdom on to my colleagues. In doing observations of open classrooms, I was concerned that my presence, particularly the way kids flocked around me as soon as I entered the classroom, was distorting the evaluation to the point where it was impossible to do good observations. Lillian laughed and suggested to me that what I was experiencing was the way those classrooms actually were. She went on to note that this was common among visitors to schools; they were always concerned that the teacher, knowing visitors were coming, whipped the kids into shape for those visitors. She suggested that under the best of circumstances a teacher might get kids to move out of habitual patterns into some model mode of behavior for as much as 10 or 15 minutes but that, habitual patterns being what they are, the kids would rapidly revert to normal behaviors and whatever artificiality might have been introduced by the presence of the visitor would likely become apparent.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 30/37

Evaluators and researchers should strive to neither overestimate nor underestimate their effects but to take seriously their responsibility to describe and study what those effects are.

Effects on the Inquirer of Being Engaged in the Inquiry A second form of reactivity arises from the possibility that the researcher or evaluator changes during the course of the inquiry. In Chapter 7, on interviewing, I offered several examples of this, including how in a study of child sexual abuse, those involved were deeply affected by what they heard. One of the ways this sometimes happens in anthropological research is when participant observers “go native” and become absorbed into the local culture. The epitome of this in a short-term observation is the legendary story of the student observers who became converted to Christianity while observing a Billy Graham evangelical crusade (Lang & Lang, 1960). Evaluators sometimes become personally involved with program participants or staff and therefore lose their sensitivity to the full range of events occurring in the setting.

Johnson (1975) and Glazer (1972) have reflected on how they and others have been changed by doing field research. The consensus of advice on how to deal with the problem of changes in observers as a result of involvement in research is similar to advice about how to deal with the reactive effects created by the presence of observers.

It is central to the method of participant observation that changes will occur in the observer; the important point, of course, is to record these changes. Field notes, introspection, and conversations with informants and colleagues provide the major means of measuring this dimension, . . . for to be insensitive to shifts in one’s own attitudes opens the way for placing naive interpretations on the complex set of events under analysis. (Denzin, 1978b, p. 200)

Inquirer-Selective Perception and Predispositions

The third concern about inquirer effects related to credibility has to do with the extent to which the predispositions or biases of the inquirer may affect data analysis and interpretations. This issue carries mixed messages because, on the one hand, rigorous data collection and analytical procedures, like triangulation, are aimed at substantiating the credibility of the findings and minimizing inquirer biases and, on the other, the interpretative and constructivist perspectives remind us that data from and about humans inevitably represent some degree of perspective rather than absolute truth. Getting close enough to the situation observed to experience it firsthand means that researchers can learn from their experiences, thereby generating personal insights; but that closeness makes their objectivity suspect. “For social scientists to refuse to treat their own behavior as data from which one can learn is really tragic” (Scriven, 1972a, p. 99). In effect, all of the procedures for validating and verifying analysis that have been presented in this chapter are aimed at reducing distortions introduced by inquirer predisposition. Still, people who use different criteria in determining evidential credibility will come at this issue from different stances and end up with different conclusions.

Consider the interviewing stance of emphatic neutrality introduced in Chapter 2 and elaborated in Chapter 7. An emphatically neutral inquirer will be perceived as caring about and interested in the people being studied but neutral about the content of what they reveal. House (1977) balances the caring, interested stance against independence and impartiality for evaluators, a stance that also applies to those working according to the standards of traditional science.

The evaluator must be seen as caring, as interested, as responsive to the relevant arguments. He must be impartial rather than simply objective. The impartiality of the evaluator must be seen as that of an actor in events, one who is responsive to the appropriate arguments but in whom the contending forces are balanced rather than non-existent. The evaluator must be seen as not having previously decided in favor of one position or the other. (pp. 45–46)

But neutrality and impartiality are not easy stances to achieve. Denzin (1989b) cites a number of scholars who have concluded, as he does, that every researcher brings preconceptions and interpretations to the problem being studied, regardless of the methods used.

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 31/37

All researchers take sides, or are partisans for one point of view or another. Value-free interpretive research is impossible. This is the case because every researcher brings preconceptions and interpretations to the problem being studied. The term hermeneutical circle or situation refers to this basic fact of research. All scholars are caught in the circle of interpretation. They can never be free of the hermeneutical situation. This means that scholars must state beforehand their prior interpretations of the phenomenon being investigated. Unless these meanings and values are clarified, their effects on subsequent interpretations remain clouded and often misunderstood. (p. 23)

Earlier I presented seven sets of criteria for judging the quality of qualitative inquiry (Exhibit 9.7, pp. 680– 681). Those varying and competing frameworks offer different perspectives on how inquirers should deal with concerns about bias. Neutrality and impartiality are expected when qualitative work is being judged by traditional scientific criteria or by evaluation standards, thus the source of House’s (1977) admonition quoted above. In contrast, constructivist analysts are expected to deal with these issues through conscious and committed reflexivity—entering the hermeneutical circle of interpretation and therein reflecting on and analyzing how their perspective interacts with the perspectives they encounter. Artistic inquirers often deal with issues of how they personally relate to their work by invoking aesthetic criteria: Judge the work on its artistic merits. Participatory and collaborative inquiries encourage the formation of meaningful and trusting relationships between researchers and those participating in the inquiry. When critical change criteria are applied in judging reactivity, the issue becomes whether, how, and to what extent the inquiry furthered the cause or enhanced the well-being of those involved and studied; neutrality is eschewed in favor of explicitly using the inquiry process to facilitate change, or at least illuminate the conditions needed for change.

Inquirer Competence Concerns about the extent to which the inquirer’s findings can be trusted—that is, trustworthiness—can be understood as one dimension of perceived methodological rigor. But ultimately, for better or worse, the trustworthiness of the data is tied directly to the trustworthiness of those who collect and analyze the data— and their demonstrated competence. Competence is demonstrated by using the verification and validation procedures necessary to establish the quality of analysis and thereby building a “track record” of quality work. As Exhibit 9.10 shows, inquirer competence includes not just systematic inquiry knowledge and skill but also interpersonal competence, reflective practice skills, situational analysis, professional practice competence, and project management. This array of competencies is being acknowledged and certified by professional evaluation associations around the world (King & Podems, 2014; Podems, 2014a). Consistent with the overall message of this chapter, especially my MQP Rumination on avoiding research rigor mortis, thinking skills also need ongoing development. An excellent resource in that regard is the Critical Evaluation Skills Toolkit (Crebert, Patrick, Cragnolini, Smith, Worsfold, & Webb, 2011).

EXHIBIT 9.10 The Multiple Dimensions of Program Evaluator Competence

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 32/37

SOURCES: Ghere, King, Stevahn, and Minnema (2006) and King, Stevahn, Ghere, and Minnema (2001).

The principle for dealing with inquirer competence is this: Don’t wait to be asked. Anticipate competence as an issue. Address the issue of competence proactively, explicitly, and multidimensionally. With quantitative methods, validity and reliability reside in tools, instruments, design parameters, and procedures. In qualitative inquiry, the competency stakes are greater because the inquirer is the instrument. Trustworthiness and authenticity are functions of systematic inquiry procedures, interpersonal (relational) dynamics in the field, and competency to engage in the challenges and deal with the ambiguities of qualitative inquiry.

Review: The Credibility of the Inquirer Because the researcher is the instrument in qualitative inquiry, the credibility of the inquirer is central to the credibility of the study. Exhibit 9.11 on the next page summarizes the issues that arise in establishing and judging the credibility of the inquirer.

EXHIBIT 9.11 The Credibility of the Inquirer: Issues and Solutions

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 33/37

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 34/37

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 35/37

MODULE

81 Generalizations, Extrapolations, Transferability, Principles, and Lessons Learned

The trouble with generalizations is that they don’t apply to particulars. —Lincoln and Guba (1985, p. 110)

Credibility and utility are linked. What can one do with qualitative findings? The results illuminate a particular situation or small number of cases. But can qualitative findings be generalized? Here, again, different qualitative frameworks based on different criteria offer different answers. The traditional scientific research criteria include generalizability. Constructivist criteria, in contrast, emphasize particularity; constructivists generally eschew, and are skeptical about, generalizability. They offer extrapolations and transferability instead. So let’s see if we can sort out these different perspectives and their implications.

Purposeful Sampling and Generalizability

Chapter 5 discussed the logic and value of purposeful sampling with small but carefully selected information- rich cases. Certain kinds of small samples and qualitative studies are designed for generalizability and broader relevance: a critical case, an index case, a causal pathway sample, a positive deviance case, and a qualitative synthesis review are examples (see Exhibit 5.8, pp. 266–272). Other sampling strategies, for example, outlier cases (exemplars of excellence or failure), a high-impact case, sensitizing concept exemplars, and principles- focused sampling, aim to yield insights about principles that might be adapted for application elsewhere. In short, the conditions for, possibility of, and relative importance attached to generalizability are determined at the design stage. To review: Purpose drives design. Design drives data collection. Data drive analysis. Purpose, design, data, and analysis, in combination, determine generalizability.

Principles of Generalizability Shadish (1995a) has made the case that certain core principles of generalization apply to both experiments and ethnographies (or qualitative methods generally). Both experiments and case studies share the problem of being highly localized. Findings from a study, experimental or naturalistic in design, can be generalized according to five principles:

1. The principle of proximal similarity: We generalize most confidently to applications where treatments, settings, populations, outcomes, and times are most similar to those in the original research. . . .

2. The principle of heterogeneity of irrelevancies: We generalize most confidently when a research finding continues to hold over variations in persons, settings, treatments, outcome measures, and times that are presumed to be conceptually irrelevant. The strategy here is identifying irrelevancies, and where possible including a diverse array of them in the research so as to demonstrate generalization over them. . . .

3. The principle of discriminant validity: We generalize most confidently when we can show that it is the target construct, and not something else, that is necessary to produce a research finding. . . .

4. The principle of empirical interpolation and extrapolation: We generalize most confidently when we can specify the range of persons, settings, treatments, outcomes, and times over which the finding holds more strongly,

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 36/37

less strongly, or not all. The strategy here is empirical exploration of the existing range of instances to discover how that range might generate variability in the finding for instances not studied. . . .

5. The principle of explanation: We generalize most confidently when we can specify completely and exactly (a) which parts of one variable (b) are related to which parts of another variable (c) through which mediating processes (d) with which salient interactions, for then we can transfer only those essential components to the new application to which we wish to generalize. The strategy here is breaking down the finding into component parts and processes so as to identify the essential ones. (pp. 424–426)

Generalizability Versus Contextual Particularity

Deep philosophical and epistemological issues are embedded in concerns about generalizing. What’s desirable or hoped for in science (generalizations across time and space) runs into real-world considerations about what’s possible. Lee J. Cronbach (1975), one of the major figures in psychometrics and research methodology in the twentieth century, devoted considerable attention to the issue of generalizations. He concluded that social phenomena are too variable and context bound to permit very significant empirical generalizations. He compared generalizations in natural sciences with what was likely to be possible in behavioral and social sciences. His conclusion was that “generalizations decay. At one time a conclusion describes the existing situation well, at a later time it accounts for rather little variance, and ultimately is valid only as history” (p. 122).

SIDEBAR

CULTURAL LIMITS ON GENERALIZABILITY

Psychological experiments have been used to study how people react to things like negotiating rewards and perceptions of whether two lines are of equal length when phony participants in the experiment say that the shorter line is really longer. The results of most such laboratory research have been interpreted as showing “evolved psychological traits common to all humans” (Watters, 2013, p. 1). However, when such experiments are repeated in other cultures, the ways in which Americans respond can be quite different from how nonliterate peoples respond. What were once thought to be tests of basic perception (how the brain works) have turned out to be culturally determined. Social scientists had assumed that lab experiments studied

the human mind stripped of culture, [that] the human brain is genetically comparable around the globe, it was agreed, so human hardwiring for much behavior, perception, and cognition should be similarly universal. No need, in that case, to look beyond the convenient population of undergraduates for test subjects. A 2008 survey of the top six psychology journals dramatically shows how common that assumption was: more than 96 percent of the subjects tested in psychological studies from 2003 to 2007 were Westerners—with nearly 70 percent from the United States alone. Put another way: 96 percent of human subjects in these studies came from countries that represent only 12 percent of the world’s population. (Watters, 2013, p. 1)

Cross-cultural research is now revealing that

the mind’s capacity to mold itself to cultural and environmental settings was far greater than had been assumed. The most interesting thing about cultures may not be in the observable things they do—the rituals, eating preferences, codes of behavior, and the like—but in the way they mold our most fundamental conscious and unconscious thinking and perception. (Watters, 2013, p. 1)

Moreover, the experiments done on American undergraduate students may be especially prone to inappropriate overgeneralizations.

It is not just our Western habits and cultural preferences that are different from the rest of the world, it appears. The very way we think about ourselves and others—and even the way we perceive reality—makes us distinct

6/2/2018 Bookshelf Online: Qualitative Research & Evaluation Methods: Integrating Theory and Practice

https://online.vitalsource.com/#/books/9781483314815/cfi/6/44!/4/2/4/2@0:0 37/37

from other humans on the planet, not to mention from the vast majority of our ancestors. Among Westerners, the data showed that Americans were often the most unusual, leading the researchers to conclude that “American participants are exceptional even within the unusual population of Westerners—outliers among outliers.”

Given the data, they concluded that social scientists could not possibly have picked a worse population [American undergraduate students)] from which to draw broad generalizations. Researchers had been doing the equivalent of studying penguins while believing that they were learning insights applicable to all birds. (Watters, 2013, p. 1)

Cronbach (1975) offers an alternative to generalizing that constitutes excellent advice for the qualitative analyst:

Instead of making generalization the ruling consideration in our research, I suggest that we reverse our priorities. An observer collecting data in a particular situation is in a position to appraise a practice or proposition in that setting, observing effects in context. In trying to describe and account for what happened, he will give attention to whatever variables were controlled, but he will give equally careful attention to uncontrolled conditions, to personal characteristics, and to events that occurred during treatment and measurement. As he goes from situation to situation, his first task is to describe and interpret the effect anew in each locale, perhaps taking into account factors unique to that locale or series of events. . . . When we give proper weight to local conditions, any generalization is a working hypothesis, not a conclusion. (pp. 124–125)

Robert Stake (1978, 1995, 2000, 2006, 2010), master of the case study, concurs with Cronbach that the first priority is to do justice to the specific case, to do a good job of “particularization” before looking for patterns across cases. He quotes William Blake on the subject: