Answer five questions in requirement . Each question requires more than 300 words, a total of 1,500 words

profileccc777
MGMT7250_TOPIC5_CritAppScience_w.pdf

Evidence-based Management

MGMT 7250

TOPIC 5:The critical appraisal of scientific evidence

Agenda for the day

• Acquiring scientific evidence • Causal inference: a reminder

• Testing causal theories and assumptions • Hypotheses • Effect direction and size • Statistical significance

• P-values • Confidence intervals • Limits of these

• Methodological appropriateness • Methodological quality • Practical relevance

How to acquire scientific evidence

Peer-reviewed journals

• Academy of Management Journal. • Administrative Science Quarterly. • Strategic Management Journal. • Journal of Management. • Journal of Applied Psychology. • Psychological Bulletin. • Journal of Management Studies. • Journal of Organizational Behavior. • Journal of Human Relations. • Journal of Vocational Behavior

ABI/INFORM Business Source Elite PsycINFO ERIC Web of Knowledge Google Scholar

Online research databases

Source: CEBMa

Online databases

• There are a number of online e-databases that provide access to full- text academic journal articles including the in journals above. • Organisations and their research teams would benefit immensely to subscribe to any of the following e-databases. • Where organisations do not have access to e-databases, they can contact academics to assist with access. • Managers undertaking executive training at universities will have access to such databases. Once you know how to use these databases, searching for scientific findings can be very expedient.

Undertaking REAs & CATs

• Expedient forms of systematic reviews • A Critically Appraised Topic (CAT) provides a quick and succinct assessment of what is known (and not known) in the scientific literature about an intervention or practical issue by using a systematic methodology to search and critically appraise primary studies • Rapid Evidence Assessments (REAs) are more extensive in scope than CATs, and need to be undertaken by groups. • Organisations can either commission REAs which are an underutilised input to decision and policy- making.

Using PICOC for search terms

• Start with the two most important PICOC terms – the intervention and outcome. • Coming up with a couple of alternative terms for the outcome and intervention may also be useful. With the example just used ‘performance’ may be used instead of ‘productivity’ for example. • Then when you have some search terms it is important to pre-test them in one database. • If the yield is too big, it is important to refine your search to increase specificity. If your search yields very few papers, consider dropping a search term.

Search strategies

1. Shotgun • Use your search terms to generate a list of articles

2. Snowball • Snowballing backward: Snowballing backwards is where you begin with a more recent publication and discover which publications were used by the author by looking at the bibliography of books or articles. • Snowballing forward: Snowballing forward is where you begin with a less recent publications and discover how often that book or article has been cited by other authors by looking at Web of Knowledge or Google Scholar.

Other tips • include the terms ‘studies’, ‘systematic review’ and ‘meta-analysis’ • tick a box if available ‘scholarly journals including peer-reviewed’. • use quotation marks (e.g. "search") to search for specific key-words. • use advanced search (access from down arrow in search bar) to exclude terms, or limit the search to a specific date range or specific publication.

• not all articles are created equal - one way to identify important articles is to look at the 'cited by...' number. The higher the number, in general, the better the article. This method is not fool proof as sometimes an article may be cited highly for the wrong reasons. Click on the 'cited by' link to view articles that cite a particular article.

• use the 'related articles' link below a specific search result to find similar articles, and the 'cite' link to show a formatted reference to the chosen article.

• use the links on the far right of the screen to access full text versions of articles. You will note not all articles have a link to a free ee full text version. You can access the original article by clicking on the article title - this will take you to the publisher's website, where you can view the abstract and buy an electronic version of the article.

Critical appraisal of scientific evidence

Critical appraisal of scientific evidence

Two criteria to determine the trustworthiness of scientific evidence: ØMethodological appropriateness

ØMethods appropriate for type of question or claim (ie., about effects) ØInternal validity – the extent to which legitimate causal inferences can be drawn.

ØMethodological quality ØMeasurement reliability and validity ØSampling and external validity

Critical appraisal 1: Methodological appropriateness

What is methodological appropriateness?

• Methodological appropriateness: certain research designs are better (more trustworthy) than others for answering particular types of research questions • Methodological appropriateness involves examining whether the research design suits the study purpose, research aims and questions. • Research designs that follow the 3 Criteria for Causality are methodologically appropriate for Cause-and-effect (Effect) questions

Causal inference: a reminder

Cause and effect: Relationships between variables • Context: How do differences in context influence the impact?

M or A Org.

performance

Size diff., cult. Diff., distance

+ or - ?

+ or - compared to status quo?

Cause and effect: Types of variables in scientific research • Variables: things that vary! Science is mainly about finding relationships between variables • Dependant/Outcome Variable (DV): of primary interest to the researcher • Independent/Explanatory Variables (IV): Factors that explain the dependant variable • Mediator: variable that mediates the relationship between IV and DV • Moderator: Variable that conditions (strengthens or weakens) the relationship between IV and DV.

IV Med DV

Mod

Testing causal theories and assumptions

Hypotheses

IV DV

Mod

+

-

The null hypothesis • Remember that science is about ‘falsification’. We’re not trying to prove effects. We’re not trying to show

for certain that an effect exists • According to the null hypothesis testing method, in testing a hypothesized effect, we’re actually checking

to see if the absence of the effect can be rejected • The null hypothesis is usually an alternative hypothesis—that the IV has no (zero) effect on the DV • If we can show that the null hypothesis is unlikely to be true, then we can show support for our main

hypothesized effect

Findings: Effect direction and effect sizes

Effect Size Small Medium Large Independent means: d, ∆, g ≤ 0.20 0.50 ≥ 0.80 Correlation: r, р ≤ 0.10 0.30 ≥ 0.50 Correlation: r2 ≤ 0.01 0.09 ≥ 0.25 ANOVA: f ≤ 0.10 0.25 ≥ 0.40 ANOVA: ƞ2 ≤ 0.10 0.06 ≥ 0.14 Simple regression: β ≤ 0.02 0.30 ≥ 0.50 Multiple regression: β ≤ 0.20 0.50 ≥ 0.80 Multiple regression: R2 ≤ 0.02 0.13 ≥ 0.26 Multiple regression: f2 ≤ 0.02 0.15 ≥ 0.35

• Findings: Estimated effects based on analysis of data

• Effect direction: Does the IV have a positive or negative effect on the DV?

• Effect size: Managers need to know whether they should bank on a finding or not. This is where effect sizes are important.

Effects and effect sizes

IV DV

Mod

+ (r2=0.6)

- (r2=0.4)

Statistical inference: Going from the sample to the population • How do we know that what we have found in the sample applies to the population • Sample representativeness • Statistical significance: How do we know that the effect will be true for the population too? • P-values • Confidence intervals

Statistical significance: p-values

• P-values: tell us how likely a finding/effect is likely to be true and not found purely by chance, ie. Is the effect statistically significant? • The probability that, while the null hypothesis is true, you would find the effect estimate that you have found • A P-value of .05 is a commonly used threshold or level of significance (this mean that the chance that the outcomes of our study are due to chance are 5% or below) and anything above the .05 level is statistically non- significant. • But this is arbitrary • Also just because there is a significant effect does not mean it is important/relevant

Statistical significance: confidence intervals

• Confidence intervals: tell us how precise the findings are • They tell us the confidence with which we can say that the true value of the population—based on the effect we’ve found in the sample, and the size of the sample—lies between a certain range of values • The smaller the confidence interval, the more likely that the population effect will be close to the sample effect • Would you make different decisions based on whether the population effect is closer to the lower boundary or upper boundary?

• Null hypothesis testing: The null hypothesis is rejected if zero does not fall in the confidence interval

Statistical significance: confidence intervals

3

concerns job satisfaction but ‘large’ when the outcome concerns fatal medical errors. When assessing impact, it is therefore important to relate the effect size directly to the outcome measured.

Effect sizes would be typically provided in the Results section of a research paper and/or a separate table. Don’t let yourself be taken in by the huge amount of numbers and symbols – just scan through the text and see if you can identify one of the effect sizes listed in the table above. In addition, if you have two studies that use different effect sizes, you can use the table to make a comparison. For example, if the first study finds a difference of d =.20 between the job satisfaction of two groups, and the second study has a difference of h2 = .01, you can conclude that the differences found in both studies are small. The same counts for effect sizes within a study. Take for example the table below.

In this table, you can see that the overall effect of goal setting on performance is d=.56, which can be considered a medium effect. When you look under ‘Goal difficulty’, however, you can see that easy goals have a small effect (d=.23), whereas difficult goals have a large effect (d=.80). Both effects are statistically significant, but only the impact of difficult goals may be practically relevant!

Many quantitative studies include a so-called correlation matrix, an overview of the correlation coefficients between the variables studied. An example is provided below. As you can see, in this study a significant correlation was found between sales training A, sales training B, sales training C and sales performance, but only the effect of sales training B is of practical relevance.

Table 3: Correlation matrix for study variables

No Variable 1 2 3 4 5 6 7 8 9 10

1 gender 1

2 age -.15* 1

3 education -.32* -.02 1

4 firm size .03* .20* .21 1

5 firm age .07 .02 .19 .34 1

6 sales training A -.02 .12 .02 .02 .02 1

7 sales training B .12 .09* .04* .03 .05 .03 1

8 sales trainng C .09 .48 .01 .22 -.06 .05* .16 1

9 experience .07 .39 .24 .27 -,05 .11* .05 .32 1

10 sales performance .06 .01 .21 .19 -.07 -.07* .62* .11* .13 1

*Significant at 0.05 level

(Courtesy CEBMa)

Cause and effect: Relationships between variables • What did CISCO find?

M or A Org.

performance

Size difference

+

Cultural difference, geographic distance

- -

Internal validity

• The extent to which legitimate causal inferences can be made – function of research design. • Do changes in the IV explain variation in the DV with everything else staying the same? • Can changes in the outcome be attributed to other causes?

• Needs to rule out “endogeneity”: reverse/mutual causality (Y causes X), and spuriousness – X and Y may be strongly correlated but they are not causally related, they may be positively correlated because of a third overlooked variable, Z (confounding variable) • Other threats to internal validity: history effects, maturation, testing effects, selection bias • Internal validity is established if the 3 Steps of Causality are followed

Critical thinking: Establishing causality The 3 criteria 1. Is there a strong correlation? • Use quantitative/statistical methods to do this

2. Does the cause precedes effect? We need to establish whether there has been before and after measurement or whether a baseline measure exists.

• But beware of “post hoc, ergo propter hoc” fallacy

3. Can we rule out alternative explanations for the correlation? • Random assignment • Control variables • Other methods: e.g. propensity score matching

Random assignment, control variables

• Since we can’t always compare similar organizations or individuals, we could use: • Random assignment: Individuals or organizations are randomly assigned to ‘treatment’ or ’control’ group. Each individual or org has an equal chance of being in either group. • This helps ensure that on average, the two groups will be similar. • Random assignment helps wash out the effects of unmeasured individual differences explaining variation in an outcome.

• Control variables: When you can’t conduct experiments, you can account for other differences among individuals and orgs in your study by using control variables, which measure the impact of other forces (than your main intervention and comparison variables) on the outcome

Randomised controlled study with before and after measurement

9

possible distorting factor is equally spread over both groups. Thus, any differences between the groups measured at the end of the study can be more confidently attributed to the variable that is expected to have an effect (the assumed cause).

3.5. The gold standard Based on the elements described above, we can now design the ‘best’ – i.e. the most appropriate – study to answer a cause-and-effect question with the lowest chance of confounders. Let’s go back to the example of the researcher who wanted to examine the effect of a stress-reduction program using on-site chair massage therapy. The first criterion of causality is demonstrating that the cause and effect correlate using valid and reliable measurements. We have already determined that the most valid and reliable way to measure stress reduction would be to measure the cortisol levels in the saliva of employees. The second criterion of causality, temporality, states that we must demonstrate that the cause (chair massage therapy) preceded the effect (stress reduction). We therefore measure the employees’ cortisol level both before and after the chair massage therapy. The third criterion concerns ruling out plausible alternative explanations for the effect found, so we randomly assign two groups of employees to a control and an intervention group. We first take a representative sample of the employees and then we randomly assign them to the two groups by flipping a coin (or, better use the random generator in Excel). Then we need to determine what the control group should do to create a valid and reliable benchmark: Continue working, take a break, or ‘fake’ massage therapy? After some discussion, we decide that continue working is an unfair comparison and a fake massage therapy is hard to realize, so the control group will take a break during which they do something to relax. This means that our research design would look like the depiction below. This design is known as a randomized controlled before-after study or randomized controlled trial (RCT), the gold standard to answer cause-and-effect questions.

Note! Random assignment is not the same as random sampling! Random sampling refers to selecting subjects in such a way that they represent a larger population, whereas random assignment concerns assigning subjects to a control group and experimental group in such a way that they are similar at the start of the study. Put differently, random selection ensures high representativeness, whereas random assignment ensures high internal validity.

This means that if a study uses the term ‘random’ or ‘randomized’, you should always determine whether this concerns random selection (as in a survey) or random assignment (as in a controlled study).

(Courtesy CEBMa)

Research designs (1) • Systematic review - Addresses a specific question, utilises explicit and transparent methods to perform a thorough literature search and critical appraisal of individual studies, and draws conclusions about what we currently know and do not know about a given question or topic (Briner & Denyer, 2012). • Meta-analysis - Uses statistical analysis techniques in a systematic review to pool the results of the individual studies numerically in order to achieve a more accurate estimate of the effect.

• Randomised control study (true experiment) - Groups compared with each other are selected entirely at random, for example by drawing lots. This means each participant (or other unit such as a team, department or company) has an equal chance of being in the intervention or control group. In this way the influence of distorting factors is spread over both groups so the groups are as comparable as possible with each other with the exception of the intervention.

• Non-randomised control study (quasi-experiment; field experiment) - Two or more groups are compared with each other, usually comprising one group in which an intervention is carried out (experimental group) and one group where no or an alternative intervention is conducted (control group). Allocation to the groups is not randomised.

Research designs (2)

• Longitudinal study: Subjects studied over many time periods • Before and after study (pre-test, post- test) - Observations are made before and after an intervention.

• Interrupted time-series - Observations are made at multiple time points before and after an intervention, to detect whether the intervention had a significantly greater effect than an underlying secular trend.

• Cohort or panel study - A group of people followed over time including before and after an intervention

• Comparative or case-control study - Compares people or cases in terms of a specific outcome.

• Cross-sectional study (case study, survey) - Independent variables and dependent variables are measured at the same point in time. • Qualitative study (interviews) - Explores and tries to understand people's beliefs, experiences, attitudes, behaviour and interactions. It generates non-numerical data. The best-known qualitative research- methods include in-depth interviews, focus groups, documentary analysis and participant observation.

Cross-sectional vs longitudinal

• In cross-sectional research, the independent variable and dependent variables are measured at the same point in time which means that reverse causation is possible. • Temporal antecedence is an important characteristic of longitudinal research. • Longitudinal research includes studies that involve repeated observations (measurements) of the same variable(s) over a certain period of time, such as cohort studies or interrupted times series. • It is possible with longitudinal research to explain changes in the dependent variable or outcome over time.

Research design and the 3 criteria for causality

Establishing correlation

Cause before effect Random assignment

Control variables, etc

Randomised controlled trials (experiments)

Y Y Y N

Field/Natural experiment

Y Y N Y

Longitudinal study Y Y N Y

Cross-sectional study

Y N N/Y Y

Case study N N N N

Methodological appropriateness for Effect Qs

Source: Center for Evidence-Based Management

38Source: CEBMa

Rating the methodological appropriateness of a study examining Cause and Effects

• Work in 3s • A large accounting firm is considering training their senior executives in mindfulness (3 day workshop) to improve their performance and leadership. If the firm goes ahead with this, how would the company evaluate the efficacy this training program?

• Which study type would you use? • Are there any potential threats to internal validity?

Activity – Internal validity and making causal claims

39

Critical appraisal 2: Methodological quality

What is methodological quality?(1)

• The trustworthiness of a study is also affected by its methodological quality. This refers to how the study was conducted. • This involves taking a closer look at other aspects of a study’s research design including: i) the study’s measurement reliability and validity; as well as ii) the sampling design and external validity. • It is very important to note the relationship between methodological quality and appropriateness. • Even if a study is methodologically appropriate (i.e. has good internal validity), it may still have quality issues regarding its measures (e.g. poor reliability and invalid measures) and sampling design (e.g. a sample unrepresentative of the target population). This will compromise the extent to which causal inferences can be made.

• In terms of rating a study’s methodological quality, we can start with the study’s methodological appropriateness which we just learned how to rate. • The Center for Evidence-based Management further suggests that, based on its number of weaknesses, the level of trustworthiness may then be downgraded by one or more levels. To determine the final level of trustworthiness you can make use of the following rules of thumb:

Ø1 weakness = no downgrade (we accept that nothing is perfect)

Ø2 weaknesses = downgrade 1 level Ø3 weaknesses = downgrade 2 levels Ø4 weaknesses = downgrade 2 levels

What is methodological quality?(2)

Measurement

Reliability

• Examining reliability involves questioning how accurately constructs or variables in a study have been measured and how good the indicators are.

• A measure is reliable when it is accurate and different attempts at measuring something converge on the same result.

• A measure of some construct (e.g., job performance) is rendered unreliable by noise and bias.

• As we learned with classical test theory, all measurements are subject to error, particularly when measuring human performance, personality, emotion, attitudes, cognitive processes, aptitudes, behaviours and capabilities.

• Consistent with classical test theory all measurements are subject to error, particularly when measuring human performance, personality, emotion, attitudes, cognitive processes, aptitudes, behaviours and capabilities. For example, managers need to evaluate employee performance periodically, what are potential sources of inaccuracy that can occur when measuring an employee’s ‘personal initiative’?

Validity

• Validity is whether enough of the right (relevant) constructs have been measured or whether critical constructs or variables have been omitted. When assessing measurement validity there are four types of validity that can be examined, including:

ØConstruct validity: The ability of a measure to reliably measure and truthfully represent a unique concept.

ØContent validity: The degree that a measure covers the breadth of a concept.

ØCriteria-related validity: The ability of a measure to correlate with established criteria. Criteria represent measures of outcomes decisions are designed to produce, and thus the validity of the decisions.

ØPredictive validity: The correlation between measures obtained before making decisions, and criterion scores obtained after making decisions.

The faces of validity

Face

Content

Construct

Criterion- related

Validity of measurement

Validity of decisions

How do we validly measure:

Leadership? Performance? Innovation?

Has something been overlooked? Did we forget to measure

something?

Sampling & External validity

• Sampling is the process of selecting the right individuals, objects, or events for a study. • Impacts study’s generalisability which affects external validity - the degree to which the results of a study can be generalised across individuals and contexts. • There are two types of sampling designs. These are: ØProbability: Known, non-zero probability for every element to be selected ØNon-probability: Probability of selecting any particular member is unknown • These two sampling designs have different implications for generalizability. That is, probability sampling is used when causal interpretability and representativeness are important

Types of external validity

• There are two dimensions of external validity, including: Øpopulation validity: sample to target population inferences; and Øecological validity: the extent to which the findings generalise across settings/contexts

Practical significance & relevance

Practical significance

• It is important to appraise just how much findings are practically important. • Effect sizes rather than p-values (Topic 4) tell us something about practical significance. • An effect may be statistically significant but not practically significant because the effect size is small. • When it comes to management interventions, it is worth investing resources where effect sizes are moderate to large.

How do you identify possible interventions or solutions in science? • Let’s think of an example where a meta-analysis by Griffeth, Hom, & Gaetner (2000) finds that organisational commitment has a moderate effect size on turnover (ρ = -.23). In the first instance the finding suggests the organisation would benefit from increasing commitment. • But what how does this translate into an intervention? 1. Many research papers have a ‘practical implications’ section. This can be helpful for identifying interventions or strategies to address a problem. 2. It would also be worth acquiring additional evidence to address a question regarding solutions: Ø “what is known in the scientific literature about drivers of commitment?”

Practical relevance

• In evidence-based practice in management it is important to think about how findings are relevant to a local context. • When it is not clear whether the effects generalise to the local context (i.e. have ecological validity) it is important to consider whether an effect or cause may or may not occur in the local context:

ØIs your organisation/division/population so different from those in the study that the study's results cannot apply?

ØHow relevant is the study to what you are seeking to understand or decide? ØHow could the research potentially benefit or harm your organisation? ØIs the intervention feasible in your setting? ØWhat are your executive’s (or client’s) concerns, preferences and expectations for the outcome you are trying to prevent and the research you are offering?