review reading scientifice psycho

profileshanta75
Describeyourreactiontothereadingsthisweek.docx

Describe your reaction to the readings this week. More specifically, what, if anything, did you find particularly interesting or surprising? What do you think is the biggest threat to scientific integrity in psychology, and what are some ways that it could be remedied?

I have copied and passed the articles at the bottom

Article:  False Positives: Fraud and Misconduct Are Threatening Scientific Research opens in new window This is an article from the popular press (i.e., not peer-reviewed) that discusses several high-profile cases of academic fraud in psychology. The article also discusses many of the factors that may lead to fraud and misconduct in psychology research.

Library Article:  Measuring the Prevalence of Questionable Research Practices With Incentives for Truth Telling opens in new window The article describes the results of a survey that was given to psychologists to assess the extent to which they engaged in what are known as questionable research practices.

https://www.theguardian.com/science/2012/sep/13/scientific-research-fraud-bad-practice

False positives: fraud and misconduct are threatening scientific research

High-profile cases and modern technology are putting scientific deceit under the microscope

· Alok Jha

·

· Alok Jha , science correspondent

·

· The Guardian , Thursday 13 September 2012 18.12 BST

Diederik Stapel

The Dutch psychologist Diederik Stapel was found to have published fabricated data in 30 peer-reviewed papers. Photograph: Hollandse Hoogte/Boxem

Dirk Smeesters had spent several years of his career as a social psychologist at Erasmus University in Rotterdam studying how consumers behaved in different situations. Did colour have an effect on what they bought? How did death-related stories in the media affect how people picked products? And was it better to use supermodels in cosmetics adverts than average-looking women?

The questions are certainly intriguing, but unfortunately for anyone wanting truthful answers, some of  Smeesters' work turned out to be fraudulent. The psychologist, who admitted "massaging" the data in some of his papers, resigned from his position in June after being investigated by his university, which had been tipped off by Uri Simonsohn from the University of Pennsylvania in Philadelphia. Simonsohn carried out an independent analysis of the data and was suspicious of how perfect many of Smeesters' results seemed when, statistically speaking, there should have been more variation in his measurements.

The case, which led to two scientific papers being retracted, came on the heels of an even bigger fraud, uncovered last year, perpetrated by the Dutch psychologist Diederik Stapel. He was found to have fabricated data for years and published it in at least 30 peer-reviewed papers, including  a report in the journal Science  about how untidy environments may encourage discrimination.

The cases have sent shockwaves through a discipline that was already facing serious questions about plagiarism.

"In many respects, psychology is at a crossroads – the decisions we take now will determine whether or not it remains a serious, credible, scientific discipline along with the harder sciences," says Chris Chambers, a psychologist at Cardiff University.

"We have to be open about the problems that exist in psychology and understand that, though they're not unique to psychology, that doesn't mean we shouldn't be addressing them. If we do that, we can end up leading the other sciences rather than following them."

Cases of scientific misconduct tend to hit the headlines precisely because scientists are supposed to occupy a moral high ground when it comes to the search for truth about nature. The scientific method developed as a way to weed out human bias. But scientists, like anyone else, can be prone to bias in their bid for a place in the history books.

Increasing competition for shrinking government budgets for  research  and the disproportionately large rewards for publishing in the best journals have exacerbated the temptation to fudge results or ignore inconvenient data.

Massaged results can send other researchers down the wrong track, wasting time and money trying to replicate them. Worse, in medicine, it can delay the development of life-saving treatments or prolong the use of therapies that are ineffective or dangerous. Malpractice comes to light rarely, perhaps because scientific fraud is often easy to perpetrate but hard to uncover.

The field of  psychology has come under particular scrutiny  because many results in the scientific literature defy replication by other researchers. Critics say it is too easy to publish psychology papers which rely on sample sizes that are too small, for example, or to publish only those results that support a favoured hypothesis. Outright fraud is almost certainly just a small part of that problem, but high-profile examples have exposed a greyer area of bad or lazy scientific practice that many had preferred to brush under the carpet.

Many scientists, aided by software and statistical techniques to catch cheats, are now speaking up, calling on colleagues to put their houses in order.

Those who document misconduct in scientific research talk of a spectrum of bad practices. At the sharp end are plagiarism, fabrication and falsification of research. At the other end are questionable practices such as adding an author's name to a paper when they have not contributed to the work, sloppiness in methods or not disclosing conflicts of interest.

"Outright fraud is somewhat impossible to estimate, because if you're really good at it you wouldn't be detectable," said Simonsohn, a social psychologist. "It's like asking how much of our money is fake money – we only catch the really bad fakers, the good fakers we never catch."

If things go wrong, the responsibility to investigate and punish misconduct rests with the scientists' employers, the academic institution. But these organisations face something of a conflict of interest. "Some of the big institutions … were really in denial and wanted to say that it didn't happen under their roof," says Liz Wager of the Committee on Publication Ethics (Cope). "They're gradually realising that it's better to admit that it could happen and tell us what you're doing about it, rather than to say, 'It could never happen.'"

There are indications that bad practice – particularly at the less serious end of the scale – is rife. In 2009, Daniele Fanelli of the University of Edinburgh carried out a meta-analysis that pooled the results of 21 surveys of researchers who were asked whether they or their colleagues had fabricated or falsified research.

Publishing his results in the journal PLoS One , he found that an average of 1.97% of scientists admitted to having "fabricated, falsified or modified data or results at least once – a serious form of misconduct by any standard – and up to 33.7% admitted other questionable research practices. In surveys asking about the behaviour of colleagues, admission rates were 14.12% for falsification, and up to 72% for other questionable research practices."

2006 analysis  of the images published in the Journal of Cell Biology found that 1% of accepted papers have at least one image that has been manipulated in a way that affects the interpretation of the data - though the authors made no conclusions about intent.

Rise in retractions

According to  a report in the journal Nature , published retractions in scientific journals have increased around 1,200% over the past decade, even though the number of published papers had gone up by only 44%. Around half of these retractions are suspected cases of misconduct.

Wager says these numbers make it difficult for a large research-intensive university, which might employ thousands of researchers, to maintain the line that misconduct is vanishingly rare.

New tools, such as text-matching software, have also increased the detection rates of fraud and plagiarism. Journals routinely use these to check papers as they are submitted or undergoing peer review. "Just the fact that the software is out there and there are people who can look at stuff, that has really alerted the world to the fact that plagiarism and redundant publication are probably way more common than we realised," says Wager. "That probably explains, to a big extent, this increase we've seen in retractions."

Ferric Fang, a professor at the University of Washington School of Medicine and editor in chief of the journal  Infection and Immunity , thinks increased scrutiny is not the only factor and that the rate of retractions is indicative of some deeper problem.

He was alerted to concerns about the work of a Japanese scientist who had published in his journal. A reviewer for another journal noticed that Naoki Mori of the University of the Ryukyus in Japan had duplicated images in some of his papers and had given them different labels, as if they represented different measurements. An investigation revealed evidence of widespread data manipulation and this led Fang to retract six of Mori's papers from his journal. Other journals followed suit.

Self-correction

The refrain from many scientists is that the scientific method is meant to be self-correcting. Bad results, corrupt data or fraud will get found out – either when they cannot be replicated or when they are proved incorrect in subsequent studies – and public retractions are a sign of strength.

That works up to a point, says Fang. "It ended up that there were 31 papers from the [Mori] laboratory  that were retracted , many of those papers had been in the literature for five-10 years," he says. "I realised that 'scientific literature is self-correcting' is a little bit simplistic. These papers had been read many times, downloaded, cited and reviewed by peers and it was just by the chance observation by a very attentive reviewer that opened this whole case of serious misconduct."

Extraordinary claims that change the paradigm for a field will elicit lots of attention and people will look at the results very carefully. But cases such as Dr Mori's – where work is flawed and falsified but the results themselves are not particularly surprising or sensational and may even corroborated by others who perform their experiments legitimately – the misconduct is difficult to detect. "It's not that the results are wrong, it's that the data are false," says Fang.

And, often, research studies are very difficult to replicate. "If someone says they did a 15-year clinical study with 9,000 subjects and they publish their results, you may have to take their word for it because you're not going to be able to run out and recruit 9,000 patients of your own and do a 15-year study just to try to corroborate something that somebody else has done," says Fang. "A number of cases recently have come to light only because the investigators didn't have institutional review board approval for their studies. Upon digging deeper, the institutions questioned whether any of the studies were done at all. This kind of misconduct is very difficult to detect otherwise."

Selective publishing

In psychology research, there is a particular problem with researchers who selectively publish some of their experiments to guarantee a positive result. "Let's say you have this theory that, when you play Mozart, people want to pay more for musical instruments," says Simonsohn. "So you do a study and you play Mozart (or not) and you ask people, 'How much would you pay for a piano or flute and five instruments?'"

If it turned out that only the price of a single type of instrument, violins, say, went up after people had listened to Mozart, it would be possible to publish a research paper that omitted the fact that the researchers had ever asked about any other instruments. This would not allow the reader to make a proper assessment of the strength of the effect that Mozart may (or may not) have on how much a person would pay for musical instruments.

Fanelli has examined this positive result bias . He looked at 4,600 studies across all disciplines between 1990 and 2007, and counted the number of papers that, after declaring an intent to test a particular hypothesis, reported a positive support for it. The overall frequency of positive supports had grown by more than 22% over this time period. In  a separate study , Fanelli found that "the odds of reporting a positive result were around five times higher among papers in the disciplines of psychology and psychiatry and economics and business compared with space science".

Culture of neophilia

This issue is exacerbated in psychological research by the "file-drawer" problem, a situation when scientists who try to replicate and confirm previous studies find it difficult to get their research published. Scientific journals want to highlight novel, often surprising, findings. Negative results are unattractive to journal editors and lie in the bottom of researchers' filing cabinets, destined never to see the light of day.

"We have a culture which values novelty above all else, neophilia really, and that creates a strong publication bias," says Chambers. "To get into a good journal, you have to be publishing something novel, it helps if it's counter-intuitive and it also has to be a positive finding. You put those things together and you create a dangerous problem for the field."

When Daryl Bem, a psychologist at Cornell University in New York,  published sensational findings in 2011  that seemed to show evidence for psychic effects in people, many scientists were unsurprisingly sceptical. But when psychologists later  tried to publish their (failed) attempts  to replicate Bem's work, they found journals refused to give them space. After repeated attempts elsewhere, a team of psychologists led by Chris French at Goldsmith's, University of London,  eventually placed their negative results  in the journal PLoS One this year.

There is no suggestion of misconduct in Bem's research but the lack of an avenue in which to publish failed attempts at replication suggests self-correction can be compromised and people such as Smeesters and Stapel can remain undetected for a long time.

In some cases, misconduct (or fraud) has grave implications. In 2006, Anil Potti and colleagues at Duke University reported in the New England Journal of Medicine that they had developed a way to track the progression of a patient's lung cancer with a device, called an expression array, that could monitor the activity of thousands of different genes. In a subsequent report in Nature Medicine, the same scientists wrote about a way to use their expression array to work out which drugs would work best for individual patients with lung, breast or ovarian cancer, depending on their patterns of gene activity. Within months of that publication, the biostatisticians Keith Baggerly and Kevin Coombes of the MD Anderson Cancer Centre in Houston had their doubts, and began uncovering major flaws in the work.

"It looked so promising that they actually started to do trials of cancer patients, they chose the chemotherapy depending on this test," says Wager. "The test has turned out to be completely invalid, so people were getting the wrong therapy, because the paper was not retracted quickly enough."

Blowing the whistle

Despite Baggerly and Coombes raising the alarm several times with the institutions involved, it was not until 2010 that Potti resigned from Duke University and several of the papers referring to his work on the expression array were retracted."Usually there is no official mechanism for a whistleblower to take if they suspect fraud," says Chambers. "You often hear of cases where junior members of a department, such as PhD students, will be the ones that are closest to the coalface and will be the ones to identify suspicious cases. But what kind of support do they have? ... That's a big issue that needs to be addressed."

In July this year, a group of the UK's main research funders and university groups published a  Concordat to Support Research Integrity . "I don't think anyone would want to see a command-control direct regulation approach here," says Christopher Hale, deputy director of policy at Universities UK. "The concordat ... outlines a framework and then identifies how people fit within that and what actions they will take forward to strengthen it." The concordat requires institutions to have a process in place for dealing with misconduct, which includes appointing a senior person at the institution who can provide the necessary leadership and oversight during investigations.

Michael Farthing, vice-chair of the UK Research Integrity Office and vice-chancellor of the University of Sussex, has been a long-time campaigner on getting institutions and funders to take research misconduct seriously. In a recent  article for Times Higher Education , Farthing said he supported the concordat but that it would not be enough. He stopped short of suggesting a statutory regulator for research but wrote: "Government and research leaders should take action to support and encourage excellence in research integrity, not sit on their hands until – as has happened in other countries – a scandal drives them towards legislation."

Statements of principle are one thing – every university and research council probably already has one applauding honourable research and deploring fraud – the key is the steps institutions take in understanding and de-incentivising misconduct.

The economics of science

The pressure to commit misconduct is complex. Arturo Casadevall of the Albert Einstein College of Medicine in New York and editor in chief of the journal mBio, places a large part of the blame on the economics of science. "What is happening in recent years is that the rewards have become too high, for example, for publishing in certain journals. Just like we see the problem in sports that, if you compete and you get a reward, it translates into everything from money and endorsements and things like that. People begin to take risks because the rewards are disproportionate."

As a PhD student in the 1980s, Casadevall says he published research in a few different journals depending on what his research was about. "Within 10 years, all you heard was, 'Where is the paper going to be published?' not 'What's in it?'. Scientists have got into this idea that where you publish determines the value of the work and that's crazy. What's important is what's in the paper."

Casadevall and Fang are aware that their spotlight on misconduct has the potential to show up scientists in a disproportionately bad light – as yet another public institution that cannot be trusted beyond its own self-interest. But they say staying quiet about the issue is not an option.

"Science has the potential to address some of the most important problems in society and for that to happen, scientists have to be trusted by society and they have to be able to trust each others' work," Fang says. "If we are seen as just another special interest group that are doing whatever it takes to advance our careers and that the work is not necessarily reliable, it's tremendously damaging for all of society because we need to be able to rely on science."

For Simonsohn, the biggest issue with outright fraud is not that the bad scientist gets caught but the corrupting effect the work can have on the scientific literature. To reduce the potential negative effects dramatically, Simonsohn suggests requiring scientists to post their data online. "That's very minimal cost and it has many benefits beyond reduction of fraud. It allows other people to learn things from your data which you were not able to learn about, it allows calibration of other models, it allows people to, three years later, reanalyse your data with new techniques."

Ivan Oransky, editor of the  Retraction Watch  blog that collects examples of retracted papers, argues: "The reason the public stops trusting institutions is when [its members] say things like, 'There's nothing to see here, let us handle it,' and then they find out about something bad that happened that nobody handled. That's when mistrust builds.The big challenges that face humanity, says Casadevall, are scientific ones – climate change, a new pandemic, the fact that most of our calories are coming from a very few plants, which are susceptible to new pests. "These are the big problems and humanity's defence against them is science. We need to make the enterprise work better."

Malpractice and misconduct

The South Korean scientist  Hwang Woo-suk, rose to international acclaim in 2004 when he announced, in the journal Science, that he had extracted stem cells from cloned human embryos. The following year, Hwang published results showing he had made stem cell lines from the skin of patients – a technique that could help create personalised cures for people with degenerative diseases. By 2006, however, Hwang's career was in tatters when it emerged that he had fabricated material for his research papers. Seoul National University sacked him and, after an investigation in 2009, he was  convicted of embezzling research funds .

Around the same time, a Norwegian researcher,  Jon Sudbø, admitted to fabricating and falsifying data. Over many years of malpractice, he perpetrated one of the biggest scientific frauds ever carried out by a single researcher – the fabrication of an entire 900-patient study, which was published in the Lancet in 2005.

Marc Hauser, a psychologist at Harvard University whose research interests included the evolution of morality and cognition in non-human primates, resigned in August 2011 after a three-year investigation by his institution found he was responsible for eight counts of scientific misconduct. The alarm was raised by some of his students, who disagreed with Hauser's interpretations of experiments that involved the, somewhat subjective, procedure of working out a monkey's thoughts based on its response to some sight or sound.

Hauser last week  admitted to making "mistakes"  that led to the findings of research misconduct. "I let important details get away from my control, and as head of the lab, I take responsibility for all errors made within the lab, whether or not I was directly involved," says Hauser in a statement sent to Nature. The doubts over Hauser's work affect a whole field of scientific work that uses the same research technique.

• This article was amended on 14 September 2012. The original referred to Liz Wager of the Committee on Public Ethics rather than Publication Ethics. This has been corrected. This article was further amended on 19 September 2012. The original stated that a 2006 analysis of the images published in the Journal of Cell Biology found that about 1% had been deliberately falsified. This has been corrected.

We have switched off comments on this old version of the site. To comment on crosswords, please  switch over to the new version to comment Read more...

· License/buy our content  

· Privacy policy  

· Terms & conditions  

· Advertising guide  

· Accessibility  

· A-Z index  

· Inside the Guardian blog  

· About us  

· Work for us  

· Join our dating site today

· © 2016 Guardian News and Media Limited or its affiliated companies. All rights reserved.

bstract

Cases of clear scientific misconduct have received significant media attention recently, but less flagrantly questionable research practices may be more prevalent and, ultimately, more damaging to the academic enterprise. Using an anonymous elicitation format supplemented by incentives for honest reporting, we surveyed over 2,000 psychologists about their involvement in questionable research practices. The impact of truth-telling incentives on self-admissions of questionable research practices was positive, and this impact was greater for practices that respondents judged to be less defensible. Combining three different estimation methods, we found that the percentage of respondents who have engaged in questionable practices was surprisingly high. This finding suggests that some questionable practices may constitute the prevailing research norm.

Although cases of overt scientific misconduct have received significant media attention recently ( Altman, 2006Deer, 2011; Steneck, 2002,  2006), exploitation of the gray area of acceptable practice is certainly much more prevalent, and may be more damaging to the academic enterprise in the long run, than outright fraud. Questionable research practices (QRPs), such as excluding data points on the basis of post hoc criteria, can spuriously increase the likelihood of finding evidence in support of a hypothesis. Just how dramatic these effects can be was demonstrated by  Simmons, Nelson, and Simonsohn (2011) in a series of experiments and simulations that showed how greatly QRPs increase the likelihood of finding support for a false hypothesis. QRPs are the steroids of scientific competition, artificially enhancing performance and producing a kind of arms race in which researchers who strictly play by the rules are at a competitive disadvantage. QRPs, by nature of the very fact that they are often questionable as opposed to blatantly improper, also offer considerable latitude for rationalization and self-deception.

Concerns over QRPs have been mounting ( Crocker, 2011; Lacetera & Zirulia, 2011;  Marshall, 2000; Sovacool, 2008; Sterba, 2006; Wicherts, 2011), and several studies—many of which have focused on medical research—have assessed their prevalence ( Gardner, Lidz, & Hartwig, 2005; Geggie, 2001;  Henry et al., 2005List, Bailey, Euzent, & Martin, 2001Martinson, Anderson, & de Vries, 2005; Swazey, Anderson, & Louis, 1993). In the study reported here, we measured the percentage of psychologists who have engaged in QRPs.

As with any unethical or socially stigmatized behavior, self-reported survey data are likely to underrepresent true prevalence. Respondents have little incentive, apart from good will, to provide honest answers ( Fanelli, 2009). The goal of the present study was to obtain realistic estimates of QRPs with a new survey methodology that incorporates explicit response-contingent incentives for truth telling and supplements self-reports with impersonal judgments about the prevalence of practices and about respondents’ honesty. These impersonal judgments made it possible to elicit alternative estimates, from which we inferred the upper and lower boundaries of the actual prevalence of QRPs. Across QRPs, even raw self-admission rates were surprisingly high, and for certain practices, the inferred actual estimates approached 100%, which suggests that these practices may constitute the de facto scientific norm.

Method

In a study with a two-condition, between-subjects design, we e-mailed an electronic survey to 5,964 academic psychologists at major U.S. universities (for details on the survey and the sample, see Procedure and Table S1, respectively, in the Supplemental Material available online). Participants anonymously indicated whether they had personally engaged in each of 10 QRPs ( self-admission rateTable 1), and if they had, whether they thought their actions had been defensible. The order in which the QRPs were presented was randomized between subjects. There were 2,155 respondents to the survey, which was a response rate of 36%. Of respondents who began the survey, 719 (33.4%) did not complete it (see Supplementary Results and Fig. S1 in the Supplemental Material); however, because the QRPs were presented in random order, data from all respondents—even those who did not finish the survey—were included in the analysis.

Table 1. Results of the Main Study: Mean Self-Admission Rates, Comparison of Self-Admission Rates Across Groups, and Mean Defensibility Ratings

 

Self-admission rate (%)

Odds ratio (BTS/control)

Two-tailed  p (likelihood ratio test)

Defensibility rating (across groups)

Item

Control group

BTS group

1. In a paper, failing to report all of a study’s dependent measures

63.4

66.5

1.14

.23

1.84 (0.39)

2. Deciding whether to collect more data after looking to see whether the results were significant

55.9

58.0

1.08

.46

1.79 (0.44)

3. In a paper, failing to report all of a study’s conditions

27.7

27.4

0.98

.90

1.77 (0.49)

4. Stopping collecting data earlier than planned because one found the result that one had been looking for

15.6

22.5

1.57

.00

1.76 (0.48)

5. In a paper, “rounding off” a  p value (e.g., reporting that a  p value of .054 is less than .05)

22.0

23.3

1.07

.58

1.68 (0.57)

6. In a paper, selectively reporting studies that “worked”

45.8

50.0

1.18

.13

1.66 (0.53)

7. Deciding whether to exclude data after looking at the impact of doing so on the results

38.2

43.4

1.23

.06

1.61 (0.59)

8. In a paper, reporting an unexpected finding as having been predicted from the start

27.0

35.0

1.45

.00

1.50 (0.60)

9. In a paper, claiming that results are unaffected by demographic variables (e.g., gender) when one is actually unsure (or knows that they do)

3.0

4.5

1.52

.16

1.32 (0.60)

10. Falsifying data

0.6

1.7

2.75

.07

0.16 (0.38)

Note: Items are listed in decreasing order of rated defensibility. Respondents who admitted to having engaged in a given behavior were asked to rate whether they thought it was defensible to have done so (0 =  no, 1 =  possibly, and 2 =  yes). Standard deviations are given in parentheses. BTS = Bayesian truth serum. Applying the Bonferroni correction for multiple comparisons, we adjusted the critical alpha level downward to .005 (i.e., .05/10 comparisons).

OPEN IN VIEWER

In addition to providing self-admission rates, respondents also provided two impersonal estimates related to each QRP: (a) the percentage of other psychologists who had engaged in each behavior ( prevalence estimate), and (b) among those psychologists who had, the percentage that would admit to having done so ( admission estimate). Therefore, each respondent was asked to provide three pieces of information for each QRP. Respondents who indicated that they had engaged in a QRP were also asked to rate whether they thought it was defensible to have done so (0 =  no, 1 =  possibly, and 2 =  yes). If they wished, they could also elaborate on why they thought it was (or was not) defensible.

After providing this information for each QRP, respondents were also asked to rate their degree of doubt about the integrity of the research done by researchers at other institutions, other researchers at their own institution, graduate students, their collaborators, and themselves (1 =  never, 2 =  once or twice, 3 =  occasionally, 4 =  often).

The two versions of the survey differed in the incentives they offered to respondents. In the Bayesian-truth-serum (BTS) condition, a scoring algorithm developed by one of the authors (Prelec, 2004) was used to provide incentives for truth telling. This algorithm uses respondents’ answers about their own behavior and their estimates of the sample distribution of answers as inputs in a truth-rewarding scoring formula. Because the survey was anonymous, compensation could not be directly linked to individual scores. Instead, respondents were told that we would make a donation to a charity of their choice, selected from five options, and that the size of this donation would depend on the truthfulness of their responses, as determined by the BTS scoring system. By inducing a (correct) belief that dishonesty would reduce donations, we hoped to amplify the moral stakes riding on each answer (for details on the donations, see Supplementary Results in the Supplemental Material). Respondents were not given the details of the scoring system but were told that it was based on an algorithm published in  Science and were given a link to the article. There was no deception: Respondents’ BTS scores determined our contributions to the five charities. Respondents in the control condition were simply told that a charitable donation would be made on behalf of each respondent. (For details on the effect of the size of the incentive on response rates, see Participation Incentive Survey in the Supplemental Material.)

The three types of answers to the survey questions—self-admission, prevalence estimate, admission estimate—allowed us to estimate the actual prevalence of each QRP in different ways. The credibility of each estimate hinged on the credibility of one of the three answers in the survey: First, if respondents answered the personal question honestly, then self-admission rates would reveal the actual prevalence of the QRPs in this sample. Second, if average prevalence estimates were accurate, then they would also allow us to directly estimate the actual prevalence of the QRPs. Third, if average admission estimates were accurate, then actual prevalence could be estimated using the ratios of admission rates to admission estimates. This would correspond to a case in which respondents did not know the actual prevalence of a practice but did have a good sense of how likely it is that a colleague would admit to it in a survey. The three estimates should converge if the self-admission rate equaled the prevalence estimate multiplied by the admission estimate. To the extent that this equality is violated, there would be differences between prevalence rates measured by the different methods.

Results

Raw self-admission rates, prevalence estimates, prevalence estimates derived from the admission estimates (i.e., self-admission rate/admission estimate), and geometric means of these three percentages are shown in  Figure 1. For details on our approach to analyzing the data, see Data Analysis in the Supplemental Material.

Fig. 1. Results of the Bayesian-truth-serum condition in the main study. For each of the 10 items, the graph shows the self-admission rate, prevalence estimate, prevalence estimate derived from the admission estimate (i.e., self-admission rate/admission estimate), and geometric mean of these three percentages (numbers above the bars). See  Table 1 for the complete text of the items. OPEN IN VIEWER

Truth-telling incentives

A priori, truth-telling incentives (as provided in the BTS condition) should affect responses in proportion to the baseline (i.e., control condition) level of false denials. These baseline levels are unknown, but one can hypothesize that they should be minimal for impersonal estimates of prevalence and admission, and greatest for personal admissions of unethical practices broadly judged as unacceptable, which represent “red-card” violations.

As hypothesized, prevalence estimates (see Table S2 in the Supplemental Material) and admission estimates (see Table S3 in the Supplemental Material) were comparable in the two conditions, but self-admission rates for some items ( Table 1), especially those that were “more questionable,” were higher in the BTS condition than in the control condition. ( Table 1 also presents the  p values of the likelihood ratio test of the difference in admission rates between conditions.)

We assessed the effect of the BTS manipulation by examining the odds ratio of self-admission rates in the BTS condition to self-admission rates in the control condition. The odds ratio was high for one practice (falsifying data), moderate for three practices (premature stopping of data collection, falsely reporting a finding as expected, and falsely claiming that results are unaffected by certain variables), and negligible for the remainder of the practices ( Table 1). The acceptability of a practice can be inferred from the self-admission rate in the control condition (baseline) or assessed directly by judgments of defensibility. The nonparametric correlation of BTS impact, as measured by odds ratio, with the baseline self-admission rate was –.62 ( p < .06; parametric correlation = −.65,  p < .05); the correlation of odds ratio with defensibility rating was –.68 ( p < .03; parametric correlation = −.94,  p < .001). These correlations were more modest when Item 10 (“Falsifying data”) was excluded (odds ratio with baseline self-admission rate: nonparametric correlation = −.48,  p < .20; parametric correlation = −.59,  p < .10; odds ratio with defensibility rating: nonparametric correlation = −.57,  p < .12; parametric correlation = −.59,  p < .10).

Prevalence estimates

Figure 1 displays mean prevalence estimates for the three types of responses in the BTS condition (the admission estimates were capped at 100%; they exceeded 100% by a small margin for a few items). The figure also shows the geometric means of all three responses; these means, in effect, give equal credence to the three types of answers. The raw admission rates are almost certainly too low given the likelihood that respondents did not admit to all QRPs that they actually engaged in. Therefore, the geometric means are probably conservative judgments of true prevalence.

One would infer from the geometric means of the three variables that nearly 1 in 10 research psychologists has introduced false data into the scientific record (Items 5 and 10) and that the majority of research psychologists have engaged in practices such as selective reporting of studies (Item 6), not reporting all dependent measures (Item 1), collecting more data after determining whether the results were significant (Item 2), reporting unexpected findings as having been predicted (Item 8), and excluding data post hoc (Item 7).

These estimates are somewhat higher than estimates reported in previous research. For example, a meta-analysis of surveys—none of which provided incentives for truthful responding—found that, among scientists from a variety of disciplines, 9.5% of respondents admitted to having engaged in QRPs other than data falsification; the upper-boundary estimate was 33.7% ( Fanelli, 2009). In the present study, the mean self-admission rate in the BTS condition (excluding the data-falsification item for comparability with  Fanelli, 2009) was 36.6%—higher than both of the meta-analysis estimates. Moreover, among participants in the BTS condition who completed the survey, 94.0% admitted to having engaged in at least one QRP (compared with 91.4% in the control condition). The self-admission rate in our control condition (33.0%) mirrored the upper-boundary estimate obtained in Fanelli’s meta-analysis (33.7%).

Response to a given item on our survey was predictive of responses to the other items: The survey items approximated a Guttman scale, meaning that an admission to a relatively rare behavior (e.g., falsifying data) usually implied that the respondent had also engaged in more common behaviors. Among completed response sets, the coefficient of reproducibility—the average proportion of a person’s responses that can be reproduced by knowing the number of items to which he or she responded affirmatively—was .80 (high values indicate close agreement; items are considered to form a Guttman scale if reproducibility is .90 or higher;  Guttman, 1974). This finding suggests that researchers’ engagement in or avoidance of specific QRPs is not completely idiosyncratic. It indicates that there is a rough consensus among researchers about the relative unethicality of the behaviors, but large variation in where researchers draw the line when it comes to their own behavior.

Perceived defensibility

Respondents had an opportunity to state whether they thought their actions were defensible. Consistent with the notion that latitude for rationalization is positively associated with engagement in QRPs, our findings showed that respondents who admitted to a QRP tended to think that their actions were defensible. The overall mean defensibility rating of practices that respondents acknowledged having engaged in was 1.70 ( SD = 0.53)—between possibly defensible and defensible. Mean judged defensibility for each item is shown in  Table 1. Defensibility ratings did not generally differ according to the respondents’ discipline or the type of research they conducted (see Table S4 in the Supplemental Material).

Doubts about research integrity

A large percentage of respondents indicated that they had doubts about research integrity on at least one occasion ( Fig. 2). The degree of doubt differed by target; for example, respondents were more wary of research generated by researchers at other institutions than of research conducted by their collaborators. Although heterogeneous referent-group sizes make these differences difficult to interpret (the number of researchers at other institutions is presumably larger than one’s own set of collaborators), it is noteworthy that approximately 35% of respondents indicated that they had doubts about the integrity of their own research on at least one occasion.

Fig. 2. Results of the main study: distribution of responses to a question asking about doubts concerning the integrity of the research conducted by various categories of researchers. OPEN IN VIEWER

Frequency of engagement

Although the prevalence estimates obtained in the BTS condition are somewhat higher than previous estimates, they do not enable us to distinguish between the researcher who routinely engages in a given behavior and the researcher who has only engaged in that behavior once. To the extent that self-admission rates are driven by the former type, our results are more worrisome. We conducted a smaller-scale survey, in which we tested for differences in admission rates as a function of the response scale.

We asked 133 attendees of an annual conference of behavioral researchers whether they had engaged in each of 25 different QRPs (many of which we also inquired about in the main study). Using a 2 × 2 between-subjects design, we manipulated the wording of the questions and the response scale. The questions were either phrased as a generic action (“Falsifying data”) or in the first person (“I have falsified data”), and participants indicated whether they had engaged in the behaviors using either a dichotomous response scale (yes/no, as in the main study) or a frequency response scale ( never, once or twice, occasionally, frequently).

Because the overall self-admission rates to the individual items were generally similar to those obtained in the main study, we do not report them here. Respondents made fewer affirmative admissions on the dichotomous response scale ( M = 3.77 out of 25,  SD = 2.27) than on the frequency response scale ( M = 6.02 out of 25,  SD = 3.70),  F(1, 129) = 17.0,  p < .0005). This result suggests that in the dichotomous-scale condition, some nontrivial fraction of respondents who engaged in a QRP only a small number of times reported that they had never engaged in it. This suggests that the prevalence rates obtained in the main study are conservative. There was no effect of the wording manipulation.

We explored the response-scale effect further by comparing the distribution of responses between the two response-scale conditions across all 25 items and collapsing across the wording manipulation ( Fig. 3). Among the affirmative responses in the frequency-response-scale condition (i.e., responses of  once or twice, occasionally, or  frequently), 64% (i.e., .153/(.151 + .062 + .023)) of the affirmative responses fell into the  once or twice category, a nontrivial percentage fell into  occasionally (26%), and 10% fell into  frequently. This result suggests that the prevalence estimates from the BTS study represent a combination of single-instance and habitual engagement in the behaviors.

Fig. 3. Results of the follow-up study: distribution of responses among participants who were asked whether they had engaged in 25 questionable research practices. Participants answered using either (a) a frequency response scale or (b) a dichotomous response scale. OPEN IN VIEWER

Subgroup differences

Table 2 presents self-admission rates as a function of disciplines within psychology and the primary methodology used in research. Relatively high rates of QRPs were self-reported among the cognitive, neuroscience, and social disciplines, and among researchers using behavioral, experimental, and laboratory methodologies (for details, see Data Analysis in the Supplemental Material). Clinical psychologists reported relatively low rates of QRPs.

Table 2. Mean Self-Admission Rate, Applicability Rating, and Defensibility Rating by Category of Research

Category of research

Self-admission rate (%)

Applicability rating

Defensibility rating

Discipline

 Clinical

27 *

2.59 (0.94)

0.56 (0.28)

 Cognitive

37 ***

2.75 * (0.93)

0.64 (0.23)

 Developmental

31

2.77 ** (0.89)

0.66 (0.27)

 Forensic

28

3.02 * (1.12)

0.52 (0.29)

 Health

30

2.56 (0.94)

0.69 (0.31)

 Industrial organizational

31

2.80 (0.63)

0.73 (0.30)

 Neuroscience

35 **

2.71 (0.92)

0.61 (0.21)

 Personality

32

2.65 * (0.92)

0.66 (0.36)

 Social

40 ***

2.89 *** (0.85)

0.73 ** (0.31)

Research type

 Clinical

30

2.61 (0.99)

0.56 (0.27)

 Behavioral

34 *

2.77 ** (0.88)

0.63 (0.28)

 Laboratory

36 ***

2.87 *** (0.86)

0.66 (0.29)

 Field

31

2.76 ** (0.88)

0.63 (0.28)

 Experimental

36 ***

2.83 * (0.87)

0.66 * (0.29)

 Modeling

33

2.74 (0.89)

0.62 (0.26)

Note: Self-admission rates are from the main study and are collapsed across all 10 items; applicability and defensibility ratings are from the follow-up study. Applicability was rated on a 4-point scale (1 =  never applicable, 2 =  sometimes applicable, 3 =  often applicable, 4 =  always applicable). Defensibility was rated on a 3-point scale (0 =  no, 1 =  possibly, 2 =  yes). For self-admission rates, random-effects logistic regression was used to identify significant effects; for applicability and defensibility ratings, random-effects ordered probit regressions were used to identify significant effects.

*

p < .05. ** p < .01. *** p < .0005.

OPEN IN VIEWER

These subgroup differences could reflect the particular relevance of our QRPs to these disciplines and methodologies, or they could reflect differences in perceived defensibility of the behaviors. To explore these possible explanations, we sent a brief follow-up survey to 1,440 of the participants in the main study, which asked them to rate two aspects of the same 10 QRPs. First, they were asked to rate the extent to which each practice applies to their research methodology (i.e., how frequently, if at all, they encountered the opportunity to engage in the practice). The possible responses were  never applicable, sometimes applicable, often applicable, and  always applicable. Second, they were asked whether it is generally defensible to engage in each practice. The possible responses were  indefensible, possibly defensible, and  defensible. Unlike in the main study, in which respondents were asked to provide a defensibility rating only if they had admitted to having engaged in a given practice, all respondents in the follow-up survey were asked to provide these ratings. We counterbalanced the order in which respondents rated the two dimensions. There were 504 respondents, for a response rate of 35%. Of respondents who began the survey, 65 (12.9%) did not complete it; as in the main study, data from all respondents—even those who did not finish the survey—were included in the analysis because the QRPs were presented in randomized order.

Table 2 presents the results from the follow-up survey. The subgroup differences in applicability ratings and defensibility ratings were partially consistent with the differences in self-reported prevalence: Most notably, mean applicability and defensibility ratings were elevated among social psychologists—a subgroup with relatively high self-admission rates. Similarly, the items were particularly applicable to (but not judged to be more defensible by) researchers who conduct behavioral, experimental, and laboratory research.

To test for the relative importance of applicability and defensibility ratings in explaining subfield differences, we conducted an analysis of variance on mean self-admission rates across QRPs and disciplines. Both type of QRP ( p < .001, η p2 = .87) and subfield ( p < .05, η p2 = .21) were highly significant predictors of self-admission rates, and their significance and effect size were largely unchanged after controlling for applicability and defensibility ratings, even though both of the latter variables were highly significant independent predictors of mean self-admission rates. Similarly, methodology was also a highly significant predictor of self-admission rates ( p < .05, η p2 = .27), and its significance and effect size were largely unchanged after controlling for applicability and defensibility ratings (even though the latter were highly significant predictors of self-admission rates).

The defensibility ratings obtained in the main study stand in contrast with those obtained in the follow-up survey: Respondents considered these behaviors to be defensible when they engaged in them (as was shown in the main study) but considered them indefensible overall (as was shown in the follow-up study).

Discussion

Concerns over scientific misconduct have led previous researchers to estimate the prevalence of QRPs that are broadly applicable to scientists ( Martinson et al., 2005). In light of recent concerns over scientific integrity within psychology, we designed this study to provide accurate estimates of the prevalence of QRPs that are specifically applicable to research psychologists. In addition to being one of the first studies to specifically target research psychologists, it is also the first to test the effectiveness of an incentive-compatible elicitation format that measures prevalence rates in three different ways.

All three prevalence measures point to the same conclusion: A surprisingly high percentage of psychologists admit to having engaged in QRPs. The effect of the BTS manipulation on self-admission rates was positive, and greater for practices that respondents judge to be less defensible. Beyond revealing the prevalence of QRPs, this study is also, to our knowledge, the first to illustrate that an incentive-compatible information-elicitation method can lead to higher, and likely more valid, prevalence estimates of sensitive behaviors. This method could easily be used to estimate the prevalence of other sensitive behaviors, such as illegal or sexual activities. For potentially even greater benefit, BTS-based truth-telling incentives could be combined with audio computer-assisted self-interviewing—a technology that has been found to increase self-reporting of sensitive behaviors ( Turner et al., 1998).

There are two primary components to the BTS procedure—both a request and an incentive to tell the truth—and we were unable to isolate their independent effects on disclosure. However, both components rewarded respondents for telling the truth, not for simply responding “yes” regardless of whether they had engaged in the behaviors. Therefore, both components were designed to increase the validity of responses. Future research could test the relative contribution of the various BTS components in eliciting truthful responses.

This research was based on the premise that higher prevalence estimates are more valid—an assumption that pervades a large body of research designed to assess the prevalence of sensitive behaviors ( Bradburn & Sudman, 1979de Jong, Pieters, & Fox, 2010; Lensvelt-Mulders, Hox, van der Heijden, & Maas, 2005;  Tourangeau & Yan, 2007Warner, 1965). This assumption is generally accepted, provided that the behaviors in question are sensitive or socially undesirable. The rationale is that respondents are unlikely to be tempted to admit to shameful behaviors in which they have not engaged; instead, they are prone to denying involvement in behaviors in which they actually have engaged ( Fanelli, 2009). We think this assumption is also defensible in the present study given its subject matter.

As noted in the introduction, there is a large gray area of acceptable practices. Although falsifying data (Item 10 in our study) is never justified, the same cannot be said for all of the items on our survey; for example, failing to report all of a study’s dependent measures (Item 1) could be appropriate if two measures of the same construct show the same significant pattern of results but cannot be easily combined into one measure. Therefore, not all self-admissions represent scientific felonies, or even misdemeanors; some respondents provided perfectly defensible reasons for engaging in the behaviors. Yet other respondents provided justifications that, although self-categorized as defensible, were contentious (e.g., dropping dependent measures inconsistent with the hypothesis because doing so enabled a more coherent story to be told and thus increased the likelihood of publication). It is worth noting, however, that in the follow-up survey—in which participants rated the behaviors regardless of personal engagement—the defensibility ratings were low. This suggests that the general sentiment is that these behaviors are unjustifiable.

We assume that the vast majority of researchers are sincerely motivated to conduct sound scientific research. Furthermore, most of the respondents in our study believed in the integrity of their own research and judged practices they had engaged in to be acceptable. However, given publication pressures and professional ambitions, the inherent ambiguity of the defensibility of “questionable” research practices, and the well-documented ubiquity of motivated reasoning (Kunda, 1990), researchers may not be in the best position to judge the defensibility of their own behavior. This could in part explain why the most egregious practices in our survey (e.g., falsifying data) appear to be less common than the relatively less questionable ones (e.g., failing to report all of a study’s conditions). It is easier to generate a post hoc explanation to justify removing nuisance data points than it is to justify outright data falsification, even though both practices produce similar consequences.

Given the findings of our study, it comes as no surprise that many researchers have expressed concerns over failures to replicate published results ( Bower & Mayer, 1985Crabbe, Wahlsten, & Dudek, 1999Doyen, Klein, Pichon, & Cleeremans, 2012, Enserink, 1999; Galak, LeBoeuf, Nelson, & Simmons, 2012;  Ioannidis, 2005a2005bPalmer, 2000Steele, Bass, & Crook, 1999). In an article on the problem of nonreplicability,  Lehrer (2010) discussed possible explanations for the “decline effect”—the tendency for effect sizes to decrease with subsequent attempts at replication. He concluded that conventional accounts of this effect (regression to the mean, publication bias) may be incomplete. In a subsequent and insightful commentary,  Schooler (2011) suggested that unpublished data may help to account for the decline effect. By documenting the surprisingly large percentage of researchers who have engaged in QRPs—including selective omission of observations, experimental conditions, and studies from the scientific record—the present research provides empirical support for Schooler’s claim.  Simmons and his colleagues (2011) went further by showing how easily QRPs can yield invalid findings and by proposing reforms in the process of reporting research and accepting scientific manuscripts for publication.

QRPs can waste researchers’ time and stall scientific progress, as researchers fruitlessly pursue extensions of effects that are not real and hence cannot be replicated. More generally, the prevalence of QRPs raises questions about the credibility of research findings and threatens research integrity by producing unrealistically elegant results that may be difficult to match without engaging in such practices oneself. This can lead to a “race to the bottom,” with questionable research begetting even more questionable research. If reforms would effectively reduce the prevalence of QRPs, they not only would bolster scientific integrity but also could reduce the pressure on researchers to produce unrealistically elegant results.

Acknowledgments

We thank Evan Robinson for implementing the e-mail procedure that tracked participation while ensuring respondents’ anonymity. We also thank Anne-Sophie Charest and Bill Simpson for statistical consulting and members of the Center for Behavioral Decision Research for their input on initial drafts of the survey items.

Competing Interests

The authors declared that they had no conflicts of interest with respect to their authorship or the publication of this article.

References

Altman L. K. (2006, May 2). For science gatekeepers, a credibility gap. The New York Times. Retrieved from  http://www.nytimes.com/2006/05/02/health/02docs.html?pagewanted=all

GO TO REFERENCE

Google Scholar

Bower G. H., Mayer J. D. (1985). Failure to replicate mood-dependent retrieval. Bulletin of the Psychonomic Society, 23, 39–42.

GO TO REFERENCE

Crossref

Google Scholar

Bradburn N., Sudman S. (1979). Improving interview method and questionnaire design: Response effects to threatening questions in survey research. San Francisco, CA: Jossey-Bass.

GO TO REFERENCE

Google Scholar

Crabbe J. C., Wahlsten D., Dudek B. C. (1999). Genetics of mouse behavior: Interactions with laboratory environment. Science, 284, 1670–1672.

image2.jpeg

image1.jpeg