Clinical Application of Epidemiology
9
Treatment
Treatments should be given “not because they ought to work, but because they do work.”
—L.H. Opie 1980
Key Words
Hypotheses
Treatment
Intervention
Comparative effectiveness
Experimental studies
Clinical trials
Randomized controlled trials
Equipoise
Inclusion criteria
Exclusion criteria
Comorbidity
Large simple trials
Practical trials
Pragmatic trials
Hawthorne effect
Placebo
Placebo effect
Random allocation
Randomization
Baseline characteristics
Stratified randomization
Compliance
Adherence
Run-in period
Crossover
Blinding
Masking
Allocation concealment
Single-blind
Double-blind
Open label
Composite outcomes
Health-related quality of life
Health status
Efficacy trials
Effectiveness trials
Intention-to-treat analysis
Explanatory analyses
Per-protocol
Superiority trials
Noninferiority trials
Noninferiority margin
Cluster randomized trials
Crossover trials
Stepped wedge cluster trials
N of 1 trials
Confounding by indication
Phase I trials
Phase II trials
Phase III trials
Postmarketing surveillance
After the nature of a patient's illness has been established and its expected course predicted, the next question is, what can be done about it? Is there a treatment that improves the outcome of disease? This chapter describes the evidence used to decide whether a well-intentioned treatment is actually effective.
IDEAS AND EVIDENCE
Ideas
Ideas about what might be a useful treatment arise from virtually any activity within medicine. These ideas are called hypotheses to the extent that they are assertions about the natural world that are made for the purposes of empiric testing.
Some therapeutic hypotheses are suggested by the mechanisms of disease at the molecular level. Drugs for antibiotic-resistant bacteria are developed through knowledge of the mechanism of resistance and hormone analogs are variations on the structure of native hormones. Other hypotheses about treatments have come from astute observations by clinicians, shared with their colleagues in case reports. Others are discovered by accident: The drug minoxidil, which was developed for hypertension, was found to improve male pattern baldness; and tamoxifen, developed for contraception, was found to prevent breast cancer in high-risk women. Traditional medicines, some of which are supported by centuries of experience,
P.143
may be effective. Aspirin, atropine, and digitalis are examples of naturally occurring substances that have become established as orthodox medicines after rigorous testing. Still other ideas come from trial and error. Some anticancer drugs have been found by methodically screening huge numbers of substances for activity in laboratory models. Ideas about treatment, but more often prevention, have also come from epidemiologic studies of populations. The Framingham Study, a cohort study of risk factors for cardiovascular diseases, was the basis for clinical trials of lowering blood pressure and serum cholesterol.
|
View Figure
|
Figure 9.1. Ideas and evidence.
|
Testing Ideas
Some treatment effects are so prompt and powerful that their value is self-evident even without formal testing. Clinicians do not have reservations about the effectiveness of antibiotics for bacterial meningitis, or diuretics for edema. Clinical experience is sufficient.
In contrast, many diseases, including most chronic diseases, involve treatments that are considerably less dramatic. The effects are smaller, especially when an effective treatment is tested against another effective treatment. Also outcomes take longer to develop. It is then necessary to put ideas about treatments to a formal test, through clinical research, because a variety of circumstances, such as coincidence, biased comparisons, spontaneous changes in the course of disease, or wishful thinking, can obscure the true relationship between treatment and outcomes.
When knowledge of the pathogenesis of disease, based on laboratory models or physiologic studies in humans, has become extensive, it is tempting to predict effects in humans on this basis alone. However, relying solely on current understanding of mechanisms without testing ideas using strong clinical research on intact humans can lead to unpleasant surprises.
Example
Control of elevated blood sugar has been a keystone in the care of patients with diabetes mellitus, in part to prevent cardiovascular complications. Hyperglycemia is the most obvious metabolic abnormality in patients with diabetes. Cardiovascular disease is common in patients with diabetes, and observational studies have shown an association between elevated blood sugar and cardiovascular events. To study the effects of tight control of blood sugar on cardiovascular events, the ACCORD Trial randomized 10,251 patients with type 2 diabetes mellitus and other risk factors for cardiovascular disease to either intensive therapy or usual care (1). Glucose control was substantially better in the intensive therapy group but surprisingly, after 3.7 years in the trial, patients assigned to intensive therapy had 21% more deaths. They also had more hypoglycemic episodes and more weight gain. Increased mortality with “tight” control, which was contrary to conventional thinking about diabetes (but consistent with other trial results), has prompted less aggressive goals for blood glucose control and more aggressive treatment of other risk factors for cardiovascular disease, such as blood pressure, smoking, and dyslipidemia.
This study illustrates how treatments that make good sense, based on what is known about the disease at the time, may be found to be ineffective when put to a rigorous test in humans. Knowledge of pathogenesis, worked out in laboratory models, may be disappointing in human studies because the laboratory studies are in highly simplified settings. They usually exclude or control for many real-world influences on disease such as variation in genetic endowment,
P.144
the physical and social environment, and individual behaviors and preferences.
Clinical experience and tradition also need to be put to a test. For example, bed rest has been advocated for a large number of medical conditions. Usually, there is a rationale for it. For example, it has been thought that the headache following lumbar puncture might result from a leak of cerebrospinal fluid through the needle track causing stretching of the meninges. However, a review of 39 trials of bed rest for 15 different conditions found that outcome did not improve for any condition. Outcomes were worse with bed rest in 17 trials, including not only lumbar puncture, but also acute low back pain, labor, hypertension during pregnancy, acute myocardial infarction, and acute infectious hepatitis (2). In another example, oxygen supplementation is common in hospitalized patients, but too much oxygen is probably harmful based on trial data. A review of 25 trials of liberal versus conservative oxygen administration in acutely ill adults found an estimated 21% increase in mortality at levels above 96% (3).
Of course, it is not always the case that ideas are debunked. The main point is that promising treatments have to be tested by clinical research rather than accepted into the care of patients on the basis of reasoning alone.
STUDIES OF TREATMENT EFFECTS
Treatment is any intervention that is intended to improve the course of disease after it is established. Treatment is a special case of interventions in general that might be applied at any point in the natural history of disease, from disease prevention to palliative care at the end of life. Although usually thought of as medications, surgery, or radiotherapy, health care interventions can take any form, including relaxation therapy, laser surgery, or changes in the organization and financing of health care. Regardless of the nature of a well-intentioned intervention, the principles by which it is judged superior to other alternatives are the same.
Comparative effectiveness is a popular name for a not-so-new concept, the head-to-head comparison of two or more interventions (e.g., drugs, devices, tests, surgery, or monitoring), all of which are believed to be effective and are current options for care. Comparison is not just for effectiveness, but also for all clinically important end results of the interventions—both beneficial and harmful. Results can help clinicians and patients understand all of the consequences of choosing one or another course of action when both have been considered reasonable alternatives.
Observational and Experimental Studies of Treatment Effects
Two general methods are used to establish the effects of interventions: observational and experimental studies. The two differ in their scientific strength and feasibility.
In observational studies of interventions, investigators simply observe what happens to patients who for various reasons do or do not get exposed to an intervention (see Chapters 6 , 7 , 8 ). Observational studies of treatment are a special case of studies of prognosis in general, in which the prognostic factor of interest is a therapeutic intervention. What has been said about cohort studies applies to observational studies of treatment as well. The main advantage of these studies is feasibility. The main drawback is the possibility that there are systematic differences in treatment groups, other than the treatment itself, that can lead to misleading conclusions about the effects of treatment.
Experimental studies are a special kind of cohort study in which the conditions of study—selection of treatment groups, nature of interventions, management during follow-up, and measurement of outcomes—are specified by the investigator for the purpose of making unbiased comparisons. These studies are generally referred to as clinical trials. Clinical trials are more highly controlled and managed than cohort studies. The investigators are conducting an experiment, analogous to those done in the laboratory. They have taken it upon themselves (with their patients' permission) to isolate for study the unique contribution of one factor by holding constant, as much as possible, all other determinants of the outcome.
Randomized controlled trials, in which treatment is randomly allocated, are the standard of excellence for scientific studies of the effects of treatment. They are described in detail below, followed by descriptions of alternative ways of studying the effectiveness of interventions.
RANDOMIZED CONTROLLED TRIALS
The structure of a randomized controlled trial is shown in Figure 9.2. All elements are the same as for a cohort study except that treatment is assigned by randomization rather than by physician and patient choice. The “exposures” are treatments, and the “outcomes” are any possible end result of treatment (such as the 5 Ds described in Table 1.2 ).
|
View Figure
|
Figure 9.2. The structure of a randomized controlled trial.
|
The patients to be studied are first selected from a larger number of patients with the condition of interest. Using randomization, the patients are then divided into two (or more) groups of comparable prognosis. One group, called the experimental group, is exposed to an intervention that is believed to be better than current alternatives. The other group, called a control (or comparison) group, is treated the same in all ways except that its members are not exposed to the experimental intervention. Patients in the control group may receive a placebo, usual care, or the current best available treatment. The course of disease is then recorded in both groups, and differences in outcome are attributed to the intervention.
The main reason for structuring clinical trials in this way is to avoid confounding when comparing the respective effects of two or more kinds of treatments. The validity of clinical trials depends on how well they have created equal distribution of all determinants of prognosis, other than the one being tested, in treated and control patients.
Individual elements of clinical trials are described in detail in the following text.
Ethics
Under what circumstances is it ethical to assign treatment at random, rather than as decided by the patient and physician? The general principle, called equipoise, is that randomization is ethical when there is no compelling reason to believe that either of the randomly allocated treatments is better than the other. Usually it is believed that the experimental intervention might be better than the control but that has not been conclusively established by strong research. The primary outcome must be benefit; treatments cannot be randomly allocated to discover whether one is more harmful than the other. Of course, as with any human research, patients must fully understand the consequences of participating in the study, know that they can withdraw at any time without compromising their health care, and freely give their consent to participate. In addition, the trial must be stopped whenever there is convincing evidence of effectiveness, harm, or futility in continuing.
Sampling
Clinical trials typically require patients to meet rigorous inclusion and exclusion criteria. These are intended to increase the homogeneity of patients in the study, to strengthen internal validity, and to make it easier to distinguish the “signal” (treatment effect) from the “noise” (bias and chance).
Among the usual inclusion criteria is that patients really do have the condition being studied. To be on the safe side, study patients must meet strict diagnostic criteria. Patients with unusual, mild, or equivocal manifestations of disease may be left out in the process, restricting generalizability.
Of the many possible exclusion criteria, several account for most of the losses:
1. Patients with comorbidity (diseases other than the one being studied) are typically excluded because the care and outcome of these other diseases can muddy the contrast between experimental and comparison treatments and their outcomes.
2. Patients are excluded if they are not expected to live long enough to experience the outcome events of interest.
3. Patients with contraindications to one of the treatments cannot be randomized.
4. Patients who refuse to participate in a trial are excluded, for ethical reasons described earlier in the chapter.
5. Patients who do not cooperate during the early stages of the trial are also excluded. This avoids wasted effort and the reduction in internal validity that occurs when patients do not take their assigned intervention, move in and out of treatment groups, or leave the trial altogether.
For these reasons, patients in clinical trials are usually a highly selected, biased sample of all patients with the condition of interest. As heterogeneity is restricted, the internal validity of the study is improved; in other words, there is less opportunity for differences in outcome that are not related to treatment itself. However, exclusions come at the price of diminished generalizability: Patients in the trial are not like most other patients seen in day-to-day care.
Example
Chronic obstructive pulmonary disease (COPD) is a common lung condition usually caused by smoking tobacco. Multiple clinical trials have shown that treatment with a combination of inhaled medication (a steroid and a beta-agonist) is effective in avoiding episodes of severe symptoms. How well do the patients in these trials represent the patients seen in the clinic? Researchers examined a database of primary care patients in the United Kingdom, and identified 36,893 patients with COPD (4). They then assessed what proportion of these patients would have been eligible in prior COPD trials. For the nine randomized trials with combination therapy, only 13% of COPD patients would have met the criteria for the trials, such as measurement of disease severity, smoking history, age, and other medical conditions ( Fig. 9.3). Not all eligible patients participate in trials, so the actual representative proportion is likely even lower.
|
View Figure
|
Figure 9.3. Exclusion criteria for trials of chronic obstructive pulmonary disease (COPD) applied to a primary care population. (Data from Halpin DM, Kerkhof M, Soriano JB, et al. Eligibility of real-life patients with COPD for inclusion in trials of inhaled long-acting bronchodilator therapy. Respir Res 2016;17(1):120.)
|
Because of the high degree of selection in trials, it may require considerable faith to generalize the results of clinical trials to ordinary practice settings.
If there are not enough patients with the disease of interest, at one time and place, to carry out a scientifically sound trial, then sampling can be from multiple sites with common inclusion and exclusion criteria. This is done mainly to achieve adequate sample size, but it also increases generalizability, to the extent that the sites are somewhat different from each other.
Large simple trials are a way of overcoming the generalizability problem. Trial entry criteria are simplified so that most patients developing the study condition are eligible. Participating patients have to have accepted random allocation of treatment, but their care is otherwise the same as usual, without a great deal of extra testing that is part of some trials. Follow-up is for a simple, clinically important outcome, such as discharge from the hospital alive. This approach not only improves generalizability, it also makes it easier to recruit large numbers of participants at a reasonable cost so that moderate effect sizes (large effects are unlikely for most clinical questions) can be detected. Pragmatic trials, discussed later in the chapter, are another type of trial that attempts to improve generalizability of results as well.
Intervention
The intervention can be described in relation to three general characteristics: generalizability, complexity, and strength.
First, is the intervention one that is likely to be implemented in usual clinical practice? In an effort to standardize the intervention so that it can be easily described and reproduced in other settings, investigators may cater to their scientific, not their clinical colleagues by studying treatments that are not feasible in usual practice.
Second, does the intervention reflect the normal complexity of real-world treatment? Clinicians regularly construct treatment plans with many components. Single, highly specific interventions make for tidy science because they can be described precisely and applied in a reproducible way, but they may have weak effects. Multifaceted interventions, which are often more effective, are also amenable to careful evaluation as long as their essence can be communicated and applied in other settings. For example, randomized trials of sepsis treatment have used multiple components, such as early administration of antibiotics and intravenous fluids, and flexible interventions to achieve goals for blood pressure and blood oxygen levels (5).
Third, is the intervention in question sufficiently different from alternative managements that it is reasonable to expect that the outcome will be affected? Some diseases can be reversed by treating a single, dominant cause. Treating hyperthyroidism with radioisotope ablation or surgery is one example. However, most diseases arise from a combination of factors acting in concert. Interventions that change only one of them, and only a small amount, cannot be expected to result in strong treatment effects. If the conclusion of a trial evaluating such interventions is that a new treatment is not effective when used alone, it should come as no surprise. For this reason, the first trials of a new treatment tend to enroll those patients who are most likely to respond to treatment and to maximize dose and compliance.
Comparison Groups
The value of an intervention is judged in relation to some alternative course of action. The question is not only whether a comparison is used, but also how appropriate it is for the research question. Results can be measured against one or more of several kinds of comparison groups.
· No Intervention. Do patients who are offered the experimental treatment end up better off than those offered nothing at all? Comparing treatment with no treatment measures the total effects of care and of being in a study, both specific and nonspecific.
· Being Part of a Study. Do treated patients do better than other patients who just participate in a study? A great deal of special attention is directed toward patients in clinical trials. People have a tendency to change their behavior when they are the target of special interest and attention because of the study, regardless of the specific nature of the intervention they might be receiving. This phenomenon is called the Hawthorne effect. The reasons are not clear, but some seem likely: Patients want to please them and make them feel successful. Also, patients who volunteer for trials want to do their part to see that “good” results are obtained.
· Usual Care. Do patients given the experimental treatment do better than those receiving usual care—whatever individual doctors and patients decide? This is the only meaningful (and ethical) question if usual care is already known to be effective.
· Placebo Treatment. Do treated patients do better than similar patients given a placebo—an intervention intended to be indistinguishable (in physical appearance, color, taste, or smell) from the active treatment but does not have a specific, known mechanism of action? Sugar pills and saline injections are examples of placebos. It has been shown that placebos, given with conviction, relieve severe, unpleasant symptoms, such as postoperative pain, nausea, or itching, in about one-third of patients, a phenomenon called the placebo effect. Placebos have the added advantage of making it difficult for study patients to know which intervention they have received (see “Blinding” in the section on Differences Arising After Randomization).
· Another Intervention. The comparator may be the current best treatment. The point of a “comparative effectiveness” study is to find out whether a new treatment is better than the one in current use.
|
View Figure
|
Figure 9.4. Total effects of treatment are the sum of spontaneous improvement (natural history) as well as nonspecific and specific responses.
|
Changes in outcome related to these comparators are cumulative, as diagrammed in Figure 9.4.
Allocating Treatment
To study the effects of a clinical intervention free of confounding, the best way to allocate patients to treatment groups is by means of random allocation (also referred to as randomization). Patients are assigned to either the experimental or the control treatment by one of a variety of disciplined procedures—analogous to flipping a coin—whereby each patient has an equal (or at least known) chance of being assigned to any one of the treatment groups.
Random allocation of patients is preferable to other methods of allocation because only randomization has the ability to create truly comparable groups. All factors related to prognosis, regardless of whether they are known before the study takes place or have been measured, tend to be equally distributed in the comparison groups.
In the long run, with a large number of patients in a trial, randomization usually works as just described. However, random allocation does not guarantee that the groups will be similar; dissimilarities can arise by chance alone, particularly when the number of patients randomized is small. To assess whether “bad luck” has occurred, authors of randomized controlled trials often present a table comparing the frequency in the treated and control groups of a variety of characteristics, especially those known to be related to outcome. These are called baseline characteristics because they are present after randomization and, therefore, should be equally distributed in the treatment groups.
Example
Table 9.1 shows some of the baseline characteristics for a randomized trial of treatment for pregnant women who have had vaginal bleeding, with the goal of increasing the chances of full-term live birth (6). The study enrolled 4,153 women who experienced vaginal bleeding early in the pregnancy, and randomized them to treatment with the hormone progesterone or placebo. The primary outcome was the birth of a live-born baby after at least 34 weeks of gestation. The characteristics shown in Table 9.1 are some of the risk factors associated with premature birth, based on prior studies or clinical experience. Each of these characteristics was similarly distributed in the treatment and placebo groups. These comparisons, at least for the characteristics that were measured, strengthen the belief that the randomization was carried out properly and actually produced groups with similar chances for live birth.
|
TABLE 9.1 Example of a Table Comparing Baseline Characteristics: A Randomized Trial of Progesterone in Women With Bleeding in Early Pregnancy |
|||||||||||||||||||||||||||||||||||||
|
It is reassuring to see that important prognostic variables are nearly equally distributed in the groups being compared. If the groups are substantially different in a large trial, it suggests that something has gone wrong with the randomization process. Smaller differences, which are expected because of chance, can be controlled for during data analyses (see Chapter 6 ).
Differences Arising After Randomization
Not all patients in clinical trials participate as originally planned. Some are found to not have the disease they were thought to have when they entered the trial. Others drop out, do not take their medications, are taken out of the study because of side effects or other illnesses, or somehow obtain the other study treatment or treatments that are not part of the study at all. In this way, treatment groups that might have been comparable just after randomization become less so as time passes.
|
View Figure
|
Figure 9.5. Diagram of stratified randomization. T, treated group; C, control group; R, randomization.
|
Patients May Not Have the Disease Being Studied
It is sometimes necessary (both in clinical trials and in practice) to begin treatment without knowing for certain whether the patient actually has the disease for which the treatment is designed.
Example
Gout, a type of arthritis caused by urate crystals in the joint fluid, usually arises rapidly and affects one or a few characteristic joints. It can be treated with medication to speed recovery. The diagnosis is confirmed with microscopic examination of joint fluid, although less accurate clinical criteria are typically used. In a multicenter international trial, 416 adults diagnosed in emergency departments with gout flares were randomized to treatment with prednisolone (a type of steroid) or indomethacin (a nonsteroidal pain medication) (7). Diagnosis was in almost all cases based on clinical criteria alone. Patients had similar improvement in pain over the next 2 weeks of follow-up, and no major complications occurred. The diagnostic approach may have meant that some of the patients in the study in fact did not have gout, although the use of clinical diagnosis mimicked most actual practice.
Randomized trials that include patients without disease could give misleading results. In such studies, because the number of patients with disease is reduced, it is more likely to miss a clinically important difference if it exists. Also, including patients without the diagnosis can lead to inefficiency in enrolling and gathering data on patients who would not contribute to the study's results (although in the gout study, confirming the diagnoses in all cases with joint analysis would likely have been more costly). Nevertheless, this kind of trial has the important advantage of providing information in a realistic setting on the consequences of a decision that a clinician must make with imperfect information.
Compliance
Compliance is the extent to which patients follow medical advice. The term adherence is preferred by some
P.150
people because it connotes a less subservient relationship between patient and doctor. Compliance is another characteristic that comes into play after randomization.
Although noncompliance suggests a kind of willful neglect of good advice, other factors also contribute. Patients may misunderstand which drugs and doses are intended, run out of prescription medications, confuse various preparations of the same drug, or have no money or insurance to pay for drugs. Taken together, noncompliance may limit the usefulness of treatments that have been shown to work under favorable conditions.
In general, compliance marks a better prognosis, apart from treatment. Patients in randomized trials who are compliant with placebo often have better outcomes than those who are not (8).
Compliance is particularly important in medical care outside the hospital. In hospitals, many factors act to constrain patients' personal behavior and render them compliant. Hospitalized patients are generally sicker and more frightened. They are in strange surroundings, dependent upon the skill and attention of the staff for everything, including their life. What is more, doctors, nurses, and pharmacists have developed a well-organized system for ensuring that patients receive what is ordered for them. As a result, clinical experience and medical literature developed on the wards may underestimate the importance of compliance outside the hospital, where most patients and doctors are and where following doctors' orders is less common.
In clinical trials, patients are typically selected to be compliant. During a run-in period, in which placebo is given and compliance monitored, noncompliant patients can be detected and excluded before randomization.
Crossover
Patients may move from one randomly allocated treatment to another during follow-up, a phenomenon called crossover. If exchanges between treatment groups take place on a large scale, it can diminish the observed differences in treatment effect compared to what might have been observed if the original groups had remained intact.
Cointerventions
After randomization, patients may receive a variety of interventions other than the ones being studied. For example, in a study of asthma treatment, they may receive not only the experimental drug but also different doses of their usual drugs and make greater efforts to control allergens in the home. If these occur unequally in the two groups and affect outcomes, they can introduce systematic differences between the groups that were not present when the groups were formed.
Blinding
Participants in a trial may change their behavior or reporting of outcomes in a systematic way (i.e., be biased) if they are aware of which patients are receiving which treatment. One way to minimize this effect is by blinding, an attempt to make the various participants in a study unaware of the treatment group patients have been randomized to so that this knowledge cannot cause them to act differently, and thereby diminish the internal validity of the study. Masking is a more appropriate metaphor, but blinding is the time-honored term.
Blinding can take place in a clinical trial at four levels ( Fig. 9.6). First, those responsible for allocating patients to treatment groups should not know which treatment will be assigned next making it impossible for them to break the randomization plan. Allocation concealment is a term for this form of blinding. Without it, some investigators might be tempted to enter patients in the trial out of order to ensure that individuals get the treatment that seems best for them. Second, patients should be unaware of which treatment they are taking so that they cannot change their compliance or reporting of symptoms because of this information. Third, to ensure physicians caring for patients in a study cannot, even subconsciously, manage patients differently, physicians should not know which treatment each patient is on. Finally, when the researchers who assess outcomes are unaware of which treatment individual patients have been offered, that knowledge cannot affect their measurements.
The terms single-blind (patients) and doubleblind are sometimes used, but their meanings are ambiguous. It is better simply to describe what was done. A trial in which there is no attempt at blinding is called an open trial or, in the case of drug trials, an open label trial.
In drug studies, blinding is often made possible by using a placebo. However, for many important clinical questions, such as the effects of surgery, radiotherapy, diet, or the organization of medical care, blinding of patients and their physicians is difficult if not impossible.
Even when blinding appears to be possible, it is more often claimed than successful. Physiologic effects, such as lowered pulse rate with beta-blocking drugs and gastrointestinal upset or drowsiness with other drugs, may signal to patients whether they are taking the active drug or placebo.
Assessment of Outcomes
Randomized controlled trials are a special case of cohort studies, and what has already been said about measures of effect and biases in cohort
P.151
studies (see Chapter 6 ) applies to them as well, as do the dangers of substituting intermediate outcomes for clinically important ones (see Chapter 1 ).
|
View Figure
|
Figure 9.6. Locations of potential blinding in randomized controlled trials.
|
Example
Clinical trials may have as their primary outcome a composite outcome, a set of outcomes that are related to each other but are treated as a single outcome variable. For example, in a study of percutaneous versus open surgery valve replacement for aortic stenosis, the composite outcome was the absence of death, stroke, or rehospitalization at 12 months after treatment (10). There are several advantages to this approach. The individual outcomes in the composite may be so highly related to each other, biologically and clinically, that it is artificial to consider them separately. The presence of one (such as death) may prevent the other (such as stroke) from occurring. With more ways to experience an outcome event, a study is better able to detect treatment effects (see Chapter 11 ). The disadvantage of composite outcomes is that they can obscure differences in effects for different individual outcomes. In addition, one component may account for most of the result, giving the impression that the intervention affects the others too. All of these disadvantages can be overcome by simply examining effect on each component outcome separately as well as together.
In addition to “hard” outcomes such as survival, remission of disease, and return of function, a trial sometimes measures health-related quality of life by broad, composite measures of health status. A simple quality-of-life measure used by a collaborative group of cancer researchers is shown in Table 9.2. This “performance scale” combines symptoms and function, such as the ability to walk. Others are much more extensive; the Sickness Impact Profile contains more than 100 items and a dozen categories. Still others are specifically developed for individual diseases. The main issue is that the value of a clinical trial is strengthened to the extent that such measures are
P.152
reported along with hard measures such as death and recurrence of disease.
|
TABLE 9.2 A Simple Measure of Quality of Life. The Eastern Collaborative Oncology Group's Performance Scale |
||||||||||||||
|
Options for describing effect size in clinical trials are summarized in Table 9.3. The options are similar to summaries of risk and prognosis but related to change in outcome resulting from the intervention.
EFFICACY AND EFFECTIVENESS
Clinical trials may describe the results of an intervention in ideal or in real-world situations ( Fig. 9.7).
First, can treatment help under ideal circumstances? Trials that answer this question are called efficacy trials or explanatory trials. Elements of ideal circumstances include patients who accept the interventions offered to them, follow instructions faithfully, get the best possible care, and do not have care for other diseases. Most randomized trials are designed in this way.
|
View Figure
|
Figure 9.7. Efficacy and effectiveness.
|
||||||||
|
TABLE 9.3 Summarizing Treatment Effects |
|||||||||
|
Second, does treatment help under ordinary circumstances? Trials designed to answer this kind of question are called effectiveness trials or pragmatic trials (and sometimes practical trials). These trials are designed to answer real-world questions in the actual care of patients by including the kinds of patients and interventions found in ordinary patient care settings. They are different from typical efficacy trials where, in an effort to increase internal validity, severe restrictions are applied to enrollment, intervention, and adherence, limiting the relevance of their results for usual patient care decisions. In pragmatic trials, patients may not take their assigned treatment. Some
P.153
may drop out of the study, and others find ways to take the treatment they were not assigned. The doctors and facilities may not be the best. In short, these trials describe results as most patients would experience them. The difference between efficacy and effectiveness has been described as the “implementation gap,” the gap between ideal care and ordinary care, and is a target for improvement in its own right.
Example
As noted in an earlier example, trials that demonstrated the efficacy of inhalers for chronic obstructive lung disease (COPD) often included a highly selective, unrepresentative patient population. In addition, the implementation of the treatments, and measurements to assess patient response, are often quite different in practice than in study settings. To see if a combination inhaler medication (fluticasone and vilanterol) was effective in a more realistic setting, investigators randomize 2,799 patients from 75 general practices to this medication or usual care (11). Few patients were excluded, and compared to prior trials they were older, more often smokers, and with more medical conditions. The investigators designed the study to reflect typical care: Patients did not need special studies for the diagnosis, medication in the treatment group could be switched back to usual care, and electronic health records were used to assess study outcomes and safety. At 1-year follow-up, the treatment group had an 8.4% decrease in the rate of severe worsening symptoms compared to usual care.
Efficacy trials often precede effectiveness trials. The rationale is that if treatment under the best circumstances is not effective, then effectiveness under ordinary circumstances is impossible. Also, if an effectiveness trial was done first and it showed no effect, the result could have been because the treatment at its best is just not effective or that the treatment really is effective but was not received.
Intention-to-Treat and Explanatory Trials
A related issue is whether the results of a randomized controlled trial should be analyzed and presented according to the treatment to which the patients were randomized or according to the one they actually received ( Fig. 9.8).
One question is: Which treatment choice is best at the time the decision must be made? To answer this question, analysis is according to which group the patients were assigned (randomized), regardless of whether these patients actually received the treatment they were supposed to receive. This way of analyzing trial results is called an intention-totreat analysis. An advantage of this approach is that the question corresponds to the one actually faced by clinicians; they either offer a treatment or not. Also, the groups compared are as originally randomized, so this comparison has the full strength of a randomized trial. The disadvantage is that to the extent that many patients do not receive the treatment to which they were randomized, differences in effectiveness will tend to be obscured, increasing the chances of observing a misleadingly small effect or no statistical effect at all. If the study shows no difference, it will be uncertain whether the problem is the treatment itself or that it was not received.
Another question is whether the experimental treatment itself is better. For this question, the proper analysis is according to the treatment each patient actually received, regardless of the treatment to which they were randomized. Trials analyzed in this way are called per-protocol analyses (also called explanatory analyses) because they assess whether actually taking the treatments, rather than just being offered them, makes a difference. The problem with this approach is that unless most patients receive the treatment to which they are assigned, the study no longer represents a randomized trial; it is simply a cohort study. One must be concerned about dissimilarities among groups, other than the experimental treatment, and must use methods such as restriction, matching, stratification, or adjustment to achieve comparability, just as one would for any nonexperimental study.
In general, intention-to-treat analyses are more relevant to effectiveness questions, whereas per-protocol analyses are consistent with the purposes of efficacy trials, although aspects of the trial other than how they are analyzed matter too. The primary analysis is usually intention-to-treat, but both are reported. Both approaches are legitimate, with the right one depending on the question being asked. To the extent that patients in a trial follow the treatment to which they were randomized, these two analyses will give similar results.
SUPERIORITY, EQUIVALENCE, AND NONINFERIORITY
Until now, we have been discussing superiority trials, ones that seek to establish that one treatment is better than another, but sometimes the most
P.154
important question is whether a treatment is no less effective than another. A typical example is when a new drug is safer, cheaper, or easier to administer than the established one and would, therefore, be preferable if it were as effective. In noninferiority trials, the purpose is to show that a new treatment is unlikely to be less effective, at least to a clinically important extent, than the currently accepted treatment, which has been shown in other studies to be more effective than placebo. The focus of the question is one directional—whether a new treatment is not worse—without regard to whether it might be better.
|
View Figure
|
Figure 9.8. Diagram of group assignment in intention-to-treat and per-protocol analyses.
|
It is statistically impossible to establish that a treatment is not at all inferior to another. However, a study can rule out an effect that is less than a predetermined “minimum clinically important difference,” also called an noninferiority margin, the smallest difference in effect that is still considered clinically important. The noninferiority margin actually takes into account both this clinical difference plus the statistical imprecision of the study. The following is an example of a noninferiority trial.
Example
Bone infections have been treated in most cases with at least 4 to 6 weeks of intravenous antibiotics, in addition to surgery to remove infected bone when possible. Compared to the pill form of antibiotics taken orally, intravenous antibiotics are more costly, less convenient, and expose patients to potentially severe complications such as deep venous clots and blood infections. Investigators conducted a noninferiority trial to assess if oral antibiotics are an acceptable alternative in these infections (12). There were 1,056 patients with bone or joint infections randomized after 7 days of intravenous antibiotics to an additional 5 weeks of oral or intravenous antibiotics. Patients were followed up for 1 year to determine rates of cure. Noninferiority was defined as the lower limit of the 95% confidence interval for the difference in cure rates (see Fig. 9.9 and Chapter 11 ) that was no more than 7.5% worse for oral than intravenous
P.155
antibiotics. The cure rate was 86.8% in the oral group and 85.4% in the intravenous group, and the difference was 1.4%. The confidence interval (4.9% to -2.2%) was above the -7.5% threshold, thus oral antibiotics met the criterion for noninferiority.
|
View Figure
|
Figure 9.9. Scenarios for a noninferiority trial.
|
Figure 9.9 shows various scenarios for the results of a noninferiority trial. For each scenario there are absolute risk difference estimates (the circle) and the range of likely values in the confidence intervals (the bars). In this illustration the noninferiority margin is -7.5%, as in the example above. The scenarios differ in whether the intervals cross the noninferiority margin, zero, or both. With superiority trials, a treatment is considered to be superior when the risk estimate and confidence intervals are greater than zero. For a noninferiority study, noninferiority is established when the risk estimate and confidence intervals are greater than the noninferiority margin (A, B, C). Conversely, noninferiority cannot be established if any results have confidence intervals that are less than the noninferiority margin (D, E, and F). The oral antibiotics trial discussed above corresponds to scenario B.
Noninferiority trials usually require a larger sample size than comparable superiority trials, especially if the noninferiority margin is small or one wants to rule out small differences. Also, any aspect of the trial that tends to minimize differences between comparison groups, such as intention-to-treat analyses in trials where many patients have dropped out or crossed over or when measurements of outcomes are imprecise, artificially increase the likelihood of finding noninferiority regardless of whether it is truly present—that is, they result in a weak test for noninferiority.
VARIATIONS ON BASIC RANDOMIZED TRIALS
P.156
care units and not in others when physicians see patients in both settings over time? For these reasons, randomizing clusters rather than patients can be the best approach in some circumstances.
There are other variations on the usual (“parallel group”) randomized controlled trials. Crossover trials expose patients first to one of two randomly allocated treatments and later to the other. If it can be assumed that effects of the first exposure are no longer present by the time of the second exposure, perhaps because treatment is short-lived or there has been a “wash-out” period between exposures, then each patient will have been exposed to each treatment in random order. This controls for differences in responsiveness among patients not related to treatment effects.
A stepped wedge cluster trial combines aspects of crossover and cluster designs. Like other cluster trials, randomization occurs at the group level. Unlike traditional cluster trials (where half of the clusters are randomized to treatment or controls throughout the trial), all clusters at the beginning of the trial are controls. Then, at regular intervals (steps), some clusters are randomized to crossover and receive the intervention, and continue the intervention thereafter. At each interval, more clusters move from control to intervention status, until by the end of the trial all clusters are exposed to the intervention.
Analysis in stepped wedge trials must account for two concerns that are not part of usual randomized trials. Like other cluster trials, patients within clusters are often more similar to each other than other clusters. In addition, over time there may be changes in patients, physician practices, and the healthcare setting that affect outcomes in ways other than the intervention. The importance of adjusting for these factors can be seen in a stepped wedge trial of a quality improvement intervention to decrease complications after myocardial infarction (13). Before adjustments, the study found that the intervention group had fewer complications than the controls. However, after accounting for both clustering and temporal trends, there was no difference. Changes in the quality of care at the study sites, in ways that had nothing to do with the intervention, may have been responsible for the improvement found before adjustments.
TAILORING THE RESULTS OF TRIALS TO INDIVIDUAL PATIENTS
Clinical trials describe what happens on average. They involve pooling the experience of many patients who may be dissimilar, both to one another and to the patients to whom the trial results will be generalized. How can estimates of treatment effect be obtained that more closely match individual patients?
Subgroups
Patients in clinical trials can be sorted into subgroups, each with a specific characteristic (or combination of characteristics) such as age, severity of disease, and comorbidity that might cause a different treatment effect. That is, the data are examined for effect modification. The number of such subgroups is limited only by the number of patients in the subgroups, which has to be large enough to provide reasonably stable estimates. As long as the characteristics used to define the subgroups exist before randomization, patients in each subgroup have been randomly allocated to treatment groups. As a consequence, results in each subgroup represent, in effect, a small trial within a trial. The characteristics of a given patient (e.g., the patient might be elderly and have severe disease but no comorbidity) can be matched more specifically to those of one of the subgroups than it can to those of one in the trial as a whole. Treatment effectiveness in the matched subgroup will more closely approximate that of the individual patient and will be limited mainly by statistical risks of false-positive and false-negative conclusions, which are described in Chapter 11 .
Effectiveness in Individual Patients
A treatment that is effective on an average may not work on an individual patient. Therefore, results of valid clinical research provide a good reason to begin treating a patient, but experience with that patient is a better reason to continue or not continue. When managing an individual patient, it is prudent to ask the following series of questions:
· Is the treatment known (by randomized controlled trials) to be efficacious for any patient?
· Is the treatment known to be effective, on average, in patients like mine?
· Is the treatment working in my patient?
· Are the benefits worth the discomforts and risks (according to the patient's values and preferences)?
By asking these questions and not simply following the results of trials alone, one can guard against ill-founded choice of treatment or stubborn persistence in the face of poor results.
N of 1 Trials
P.157
(or N = 1), is an improvement over the time-honored process of trial and error. A patient is given one treatment or another, such as an active treatment or placebo, in random order, each for a brief period of time. The patient and physician are blinded to which treatment is given. Outcomes, such as a simple preference for a treatment or a symptom score, are assessed after each period. After many repetitions patterns of responses are analyzed statistically, much as one would for a more usual randomized controlled trial. This method is useful for deciding on the care of individual patients when activity of disease is unpredictable, response to treatment is prompt, and there is no carryover effect from period to period. Examples of diseases for which the method can be used include migraine headaches, asthma, and fibromyalgia. For all their intellectual appeal, however, N of 1 trials are uncommon and even less often published.
ALTERNATIVES TO RANDOMIZED CONTROLLED TRIALS
Randomized controlled trials are the gold standard for studies of the effectiveness of interventions. Only large randomized trials can definitively eliminate confounding as an alternative explanation for observed results.
Limitations of Randomized Trials
However, the availability of several well-conducted randomized controlled trials does not necessarily settle a question. For example, after five decades and dozens of randomized controlled trials the effectiveness of corticosteroids for septic shock remains controversial. The heterogeneity of trial results seems partly related to differences in dose and duration of the drug, the proportion of patients with relative adrenal insufficiency, and whether the outcome is reversal of shock or survival. That is, the trials were of the same general questions but very different specific questions.
Clinical trials also suffer from practical limitations. They are expensive; major drug trials often cost tens of millions of dollars. Logistics can be daunting, especially in maintaining similar methods across sites in multicenter trials and in maintaining the integrity of allocation concealment. Randomization itself remains a hindrance, if not to the conduct of trials at all then to full, unbiased participation. It is particularly difficult to convince patients to be randomized when a practice has become well established in the absence of conclusive evidence of its benefit.
For these reasons, clinical trials are not available to guide clinicians in many important treatment decisions, but clinical decisions must be made nonetheless. What are the alternatives to randomized controlled trials, and how credible are they?
Observational Studies of Interventions
In the absence of a consensus favoring one mode of treatment over others, various treatments are given according to the preferences of each individual patient and doctor. As a result, in the course of ordinary patient care, large numbers of patients are treated in various ways and go on to manifest the effects. When experience with these patients is captured and properly analyzed, it can complement the information available from randomized trials, suggest where new trials are needed, and provide answers where trials are not yet available.
Example
Non-small cell lung cancer is common in the elderly. Palliative chemotherapy for advanced disease prolongs life, yet the elderly receive chemotherapy less often than younger patients, perhaps because of concern that they would be more likely to develop adverse events such as fever and infections, nerve damage, or deep venous thrombosis. To see if elderly patients in the community really are more likely to experience adverse events, everything else being equal, investigators analyzed data from a cohort study of the care and outcomes of lung cancer (14). There was consistent evidence that the older patients who received chemotherapy were selected to be less likely to have adverse events: Fewer got chemotherapy, they had fewer adverse events before treatment, and they received less aggressive therapy. Even so, adverse event rates after chemotherapy were 34% to 70% more frequent in older patients than in younger ones. This effect could not be attributed to other diseases more common with age—they persisted after adjustment for comorbidity—and is more likely to be the result of age-related decrease in organ function. Whether the benefits of treatment outweigh the increased rate of adverse events is a separate question.
Unfortunately, it is difficult to be sure that observational studies of treatment are not confounded. Treatment choice is determined by a great many factors including severity of illness, concurrent diseases, local preferences, and patient cooperation. Patients
P.158
receiving the various treatments are likely to differ not only in their treatment but in other ways as well.
Especially troubling is confounding by indication (sometimes called “reverse causation”), which occurs when whatever prompted the doctor to choose a treatment (the “indication”) is a cause of the observed outcome, not just the treatment itself. For example, patients may be offered a new surgical procedure because they are at a good surgical risk or have less aggressive disease and, therefore, seem especially likely to benefit from the procedure. To the extent that the reasons for treatment choice are known, they can be taken into account like any other confounders.
Example
Patients with attention-deficit/hyperactivity disorder (ADHD, a condition that usually starts in childhood and is characterized by hyperactivity, impulsivity, and inattention) are often treated with medication. Observational studies have found that patients with ADHD are at increased risk of suicide compared to the general population, and that medication-treated ADHD patients have higher suicide rates than those not on medication. However, it is likely that patients with more severe ADHD have a higher suicide risk, and also are most likely to be prescribed medication. Investigators studied suicide risk in a cohort of 37,936 patients with ADHD in Sweden (15). Patients with ADHD on medication had a 31% increased risk of suicide compared to unmedicated patients, even after accounting for other measured differences. However, the measured differences may not have picked up ways that medicated patients have more severe ADHD. A second analysis attempted to account for this by analyzing only those ADHD patients prescribed medication. This analysis compared suicide rates when these patients were off medication to times they took the medicine; each patient served as his or her own control. There was no increased suicide risk with medication use overall in this within-patient comparison, and decreased suicide rates with one class of medication (stimulants such as methylphenidate).
Clinical Databases
Sometimes, databases are available that include baseline characteristics and outcomes for a large number of patients. Clinicians can match characteristic of a specific patient to similar patients in the database and see what their outcomes were. When use of the database is not part of formal research, there is no accounting for confounding and effect modification, but the predictions do have the advantage of being about real-world patients, not those as highly selected as in most clinical trials.
Randomized Versus Observational Studies?
Are observational studies a reliable substitute for randomized controlled trials? With controlled trials as the gold standard, most observational studies of most questions get the right answer. However, there are dramatic exceptions. For example, observational studies have consistently shown that antioxidant vitamins are associated with lower cardiovascular risk, but large randomized controlled trials have found no such effect. Therefore, clinicians can be guided by observational studies of treatment effects when there are no randomized trials to rely on, but they should maintain a healthy skepticism.
Well-designed observational studies of interventions have some strengths that complement the limitations of usual randomized trials. They count the effects of actual treatment, not just of offering treatment, which is a legitimate question in its own right. They commonly include most people as they exist in naturally occurring populations, in either clinical or community settings, without severe inclusion and exclusion criteria. They can often accomplish longer follow-up than trials, matching the time it takes for disease and outcomes to develop. By taking advantage of treatments and outcomes as they happen, observational studies (especially case control studies and historical cohort studies using health records) can answer clinical questions more quickly than it takes to complete a randomized trial. Of course, they are also less expensive.
Because an ideal randomized controlled trial is the standard of excellence, it has been suggested that observational studies of treatment be designed to resemble, as closely as possible, a randomized trial of the same question (16). One might ask, if the study had been a randomized trial what would be the inclusion and exclusion criteria (e.g., excluding patients with contraindications to either intervention), how would exposure be precisely defined, and how should drop-outs and crossovers be managed? The resulting observational study cannot be expected to avoid all vulnerabilities, but at least it would be stronger.
PHASES OF CLINICAL TRIALS
For studies of drugs, it is customary to define three phases of trials in the order they are undertaken. Phase I trials are intended to identify a dose range
P.159
that is well tolerated and safe (at least for high-frequency, severe side effects) and include very small numbers of patients (perhaps a dozen) without a control group. Phase II trials provide preliminary information on whether the drug is efficacious and the relationship between dose and efficacy. These trials may be controlled but include too few patients in treatment groups to detect any but the largest treatment effects. Phase III trials are randomized trials and can provide definitive evidence of efficacy and rates of common side effects. They include enough patients, sometimes thousands, to detect clinically important treatment effects and are usually published in biomedical journals.
Phase III trials are not large enough to detect differences in the rate, or even the existence, of uncommon side effects (see discussion of statistical power in Chapter 11 ). Therefore, it is necessary to follow up very large numbers of patients after a drug is in general use, a process called postmarketing surveillance.