Showing posts with label Randomized Trials. Show all posts
Showing posts with label Randomized Trials. Show all posts

Monday, May 6, 2013

What the Oregon health study shows or doesn't show

Some thoughts about the Oregon Health Study - which is being dissed by MR here and here.

First off the comparison of this study to Reinhart and Rogoff is way off base. Reinhart and Rogoff were recalcitrant in sharing their data. Had they been more forthcoming in the first place it is possible that someone would have pointed out the Excel error much earlier. Instead they left all the researchers puzzled over their attempts to replicate their findings. This is NOTHING like Reinhart Rogoff.

What if they had found positive effects? Then I think the debate would have shifted to size effects and whether the results were meaningful in any sense. The critics are leaping to the "no result" finding as a way to justify their opposition. They would probably have leaped to a positive result by pointing out the lack of meaningful results regardless.

RCTs are pretty useless in this debate. Think of their outcomes - blood pressure, hypertension and cholesterol levels. How can having insurance by itself promote better outcomes? To actually change these outcomes, insurance holders have to actually change their behavior. There is a lot of behavioral economics evidence out there to argue against anyone actually doing anything about this regardless of health insurance status.

This is an example of RCTs taking away the analysts' ability to think straight. Retrospection would have predicted the null effect. The low income population makes them even more constrained in their ability to affect lifestyle changes that would positively affected the outcome measures. I have some doubt as to whether even a high income population would have been able to affect the lifestyle changes necessary to have a detectable effect.

This is an example of RCT measuring the wrong outcome. We only measure what we can see or even worse - what is easiest to measure. By some kind of fluke these are the outcomes that are hardest to change. I have to change my diet, my exercise regimen, my sleep habits, and so forth. If this is hard for a middle class person how much harder would it be for a low income group?

Effects of the recession on health? Is it possible that the stress of the recession and the job market is driving the results in both treatment and control?

This randomized study should have been titled something like "How much do people take their doctors' advice in changing their lifestyles? An analysis of low income population using Medicaid lotteries" or something similar. Would the findings have made any headlines then?

Alternatively, if the study participants were given medications to treat their hypertension, cholesterol, etc. then the study would have been about the efficacy of drugs on the low income population. Or perhaps prescription drug compliance among the low income population.

The comeback "You still buy health insurance don't you?" is a reasonable response contrary to what Tyler Cowen thinks. And this returns to the possibility that the benefits of ACA and expansion of Medicaid was oversold. Why do we buy health insurance - surely not to become healthier. We buy insurance to preserve the status quo, i.e. I don't want my health to get worse and if it gets worse then at least I can do something about it. I have health insurance and I can go to the doctor. And when we think about health we generally don't think in terms of cholesterol or blood pressure. We think of actual illnesses, pain, discomfort, etc. and again evidence from behavioral economics may be useful here.

What is going here I think is that possibly and perhaps in fact, very likely that supporters of ACA felt they had to justify the expansion of Medicaid by some evidence based reason and like their opponents who will grab at anything they can to fight against it they fell onto the belief that health insurance improves overall health - not specific illnesses, but overall health because it was easier to measure in an RCT.

Instead of making arguments for ACA on grounds of equity and perhaps even a right to adequate care and a right not to impose the costs of emergency room visits on others, etc. they have backed themselves into a corner with this null finding.

Wednesday, July 4, 2012

Does correcting for self selection change the policy question

Consider an experiment of whether job training after layoff increases the probability of being re-employed. A ‘naive’ treatment effect would be to compare the effects of those who enrolled in job training and those who didn’t. The estimated effect would then be the difference in likelihood of being employed for those with job training and those without. The policy question addressed here is whether job training increases the likelihood of employment.

But the econometrician would argue that those who did not enrol in job training are different from those who did and that these characteristics are unobservable to him (the econometrician, e.g. motivation might be unobservable). In order to accurately estimate the impact of job training one would have to compare apples to apples, i.e. those who applied for job training but were (randomly) rationed out of the program. This gives the correct estimated impact. But the policy question now seems to be whether those who applied for job training but were not denied increases the likelihood of being employed. I would argue that this is NOT the policy question of interest.

The policy instrument is to shift people into job training - assuming that the impact is or can be positive. But by estimating the impact only for “motivated” people this naturally assumes that the unmotivated will not be treated. Suppose the following:


  • A randomized control trial of a job training program is run and impacts estimated.
  • The impacts are found to be large and cost benefit analysis shows that the benefits are positive on net.
  • What happens when the program is scaled up, i.e. rolled out to the entire population of unemployed (instead of just to the treatment and control who were "similar" in the trial)? Should we assume that the impacts would still be the same as in the RCT? Are the participants on the now scaled up program still similar? An RCT advocate would argue yes - but - isn't the original intent of scaling up a program to get as many people as possible to participate regardless of the original composition of the treatment and control groups?
  • Should the scaled up program be the same as the RCT, i.e. a static program that doesn't enroll anyone but just allows the "motivated" to enroll themselves? What if there was an effort to try to get the recalcitrant unemployed into the program - after all since the benefits are positive, don't we want to extend the benefits to as many as possible? If there were such an effort would the estimated impacts still be the same as in the trial?
  • Suppose that after the completion of the trial we find that the population of unemployed has changed so that there are now more women than men? Do we deny one gender the treatment because it is no longer the same as those in the randomized trial?



Friday, June 29, 2012

What RCTs can reveal

From the NBER WP:
Up in Smoke: The Influence of Household Behavior on the Long-Run Impact of Improved Cooking Stoves by Rema Hanna, Esther Duflo, Michael Greenstone

Abstract:
It is conventional wisdom that it is possible to reduce exposure to indoor air pollution, improve health outcomes, and decrease greenhouse gas emissions in the rural areas of developing countries through the adoption of improved cooking stoves. This belief is largely supported by observational field studies and engineering or laboratory experiments. However, we provide new evidence, from a randomized control trial conducted in rural Orissa, India (one of the poorest places in India), on the benefits of a commonly used improved stove that laboratory tests showed to reduce indoor air pollution and require less fuel. We track households for up to four years after they received the stove. While we find a meaningful reduction in smoke inhalation in the first year, there is no effect over longer time horizons. We find no evidence of improvements in lung functioning or health and there is no change in fuel consumption (and presumably greenhouse gas emissions). The difference between the laboratory and field findings appear to result from households’ revealed low valuation of the stoves. Households failed to use the stoves regularly or appropriately, did not make the necessary investments to maintain them properly, and usage rates ultimately declined further over time. More broadly, this study underscores the need to test environmental and health technologies in real-world settings where behavior may temper impacts, and to test them over a long enough horizon to understand how this behavioral effect evolves over time.

Thoughts:
An incredible amount of work went into this in terms of data collection.
Forcing technological adoption when the population isn’t ready for it will not lead to any measurable impact.

Friday, June 4, 2010

What is the right counterfactual

It's not always obvious and the obvious - no treatment versus treatment is not always right. This post on airline deregulation was a good reminder:

The catch is that all such economic comparisons must be counterfactual: they must show an improvement not with respect to CAB [Civil Aeronautics Board]-set fares of the late-1970s, but rather with respect to what reasonably competent regulation could have produced under the other circumstances of the deregulated era. ... If the comparison exercise is tough by the (inappropriate) historical yardstick thanks to declines in (average) service quality and the airline industry’s trail of fleeced stakeholders, then the counterfactual comparison is going to be tougher still thanks to a couple of factors that should have produced large declines in airline costs and hence fares even in the absence of deregulation.

The factors of note are a pair of technological advancements — the development of high bypass ratio turbofans suitable for shorter-haul airliners and the demise of the flight engineer’s job thanks to cockpit automation, both of which have origins predating deregulation — and the long secular decline in oil prices through the deregulated era’s zenith prior the crash of the 1990s stock market bubble.

Saturday, May 1, 2010

Economic models redux

There is widespread agreement that economic and statistical models have failed us. (See some old posts here here, here, here, and here.) Some recent posts in the blogosphere continue to remind us of this but have not provided any alternatives. The main point that all model users fail to remember is that ALL MODELS ARE FALSE. It is when they decide to drink their own model elixir and to substitute models for reality that instead of living in the real world they begin to live in a dream like state. (If I could substitute my spouse for a model I'd probably live in a dream like state as well.)

First off, John Cassidy is absolutely correct:
To repeat myself, the problem wasn’t so much with the models themselves, but with how they were utilized. Rather than being used to discipline individual traders and trading desks, they were used to justify bigger and bigger speculative positions, and more and more leverage.

I take exception however to the tone of the article that blames model builders without considering the incentives. What incentive is there to tell a CEO or a trader that he should not do a trade even when the payoffs are huge because even though the model may be right, there is a small chance that it could be wrong?

Another recent post attacks utility maximization. I am sympathetic to to the idea that life is not all utility maximization. My more general attack on microeconomics here - in fact, the failure of contract theory, principal-agent models and pay for performace indicates a failure of the expected utility maximization. These models typically assume that utility is unbounded from below - yet as the blogger cites Carol Graham:

"People seem to be able to adapt to high levels of adversity, poor health and all kinds of things and retain their natural cheerfulness or their natural happiness…People really can adapt to adversity."

I would also indicate that Herbert Simon's concept of satisficing has been around a long time (and surprisingly did not make it into the bloggers comments) and for all purposes appear to be a possible alternative but has not made it into mainstream economics.

While there may be a case for starting anew, it is still possible to retain the EU framework which almost all economists love. Hyperbolic discounting is one that comes to mind. Moreover, it is also possible to re-cast Carol Graham's comment into a framework of binding budget constraints. And while it is easy to talk about models in lyrical terms or using analogies, the preferred language of economists is mathematics. Whether this is restraining economists is a debate for another post, i.e. are economists constrained by their tools (mathematics) or by their lack of imagination? For instance, economists (still?) struggle to put Nelson and Winter's evolutionary approach into mathematical models.

This point is brought recently to force by Rajiv Sethi's discussion on John Geanakoplos' Leverage Cycle. Rajiv blogs:

David at Deus Ex Macchiato agreed that the work is important, but added:

What astonishes me however is that this is in any way news to the economics community. Ever since Galbraith’s account of the importance of leverage in the ‘29 crash, haven’t we known that leverage determines asset prices, and that the bubble/crash cycle is characterised by slowly rising leverage and asset prices followed by a sudden reverse in both?

I would add that it is one thing for Galbraith to articulate an idea but what is more important in PhD programs these days is not exploring ideas but model-building and sometimes putting ideas into models is easier said than done. One idea dating at least to Zarnowitz tha thas been challenging is that every boom sows the seeds for its eventual bust. It is possible that economists have not been re-reading popular works as much as they should (after all these are frowned upon since they should be building models or extending DSGE models) and Rajiv agrees:

Implicit in David's question is the accusation that the training of professional economists has become too narrow, and on this point I believe that he is absolutely correct.

Econbrowser's Jim Hamilton agrees that economists think too narrowly:

We're fond of building models of rational people reacting in a predictable way to the incentives they face; if their behavior changes, we look for an explanation in terms of changed incentives. It turned out to be in the fund managers' short-term interests to go with the more aggressive strategy, with disastrous longer-run consequences. Was the manager rational before 2006 and irrational after 2006, or did the incentives fundamentally change?

One of the explanations I sometimes hear is a story about "search for yield," which appears to be a combination of the two interpretations, attributing some of the altered risk-taking strategy to the period of very low interest rates in the preceding years. If this indeed accounts for some of the changed behavior by lenders, it is a channel for the transmission of monetary policy to the economy that's left out of the Fed's standard models, and another reason to be cautious about overestimating the benefits that are practical to achieve from a stimulative monetary policy.

Finally, Sciencenews reminds us:

The “scientific method” of testing hypotheses by statistical analysis stands on a flimsy foundation. Statistical tests are supposed to guide scientists in judging whether an experimental result reflects some real effect or is merely a random fluke, but the standard methods mix mutually inconsistent philosophies and offer no meaningful basis for making such decisions. Even when performed correctly, statistical tests are widely misunderstood and frequently misinterpreted. As a result, countless conclusions in the scientific literature are erroneous, and tests of medical dangers or treatments are often contradictory and confusing.

Even randomized control trials should be viewed with skepticism:

Statistical problems also afflict the “gold standard” for medical research, the randomized, controlled clinical trials that test drugs for their ability to cure or their power to harm. Such trials assign patients at random to receive either the substance being tested or a placebo, typically a sugar pill; random selection supposedly guarantees that patients’ personal characteristics won’t bias the choice of who gets the actual treatment. But in practice, selection biases may still occur, Vance Berger and Sherri Weinstein noted in 2004 in ControlledClinical Trials. “Some of the benefits ascribed to randomization, for example that it eliminates all selection bias, can better be described as fantasy than reality,” they wrote.

Randomization also should ensure that unknown differences among individuals are mixed in roughly the same proportions in the groups being tested. But statistics do not guarantee an equal distribution any more than they prohibit 10 heads in a row when flipping a penny. With thousands of clinical trials in progress, some will not be well randomized.

Friday, April 23, 2010

Long term impact evaluation

Michael Clemens shows healthy skepticism toward the impact of the Millenium Villages Project. I concur with his views although what he is asking for is highly unrealistic - a long term impact evaluation beyond 5 years. What he asks for is so unrealistic that if the MVP does not comply (which I predict, it will not) that he has more or less set MVP up to fail.

There are literally no impact evaluations beyond a short time frame - 5 years may even be stretching it. Even clinical trials have problems beyond the 3-year period due to sample attrition. Even the Head Start Impact Study is not scheduled to go beyond 5 years and I would argue that this is an important policy that needs to be carefully studied.

There is also a problem with the MVP project that a randomized trial cannot answer:
"The project deploys a broad package of interventions for five years in each village, including distribution of fertilizer and insecticide-treated bednets, school construction, HIV control, microfinance, electric lines, road construction, piped water and irrigation lines, mobile phones, and several others."

Any time a treatment consists of varying sub-treatments and dosage whose levels are hard to measure you can bet that the even if the impacts were positive, the causal effect remains a black box. Which sub-treatment was more effective? At what dosage? A village level randomized trial will not be able to conclusively answer this because the object being randomized is a village and the sample will be small (even though the number of people in a village may be large).

I would still advocate a long term study. I would not bother with trying to find decent controls at the the time of randomization because the study at this level will have low power. I would however, collect as much data as I can from as many villages as I can. In terms of data collection (if it were in an advanced country) I would try to attain what has been done by the PSID although I would have more observations. Being that this is in Africa, it will be hard to get this done. Does this mean that even if we fall short of the data collected that we should not do it? Absolutely not. The funding requirements are such that someone else besides MVP may have to step in.

Monday, April 19, 2010

What if our randomized trial was incorrectly implemented

In an interesting article on estrogen in the NYT:

... the Women’s Health Initiative, or W.H.I. It was a federally financed examination of adult women’s health, extraordinary in scale and ambition, that started up in the early 1990s; one of its drug trials enrolled more than 16,000 women for a multiyear comparison of hormone pills versus placebos. On July 9, 2002, W.H.I. investigators announced that they had ended the trial three years early, because they were persuaded that it was dangerous to the hormone-taking participants to let them continue. ...

First of all, ... there are different forms of estrogenic molecules — ... estradiol. It’s [Estradiol] not the estrogen used in the W.H.I. study. Pharmaceutical estradiol like mine comes from plants whose molecules have been tweaked in labs until they are atom for atom identical to human estradiol, the most prominent of the estrogens premenopausal women produce naturally on their own. The W.H.I. estrogen, by contrast, was a concentrated soup of a pill that is manufactured from the urine of pregnant mares. ...

The progesterone he prescribed ... , like the estradiol, is a molecular replica of the progesterone women make naturally. It’s different from the progesteronelike synthetic hormone that was used for the W.H.I. study that ended in 2002. That medication was a formulation whose multisyllabic chemical name shortens to MPA and which has a problematic back story of its own: MPA takes care of the uterine-cancer risk, but there’s reason to suspect it may be a factor in promoting breast cancer. And it’s ingested as a pill, which means that like equine estrogens ... MPA metabolizes through the liver, possibly creating additional complications en route, before going about its business.

The biggest difference between me and the W.H.I. women, though, has to do with age and timing. I started on the patches while my own estrogen, pernicious though its spikes and plummets may have been, was still floating around at more or less full strength. The average age of the W.H.I. women was just over 63, though the study accepted women as young as 50. More significant, though, most of them were many years past their final menstrual period, which is the technical definition of menopause, when they began their trial hormones. The bulk of the group was at least 10 years past; factoring in the oldest women, the average number of years between the volunteers’ menopause and their start on the trial medications was 13.4.


The bottom line:
... one undiplomatic critic sum up the W.H.I. as “the wrong drugs, tested on the wrong population,”

Gelman's blog also discusses another randomized trial that did not fully answer the question on PSA screening. More here.

Wednesday, March 31, 2010

Questioning the flu vaccine

I enjoyed Shannon Brownlee and Jeanne Lenzer's boldness in questioning everything we thought we knew about the flu vaccine. I felt that they did not go far enough. The main distraction was the focus on mortality outcomes. This is an example of an irrelevant outcome of a study. Many healthy people get vaccinated not to prevent dying from the flu but to prevent the flu in the first place! I get a flu shot because getting the flu is a pain the ass, not because I'm afraid of dying from it.

Moreover, if vaccination is effective, it also prevents other people from getting sick. It is possible, for instance, to have an whole unit or group out sick just because one person did not get a vaccination. (externalities)

However, they only touch on the general efficacy of the flu vaccine, i.e.:

... vaccine “mismatches” occurred in 1968 and 1997: in both years, the vaccine that had been produced in the summer protected against one set of viruses, but come winter, a different set was circulating. In effect, nobody was vaccinated. Yet death rates from all causes, including flu and the various illnesses it can exacerbate, did not budge.

I would like to be able to replace the phrase "death rates" with "infection rates".

... Studies show that young, healthy people mount a glorious immune response to seasonal flu vaccine, and their response reduces their chances of getting the flu and may lessen the severity of symptoms if they do get it. But they aren’t the people who die from seasonal flu. ... Is vaccine necessary for those in whom it is effective, namely the young and healthy?

Focusing only on the young and healthy, a randomized trial would not have the same ethical problems as a randomized trial of the elderly. They also do not consider the externality effects of vaccination (if vaccination is indeed effective):

From a paper I have yet to read:

Vaccination provides indirect benefits to the unvaccinated. Despite its important policy implications, there is little analytical or empirical work to quantify this externality, nor is it incorporated in a number of cost-benefit studies of vaccine programs. We use a standard epidemiological model to analyze how the magnitude of this externality varies with the number of vaccinations, vaccine efficacy, and disease infectiousness. We also provide empirical estimates using parameters for influenza and mumps epidemics. The pattern of the externality is complex and striking, unlike that suggested in standard treatments. The size of the externality is not necessarily monotonic in the number vaccinated, vaccine efficacy, nor disease infectiousness. Moreover, its magnitude can be remarkably large. In particular, the marginal externality of a vaccination can be greater than one case of illness prevented among the nonvaccinated, so its omission from policy analyses implies serious biases.

Monday, January 25, 2010

Replicating experiments with propensity score matching

I've been trying to learn some propensity score matching and consequently have been perusing some papers. The Smith-Todd paper "Does Matching Overcome Lalonde's Critique of Nonexperimental Estimators?" was a good useful starting point for me. I was a little perplexed by the desire of the authors to "hit" the experimental estimate though. Presumably, if the experiement were repeated, it would not achieve the same treatment effect as the original - the treatment effect has a standard error or confidence interval around it.

Agodini and Dynarski's paper "Are Experiments the Only Option? A Look at Dropout Prevention Programs" was also interesting. The authors don't try to match the experimental effects but ask if the direction of the experimental effect can be concluded based on propensity score methods. Also interesting was the whole question surrounding the standard error of the propensity score estimate - whether the simple random sample estimate is "close" to the bootstrapped estimate or not since there have been claims that the standard error from the propensity score estimator is from an estimate based on nonlinear methods and hence not reliable. They find that the bootstrapped estimates are similar to an SRS standard error.

Wednesday, January 28, 2009

Crime and Section 8

The Atlantic article on crime and the spread of Section 8 vouchers to the suburbs which in turn also spreads crime was compelling though not entirely convincing. For those who believe that you can take the poor out of the crime ghettos but not the crime out of the poor ghettos this article provided the ammunition.

On a theoretical level the idea is that as more poor people with Section 8 vouchers move out of the inner city and as they begin locate closely to one another again they form a new pocket of crime. While one or two househoulds with a Section 8 voucher in a suburb may not result in an increase in crime, maybe ten or more households might be sufficient to increase crime because it is more likely that at least one or two households have a criminal past and are more likely to vicitimize one another and others in the new neighborhood. I actually have a strong prior on this though it is only theoretical.

This paper by Jeff Kling and Jens Ludwig ("Is Crime Contagious") uses data from MTO randomization/demonstration program finds results that I would characterize as mixed. I called MTO a demonstration program because the HUD site indicates this is what it was although the analysts who ran it call it a randomized trial. I hestitate to call this a randomized trial because if memory serves they had difficulty recruiting households to participate.

The authors conclude:
Our results are not consistent with the idea that contagion explains as much of the across neighborhood variation in violent crime rates as previous research suggests. We do not find any statistically significant evidence that MTO participants are arrested for violent crime more often in communities with higher violent crime rates. Our estimates enable us to rule out very large contagion effects, but not more modest associations. This general finding holds for our full sample of MTO youth and adults as well as for sub-groups defined by gender and age, and it also holds when we simultaneously instrument for neighborhood racial segregation or poverty rates.

I don't know how this translates into rising "local" crime rates that are more of interest to police in the Atlantic article. The article is concerned that crime rates rise in suburbs and the Kling and Ludwig do not test that crime rates rise in the neighborhoods in which the MTO participants move into. The definition of "neighborhood" is one difficulty althought it may be possible to test for the significance of whether MTO participants are arrested for violent crime more often in communitites with lower violent crime rates. It is possible that there was not enough variation in the data to perform this test.

Wednesday, January 14, 2009

Thoughts on randomization

Mostly triggered by Chris Blattman's advice to PhDs:
"The randomized evaluation is just one tool in the knowledge toolbox. It's currently the rage, but that means it will probably be old news by the time you finish your PhD."

One of the problems with randomized trials is that is is a black box. We understand very little or we may think we understand a lot. There is also a lot of potential subgroup interaction that needs to be tested.

All this points to the fact that if we have to do a randomized trial then we don't really understand the mechanism of how the treatment works (and this also applies to medical "science"/drug therapy, etc.). And if it does work to our expectations then it validates our priors and perhaps advances the field a little. But does it really advance our understanding of the causal underlying mechanism? All we can point to are suggestions that our limited understanding is validated but we could still be spectacularly wrong.

Another problem with randomized trials is that it usually is never the last word. (Perhaps repeated randomized trials can provide the last word, but rarely one randomized trial.) Again, this is because if a theory accords with my priors and the results of my hypotheses are rejected it doesn't seem to lower my priors as much as it should - mainly because the "theory" sounds so sensible and plausible. So it must be something with the way the trial is conducted. For instance, the effects of Head Start on children and the disappointing First Year results - yet the underlying premise of Head Start is so strong that it will not go away.

Randomized trials also do not address the question: How will it work for me? And this is particularly true for drugs. I really do not care about average treatment effects of the average treatment effects on my subgroup. And it is this thinking that leads to experimentation and continuing treatment using "less than acceptable" methods or alternative methods. This would be the test of our understanding - if we can predict individual results then we can claim to have the final word on causality.

Wednesday, September 17, 2008

What randomized trials do not reveal

From the same story in this post:

In 1983 the first rotavirus vaccine was ready for testing. ... From all vantages, the first trial, conducted in Finland was a landmark success: the vaccine reduced the changes that a vaccinated child would get severe rotavirus by 88 percent demonstrating that immunity could be induced with a live oral vaccine. Moreover, the vaccine had no troubling side effects.

Encouraged, Smith Kline-RIT (now GlaxoSmithKline Biologicals) launched trials in other countries, and by the late 1980s the end of rotavirus-related deaths seemed at hand. But then the results from trials in Africa and Peru proved inconsistent and disappointing. Lacking certainty about the reasons for the troubles - although poor health untreated infections, malnutrition and parasites are known to affect a child's immune response to vaccines - the company put its rotavirus program on hold.

I highlight this and the earlier post because development economists have begun to embrace randomized trials as a solution to finding out what kind of development aid works. Randomized trials have had a long history in medicine and economists have the advantage of adopting the best practices from those experiences. While they have not ignored the main criticism that randomized trials are a 'black box' the enthusiasm that has been coursing through the veins of development economists is palpable. Why have economists so ardently embraced randomized trials?

1. It gets at the question of causation. No more instrumental variables! The treatment-control differences are a clear cut answer as to what works and what doesn't. It doesn't matter if we do not know why it doesn't work -- accountability is ensured when a program is thrown out on the basis of a randomized trial. Yet as the above story shows, it can matter to know why a program works in one location but not another. As Atul Gawande writes so eloquently in the New Yorker, even knowing what works is not sufficient. We need to get at the question: why does it work better in one place and not another. While it is not a story about randomized trials, it is a story about treatment.

Over the phone, the doctor told Honor that her daughter’s chloride level was far higher than normal. Honor is a hospital pharmacist, and she had come across children with abnormal results like this. “All I knew was that it meant she was going to die,” she said quietly when I visited the Pages’ home, in the Cincinnati suburb of Loveland. The test showed that Annie had cystic fibrosis. ...

The one overwhelming thought in the minds of Honor and Don Page was: We need to get to Children’s. Cincinnati Children’s Hospital is among the most respected pediatric hospitals in the country. It was where Albert Sabin invented the oral polio vaccine. The chapter on cystic fibrosis in the “Nelson Textbook of Pediatrics”—the bible of the specialty—was written by one of the hospital’s pediatricians. The Pages called and were given an appointment for the next morning. ...

The one thing that the clinicians failed to tell them, however, was that Cincinnati Children’s was not, as the Pages supposed, among the country’s best centers for children with cystic fibrosis. According to data from that year, it was, at best, an average program. This was no small matter. In 1997, patients at an average center were living to be just over thirty years old; patients at the top center typically lived to be forty-six. By some measures, Cincinnati was well below average. The best predictor of a CF patient’s life expectancy is his or her lung function. At Cincinnati, lung function for patients under the age of twelve—children like Annie—was in the bottom twenty-five per cent of the country’s CF patients. And the doctors there knew it.

2. The earlier post highlighted that even within a randomized trial, subgroup interactions can be important. In that case, it was infants under 3 months. The search for subgroup effects (i.e. is the treatment most effective for this particular subgroup?) can easily become a fishing expedition or is sometimes referred as data mining. The number of hypotheses tested can easily reach into the hundreds. This type of analysis essentially puts the analyst back where they were before the randomized trial.

Because a randomized trial is so expensive to run, finding a positive treatment effect on a particular subgroup, means that another randomized trial on that subgroup cannot be repeated. The analyst has to stand by the statistical analysis which raises issues of multiple comparisons, sample power, independence and a whole host of other issues that the analyst had tried to avoid in the first place by putting his faith on the results of a randomized trial.

3. One of the main reasons for using randomized trials is to get at the selection problem. Looking at the results of a program by comparing those in and out of the program is biased because some participants self select into the program. However, the selection problem has only been pushed back one step in a randomized trial. There are experiments where those selected to receive the treatment refuse the treatment (known as "non-compliance") while those who are selected to receive the placebo somehow manage to circumvent the experimental controls to receive treatment ("crossovers"). In randomized trials and in econometrics, the selection problem has given rise to a whole host of estimators: ATE (average treatment effects), ITT (intent to treat), TOT (treatment on treated) and some others I'm not familiar with. Not surprisingly, the compliance problem has own estimator: CACE (complier average causal effect).

Just as growth econometrics has abandoned its search for causes of economic growth and settled for correlations, development economists seem to have abandoned its search for explanations and settled for finding out what works without knowing fully why it works.

This is meant to only to be a note of caution and nothing more but economists need to look at the experiments they are conducting and subject themselves to a cost-benefit analysis: Are the costs of running randomized trials greater than the benefits received by the participants? What about the benefits of the data that is gathered - can they be used to advance the knowledge in the field or are the results of these 'black box' experiments to be filed away and forgotten when the fad is over?

Why cost benefit analysis will always be unacceptable

A rotavirus vaccine is developed and the vaccine goes to clinical trials. (Full details here - preview only.)

In 1991 the Food and Drug Administration granted the pharmaceutical company Wyeth Ayest (later Wyeth Pharmaceuticals) permission to make and test this vaccine, which they named RotaShield. Over the next five years it launched large-scale clinical trials in the U.S., Finland and Venezuela, verifying RotaShield's safety, ability to induce a protective immune response and lasting efficacy. ... Over the next nine months months more than 600,000 children received an estimated 1.2 million does of RotaShield. ...

The disaster struck. In 1999 several infants suffered a serious complication within two weeks of receiving the vaccine: a segment of the intestine folded into a nearby region (like part of a telescope collapses into another), creating a blockage called intussusception. The condition can be excruciatingly painful and must be quickly reversed with either an air or fluid enema or fixed surgically. In rare cases, the intestine perforates and the infant dies. The CDC, which was monitoring experience with RotaShield, called for an immediate halt to the immunization program, thereby sinking a vaccine that had taken 15 years and several hundred million dollars to launch.

The agency initially estimated the risk to be one intussusception in 2,500 vaccine recipients, which was considered unacceptable. Later studies pegged the probability at only one in 11,000. Then Lone Simonsen of the NIH correlated risk with age: infants younger than three months were in less danger than older ones. If the vaccine were given only to young babies, the likelihood of intussusception could drop 10-fold to perhaps one in 30,000.

The new data raised new questions. Was this risk acceptable in the U.S., where children are often hospitalized but rarely die of rotavirus? Were the odds more palatable in the developing world, where one child in 200 dies of rotavirus? If 150 lives could be saved for each complication from intussusception, might the risk be justified? Given these statistics, was it unethical, in fact, to withhold a vaccine that might save half a million lives a year? Or no matter what the risk-benefit analysis showed, was it unethical to market a vaccine in the developing world that had been withdrawn from use in the U.S.?

The CDC and the WHO called a meeting of policymakers from developing countries. After heated discussion, science bowed to politics. As a high-ranking Indian official said, "I know this vaccine would save 100,000 children in my country. But when the first case of intestinal blockage occurred, I would not be forgiven for allowing a vaccine that had been withdrawn in the United States to be used in my country."