Evidence-based policy: Why one good trial isn't enough
Share
Subscribe
Medicine has run rigorous clinical trials since the 1940s. Development economists borrowed the structure of the trials during the credibility revolution. Chris Cotton (Queen's University) argues that they misunderstood an important part of this revolution. One successful study is not a mature evidence base, and policies built on too little evidence tend to disappoint at scale.
His answer is a framework of Policy Evidence Readiness Levels, or PERLs, just published in Science. PERLs 1 through 4 categorise the weight of evidence behind an intervention, and are inspired by the scale NASA uses to judge whether a technology is ready for a space mission. He tells Tim Phillips which development programmes have already climbed to the top rung, which were implemented at scale on too little evidence (and what happened when they were), and why his framework, by his own admission, is still a modest PERL level 1.
The research behind this episode
Cotton, Christopher. 2026. "Rigor Is Not Readiness: The PERL Framework for Evidence-Based Policy." Science 393 (6810): 463-465.
To cite this episode
Phillips, Tim, and Christopher Cotton. 2026. "Evidence-based policy: Why one good trial isn't enough." VoxDev Talks (podcast).
About the guest
Christopher Cotton is Professor of Economics at Queen's University in Kingston, Ontario, where he holds the Jarislowsky-Deutsch Chair in Economic and Financial Policy and directs the John Deutsch Institute for the Study of Economic Policy. He is cross-appointed to the School of Policy Studies and the Department of Medicine. His research spans evidence-based policy, the economics of science, education and public health, and how organisations use expertise when they make decisions. He is a co-owner of and research adviser at Limestone Analytics, a policy advisory firm, and has advised on impact evaluation and evidence-based policy for the UK and US governments, the World Health Organization, and others.
Research cited in this episode
Evidence-based medicine. The movement that took hold in the 1990s, decades after the clinical trial became standard. Its classic statement is Sackett, David L., William M. C. Rosenberg, J. A. Muir Gray, R. Brian Haynes, and W. Scott Richardson. 1996. "Evidence Based Medicine: What It Is and What It Isn't." BMJ 312 (7023): 71-72. It set out how clinicians should combine the best research evidence with clinical judgement. Cotton's point is that medicine's evidence hierarchy puts systematic reviews above any single trial; the social sciences adopted the trials but not the hierarchy.
Technology Readiness Levels. NASA's nine-level scale for the maturity of a technology, from basic principles observed (TRL 1) to flight proven on a successful mission (TRL 9). It gives engineers a shared vocabulary for readiness without dictating how each component is tested. PERLs borrow the idea and apply it to claims about policy.
The four PERLs. PERL 1 is a hypothesis, backed by theory or observational data. PERL 2 is efficacy, shown in controlled trials under favourable conditions. PERL 3 is targeted adoption, with at least one rigorous causal evaluation under real-world implementation. PERL 4 is generalised guidance, built on systematic reviews that address how effects vary across contexts. Cotton's examples of PERL 4 evidence in development include insecticide-treated bed nets for malaria, cash transfers for school enrolment, and iron and folic acid supplements for pregnant women.
The science of scaling. A growing literature on why promising pilots disappoint at scale; Cotton cites Rasul, Imran. 2026. "From Field Experiments to Policy Interventions at Scale." Science 391 (6790). Much of this work asks researchers to test programmes under real-world conditions. In PERL terms, that moves evidence from level two to level three; Cotton argues it is necessary but not sufficient.
Evidence synthesis organisations. Cochrane produces systematic reviews in health; the Campbell Collaboration does the same for social policy; 3ie, the International Initiative for Impact Evaluation, funds and synthesises impact evaluations in development. Together with J-PAL's evidence reviews and WHO guidelines, Cotton identifies them as where PERL 4 evidence lives.
Microcredit. Cotton's main cautionary tale. Early pilots produced promising causal evidence under specific conditions, and donors, NGOs, and the media turned it into a claim that microcredit would transform global poverty. The later randomised evidence is summarised in Banerjee, Abhijit, Dean Karlan, and Jonathan Zinman. 2015. "Six Randomized Evaluations of Microcredit: Introduction and Further Steps." American Economic Journal: Applied Economics 7 (1): 1-21. Across six countries, the studies found some increase in business activity but no evidence of a reduction in poverty.
The graduation approach. A package of a productive asset, training, coaching, savings, and consumption support for ultra-poor households, first developed by the Bangladeshi NGO BRAC in 2002. The landmark multi-country study is Banerjee, Abhijit, Esther Duflo, Nathanael Goldberg, Dean Karlan, Robert Osei, William Parienté, Jeremy Shapiro, Bram Thuysbaert, and Christopher Udry. 2015. "A Multifaceted Program Causes Lasting Progress for the Very Poor: Evidence from Six Countries." Science 348 (6236): 1260799. J-PAL's Policy Insight now draws on 20 randomised evaluations. Cotton uses it as an example of evidence that matured as implementation expanded.
The reading wars. Cotton's example from outside development. Many education systems adopted whole-language and three-cueing approaches to reading, based on appealing theory and small, heavily resourced studies, and pushed phonics aside. The evidence is reviewed in Castles, Anne, Kathleen Rastle, and Kate Nation. 2018. "Ending the Reading Wars: Reading Acquisition from Novice to Expert." Psychological Science in the Public Interest 19 (1): 5-51.
Quasi-experimental methods. Techniques such as difference-in-differences, which compare changes over time between places or groups that did and did not receive a policy. Cotton argues that macro policies and institutional reforms, which cannot be randomised, can still climb to PERL 3 and even PERL 4 by combining such studies with well-tested theory and the historical record.
More VoxDev Talks episodes
How do policymakers interpret different types of evidence? Eva Vivalt on how policymakers update their beliefs when they see new evidence, and the biases that get in the way.
Rethinking evidence and refocusing on growth in development economics. Lant Pritchett makes the sceptic's case against relying on RCTs and systematic reviews as the main guide to policy.
How AI can put 20 years of development evidence to work. Iqbal Dhaliwal of J-PAL on building AI programmes on the existing evidence base, and the scaling problems AI inherits.
Related reading on VoxDev.org
Scaling policy ideas in developing countries, in which John List argues that researchers should generate policy-based evidence before decision-makers commit to scale.
Insights from first generation of microcredit RCTs, from the VoxDevLit on microfinance, which reviews seven randomised evaluations of microcredit programmes.
