PERL-Framework for Evidence-Based Policy
Policy Evidence Readiness Levels (PERLs) classify how mature the evidence behind a claim is, and what action that maturity supports. They give funders, evaluators, analysts, and evidence synthesists a common language for matching decisions, and the claims used to justify them, to the strength of the underlying evidence. A single decision usually rests on several claims, and each can be assessed separately.
Addressing an evidence gap
While the policymakers and practitioners increasingly recognizes the benefits of rigorous causal impact evaluations, misperceptions remain about how to translate evaluation results into action, program design, and policy. Many fields still lacks a shared way to express how mature the evidence behind an intervention actually is, and what that maturity justifies. A single well-identified study and a synthesized body of science accumulated over many studies represent very different levels of evidence maturity and should support very different actions. Yet they are routinely treated as similarly credible proof of effectiveness in funding decisions, scaling choices, and appeals to “follow the science.”
The consequences of treating preliminary evidence as conclusive are substantial. Interventions get scaled, and millions of dollars invested, on promising early results that later fail to replicate. Evidence-backed claims are produced without signaling whether they rest on an evolving or a mature evidence base, leaving their weight for decisions unclear. And caution is too easily framed as anti-evidence or skeptical of science, when such skepticism is often justified given what the evidence actually shows.
Policy Evidence Readiness Levels provide the missing signal. Modeled on NASA’s Technology Readiness Levels, which classify whether a technology is ready for deployment, PERLs classify whether an evidence base is ready to support experimentation, targeted adoption, or broad scaling. They are a technical assessment of evidence maturity. Whether a policy is desirable, affordable, or politically viable is a separate question. PERLs are also no barrier to action while evidence is still evolving; acting on early-stage evidence is often necessary. What they require is that the claims and commitments made on that evidence match what it can actually support.
PERLs at a glance
Evidence matures in stages. A decision-maker can act at any stage, but what the evidence can honestly be claimed to show, and how far it justifies committing, changes as it matures.

The PERL staircase: what the evidence supports, and what “following the science” means, at each level of maturity. Adapted from Cotton (2026), Science.
The levels in detail
| Level | What the evidence shows | What following the science means |
|---|---|---|
| PERL 1: Preliminary Insight (“It might work”) | A theoretical or observational basis for expecting an intervention may be effective, but no rigorous evidence of impact. | Support research design and piloting rather than implementation. Present findings as hypotheses to be tested, and never as evidence that the intervention works. |
| PERL 2: Emerging evidence (“It can work”) | Rigorous evidence of potential effectiveness, from a controlled trial under specific or favorable conditions, or comprehensive field research documenting impact without establishing causation. Not yet tested causally under representative implementation conditions. | Fund effectiveness trials under representative conditions. Present results as promising but preliminary: evidence that the intervention can work, which falls short of evidence that it will work across contexts or at scale. |
| PERL 3: Evolving evidence (“It works for some”) | Robust evidence of real-world impact, including at least one rigorous causal evaluation under representative delivery conditions, but not yet mature enough to support general guidance on cross-context robustness. | Adopt in contexts matching the successful studies; scale into new contexts incrementally and adaptively. Frame recommendations as conditional rather than universal, noting where evidence gaps remain. |
| PERL 4: Mature evidence (“It works, broadly”) | A systematically synthesized evidence base, with rigorous reviews and meta-analyses that address heterogeneity and cross-context effectiveness, supporting generalized guidance, including which variations work best and for whom. | Act on the synthesized guidance and scale broadly. Continue to monitor fidelity and revisit the evidence base, as effectiveness can change as conditions evolve. |
At every level, the danger is the same: acting as if the evidence were more mature than it is. PERLs do not tell decision-makers to wait. They align claims and commitments with what the evidence can actually support, and shift the burden of justification onto those arguing to implement before the evidence has matured.
Apply PERLs to each claim underlying a decision
Science-to-policy translation involves two stages. First, existing research leads to concrete claims about an intervention’s likely costs and benefits. Second, decision-makers weigh those impacts to decide what to do. PERLs sit between the two. They rate the maturity of the evidence behind each claim. They do not prescribe which methods should generate the claims (meta-analysis, structured weight-of-evidence review, expert elicitation, or anything else), and they do not say how the claims should be weighed against one another once produced. Both of those choices matter, and PERLs complement whichever approach is used.
This matters because most policy decisions rest on many claims at once. A cost-benefit analysis or feasibility study is built from a stack of “studies show” statements: the main benefit, secondary benefits, costs, and spillovers. Each of these is its own claim, with its own evidence base, at its own level of maturity. A program’s main benefit may rest on replicated evidence across populations (PERL 4), its spillovers on theory alone (PERL 1), and its costs and other benefits on evidence from idealized or distant implementations (PERL 2 or 3).
Assigning a separate readiness level to each claim makes the uncertainty behind each input transparent. It shows where the case for a program is strongest and where it is most likely to break down. That, in turn, tells analysts where sensitivity analysis is most needed, tells funders where monitoring and evaluation should concentrate, and tells decision-makers which parts of a promising case to treat with confidence and which to treat as provisional.
Application: delivery claims
Public health researchers have long recognized that a program’s effect in a trial is only one of the things that determines its impact in practice. The RE-AIM framework (Glasgow, Vogt, and Boles, 1999) breaks real-world impact into five components: reach (who receives the program), effectiveness (what it does for them), adoption (which organizations take it up), implementation (whether it is delivered as intended), and maintenance (whether effects and delivery persist over time). Research is abundant on effectiveness and far thinner on the other four.
PERLs give this pattern a language. Each of these components is a claim, and each claim’s evidence can be rated. Reach, adoption, and implementation describe how a program is delivered, so each gets one readiness level per program. Effectiveness is different: most social policy, development, and health interventions target several outcomes, and the evidence behind each outcome can sit at a different level, so each gets its own effectiveness PERL. Maintenance splits, with institutionalization rated once at the program level and sustained behaviour change rated per outcome. A program with several intended impacts will therefore carry more than five readiness assessments, and the pattern across them is often more informative than any single one.
Putting PERLs in practice
PERLs are designed to fit existing workflows rather than replace them. They function as a meta-standard. Like NASA’s Technology Readiness Levels, they provide a shared scale without prescribing how any single study should be conducted or assessed.
- Funders and program officers can require a PERL classification in proposals and evaluations, reserving broad implementation funding for interventions supported by mature, synthesized evidence while directing innovation funds toward testing and incremental scaling of less-ready candidates.
- Evaluators can report a PERL classification alongside findings, making explicit what the evidence supports (and does not yet support) for scaling.
- Evidence synthesists can state the PERL of the evidence base a review assesses. A systematic review of an evolving evidence base and one of a mature base can support very different actions; a PERL classification makes that distinction legible. This becomes increasingly important as AI makes on-demand evidence synthesis routine.
- Policymakers, journalists, and advocates can evaluate claims about what “the science” supports against a common, contestable standard.
Assessing evidence
The same logic can be applied quickly to any claim that an intervention “works.” Start by identifying which claim you are assessing. A claim that a program raises test scores, a claim that families will enrol, and a claim that it costs a given amount per child are three different claims, and each gets its own answers to the four questions below.
- Is it one study, or a body of research? A single result, however rigorous, is a starting point; confidence comes from findings that repeat.
- Was it tested under representative conditions, or ideal ones? Effects established under favorable conditions or by expert teams often shrink in ordinary delivery.
- Does it hold beyond where it was first shown to work? Success in one setting or population does not guarantee it transfers to others.
- Has the whole evidence base been independently synthesized? A systematic review weighing all the studies together — not a single team’s read of its own work — is the mark of mature evidence.

A reference card for assessing how mature the evidence behind a claim is. Adapted from Cotton (2026), Science.
Why I developed PERLs
In addition to my academic research, I also have substantial experience advising governments, firms, and international organizations on evidence-based policy, cost-benefit analyses, feasibility studies, and impact evaluations. In that work, the same problem recurs. Arguments in favor of a proposal are often made with confidence, backed by claims about “the science” and what “studies show.” Yet the actual evidence supporting them varies enormously: a theory, a correlation in a large dataset, a small pilot, or a widely replicated result across a well-established literature. Feasibility studies and decision processes often treat these as equally conclusive, leading to overconfidence that a program or policy “will work.” PERLs help us define and acknowledge how strong the evidence behind our assumptions is, while recognizing that a methodologically strong study is not the same as a generalizable result. They are intended to improve transparency throughout the research-to-policy pipeline.
Reference and reuse
PERLs are introduced in a peer-reviewed article in Science. The framework and the figures above are free to use with attribution, including in evaluations, funding and procurement guidelines, systematic review protocols, teaching, and reporting.
Cite as: C. Cotton (2026). Rigor is not readiness: The PERL framework for evidence-based policy, Science 393 (6810): 463–5. doi:10.1126/science.aef3529
Attached to the online article at Science is an eLetter discussion, including my correspondence:
C. Cotton (2026). Policy Evidence Readiness Levels sit between evidence and decision, Science, eLetter. doi:10.1126/science.aef3529
Read the original article and the eLetters: doi.org/10.1126/science.aef3529. Please credit figures as “Adapted from Cotton (2026), Science.” If your organization is considering incorporating PERLs into evaluation, funding, or evidence-synthesis guidelines, please get in touch.
In the news:
Global News live radio interview
VoxDev Talk podcast – coming soon.