Chapter 20: Impact Assessment
When you are done with this section, you will be able to...
Define impact assessment and explain attribution in evaluation.
Distinguish between observed participant improvement and true program impact.
Describe the strengths and limitations of quasi-experimental and experimental designs.
Identify how to select the most appropriate impact assessment design for your intervention.
INTRODUCTION
After tracking participant outcomes, organizations must address a deeper question: whether the changes observed within a targeted population can be attributed to their intervention. While outcome measurement documents what changed as a whole, impact assessment focuses on understanding why a change occurred and how much of it can be directly ascribed to an intervention. Once outside influences like economic conditions, community resources, and participants’ personal resilience are accounted for, those implementing the intervention can gain valuable insight into how their product or program is succeeding or failing.
This chapter introduces impact assessment as the next step in a comprehensive evaluation approach. It teaches you how to apply the concepts and designs needed to isolate a program’s contribution to observed outcomes and helps distinguish between general participant improvement and improvements that occur as a result of the intervention.
WHAT IS IMPACT ASSESSMENT?
Impact assessment is the process of determining whether an intervention directly caused the observed changes within a target population. These changes may be positive or negative, but the central question remains the same: did the intervention itself produce the observed outcomes, or would those changes have occurred regardless because of external influences?
This focus on attribution represents the core challenge of evaluation. While outcome measurement can demonstrate that change occurred after an intervention was implemented, impact assessment goes a step further by determining whether and to what extent the intervention was responsible for that change. Establishing this distinction is critical not only for demonstrating the value of an intervention but also for strengthening its credibility as an effective social impact solution. As a result, social problem-solving organizations (SPSOs) rely on impact assessment to generate reliable evidence that their efforts are truly driving meaningful change.
To answer questions of attribution, impact assessment builds upon the data collected during outcome measurement and uses evaluation designs capable of isolating the effects of the intervention from other contributing factors. This often requires quantitative data that measures how much change can reasonably be attributed to the program itself. Two of the most common approaches used to accomplish this are quasi-experimental designs and randomized controlled trial designs.
HOW CAN YOU DETERMINE YOUR PROGRAM’S IMPACT?
Quasi-experimental design (also referred to as comparison group design) and randomized controlled trials (RCTs) are two effective evaluation methods for clarifying impact. Both approaches aim to establish a cause-and-effect relationship between the intervention and the documented outcomes. Each demands a different level of rigor and comes with its own set of practical considerations.
1. Quasi-Experimental Design (Comparison Group)
The comparison group design strengthens causal claims by comparing your treatment group, participants who received your intervention, against a similar group that did not receive it. If the treatment group improves significantly more than the comparison group over the same period, that difference can be reasonably attributed to your intervention rather than to an external factor or the passage of time.
How It Works:
Identify two groups: a treatment group and a comparison group. For example, you might compare adults in your 12-week substance abuse detox program (treatment group) against similar adults on a waitlist for your program who haven’t yet received services (comparison group). Measure both groups prior to the intervention period and again after, observing if one group displays any significant differences.
Real World Example:
In 2018, Central City Community College in Ohio partnered with the nonprofit JobsFirst to evaluate its new 10‑week, workforce development program for unemployed adults.1 The college enrolled 200 participants into the program (treatment group) and, due to limited capacity, placed another 200 eligible applicants on a waitlist (comparison group).
Central City and JobsFirst administered a baseline survey to both groups early in the program and a follow‑up survey six months after the program ended. After the intervention was implemented, they found that 62% of program participants were in stable employment (at least 30 hours per week for 12 consecutive weeks), compared with 38% of individuals on the waitlist. This led to a difference of 24 percentage points, which the evaluation team attributed to the program, after adjusting for baseline differences in age, education, and prior work history.2
Strengths:
By measuring two groups within a similar geographic area, the JobsFirst example controls for multiple external factors, like economic shifts or the passage of time, because both affect the groups equally. The similarity of external factors between the two groups can help in identifying and attributing changes in outcome. Additionally, a quasi-experimental design is often more feasible and affordable than a full randomized trial, making it a practical choice for many social impact organizations.
Limitations:
The main limitation of the quasi-experimental design is its lack of random assignment, meaning that individuals in the treatment group may be inherently different from those in the comparison group, even before the program begins. As a result, selection bias can occur, which affects the validity of the results and makes attribution unclear. Some changes found may reflect pre-existing changes, rather than the intervention. This is true, even when external factors are taken into account.
Common sources of selection bias include: self-selection (individuals who choose to enroll might have more motivation, stronger family support, or fewer barriers than those who didn’t enroll), geographic differences (participants in one area of an implementation may differ from those in another when it comes to resources or other factors that affect outcomes), and timing differences (conditions that may have changed and affected participants such as funding, staff expertise, or economic conditions).
The Bottom Line:
Comparing outcomes from a treatment group and a comparison group can provide reasonable evidence of the impact of your intervention. However, selection bias may interfere with the validity of data. This limitation drives the need for other, more complex evaluation designs.
Imagine you’re running a job training program and want to use a comparison group design. You decide to compare the employment rates of your program graduates against people currently on your waitlist who haven’t yet received services.
Why might the people on the waitlist NOT be a perfect comparison?
What differences might exist between those who got into your program immediately and those who are still waiting?
How might these differences affect your findings?
What could you do to strengthen your comparison group design?
2. Experimental Design (Randomized Controlled Trials -RCTs)
Another methodology for establishing causal evidence is experimental design or RCTs, which uses random assignment to eliminate selection bias and provide solid evidence of an intervention’s impact. Random assignment refers to the process of selecting and allocating individuals, at random, to be part of the treatment or control group. Because assignment is random and not based on any characteristics of the participants, the two groups should be statistically identical at the start in both measured characteristics (e.g., age, income, and education) and unmeasured characteristics (e.g., motivation, family support, and resilience). In the two groups, the only systematic difference between them is whether they received your intervention. With only one difference, the outcomes at the end can be confidently attributed to your program rather than to pre-existing differences, external factors, or alternative explanations.
How It Works:
Randomly assign participants to either a treatment group that receives your intervention or a control group that does not receive your intervention. This control group may still be offered a form of standard treatment or be given a placebo. Statistical principles show that randomly assigning thirty or more participants to each group increases the likelihood of equivalent groups. Compare the before-and-after results from each group to evaluate the impact of your intervention.
Real World Example:
The Nurse-Family Partnership (NFP), a home-visiting program for first-time mothers, has used randomized controlled trials extensively to prove its impact. In a Memphis study, 743 randomly assigned pregnant women received either home visits from trained nurses during pregnancy and the first two years of the child’s life (treatment group) or were offered the standard prenatal and pediatric care available in the community (control group). Follow-up studies found that children from the treatment group had 56% fewer emergency room visits in their first two years of life, and the mothers from the treatment group had 44% fewer maternal behavioral problems related to substance abuse, and longer intervals between births. Because the groups were randomized, researchers could confidently conclude that NFP caused these improvements rather than attributing them to the mother’s inherent parenting skills or other external factors.3
Strengths:
If RCTs are implemented correctly, evaluators can reliably identify whether the program or intervention is the cause of the observed changes within a targeted population. This results from the design controlling all potential confounding variables, both measured and unmeasured, to ensure the two groups are equivalent at the start of the study. This makes it possible for RCTs to provide policymakers, funders, and the research community with credible evidence of the intervention’s impact. As a result, this design confirms whether a new intervention is successful and acts as a necessary precursor to scaling an intervention or advocating policy adoption. When resources are limited, and the intervention is unable to serve everyone in the targeted population, random assignment can also be an ethical way to allocate scarce services fairly.
Limitations:
While RCTs are often considered one of the strongest methods for evaluating effectiveness, their practical challenges, such as contamination—when individuals in the control group seek similar services elsewhere—and non-compliance—when those assigned to treatment do not participate as intended—can weaken the reliability of the findings.
Furthermore, even when an RCT produces strong results under highly controlled conditions, those outcomes may not translate equally well across different contexts or at larger scales. As a result, a program that appears effective in one setting may not achieve the same impact when implemented in more complex real-world environments. RCTs also face ethical dilemmas when the random assignment of individuals treated
The structure of an RCT also demands a significant number of resources, requiring larger sample sizes, longer timeframes, sophisticated data systems, and the use of external evaluators. These resources are expensive to obtain and can cost between $100,000 and $1 million, sometimes more. This is further exacerbated by the length of time these resources are needed. RCTs often require years to show long-term impact, thereby necessitating the need to sustain contact with both groups throughout the study period and maintain the funds to do so. This can make it difficult to evaluate an intervention via RCT if there are significant financial limitations or restricted time frames.
The Bottom Line:
Randomized controlled trials provide strong evidence of the impact of your intervention by proving the causation of evaluation results. RCTs are often the preferred form of evaluation used before scaling an intervention, advocating policy adoption, or contributing to social impact at a broader level. However, while highly beneficial, RCTs are not always feasible because of cost and time constraints.
Given the high costs, long timelines, and resource demands of RCTs, do you think they should always be used before scaling a program or advocating for policy change? Why or why not?
HOW DO YOU CHOOSE AN APPROPRIATE IMPACT ASSESSMENT DESIGN?
Choosing the appropriate impact assessment design begins with determining how certain you need to be that the intervention itself caused the observed outcomes. Different evaluation designs provide different levels of confidence in establishing causation. For example, if the primary goal is internal learning or program improvement, a less rigorous and less resource-intensive design may be sufficient to identify patterns, trends, or areas for adjustment. However, if the findings will be used to persuade funders, influence policymakers, justify large-scale investment, or support broader adoption of an intervention, stronger evidence of attribution is often required. In these situations, more rigorous designs are necessary to demonstrate that the outcomes observed were caused by the intervention rather than by external factors or coincidence.
Because impact assessment is specifically concerned with attribution, the design selected should reflect both the purpose of the evaluation and the expectations of the intended audience. Audiences that will use the findings to make significant funding, policy, or implementation decisions typically require a higher standard of evidence and greater methodological rigor. In contrast, audiences focused primarily on organizational learning or program refinement may prioritize timely and practical insights over definitive causal proof. As a result, selecting an impact assessment design involves balancing the level of evidence needed with the practical realities of time, cost, capacity, and context.
HOW DO REAL SPSOS EVALUATE THEIR IMPACT?
Evaluation in Action: The Housing First Example
Let’s apply these principles to a real-world scenario addressing homelessness in Utah. This example will walk you through how to design an impact assessment from start to finish.
- Social Issue: Chronic homelessness among families in Salt Lake City, Utah.
- Intervention: Housing First Program, a model that provides permanent housing immediately to families experiencing homelessness without preconditions like sobriety, employment, or treatment compliance, combined with voluntary supportive services.
- Evaluation Question: Does the Housing First program increase housing stability and improve economic outcomes for participating families?
Housing First’s Evaluation Design Matrix
| Design Component | Evaluation Plan Details | Rationale |
| WHO is studied? (The Groups) |
| The waitlist provides a "counterfactual"—it shows what happens to similar families who don’t receive the program. Because both groups applied and were deemed eligible, they’re more comparable than if we compared the participants to families who never applied (who might be less motivated or have different needs). |
| WHEN is data collected? (The Timing) |
| We measure what matters to families and policymakers. These indicators align with our program goals (housing stability and economic improvement) and are measurable, specific, and relevant to proving our program works. |
| WHAT is measured? (The Indicators) | Primary Outcomes:
Secondary Outcomes:
Outputs:
| We measure what matters to families and policymakers. These indicators align with our program goals and are measurable, specific, and relevant to proving our program works. |
| HOW is impact proven? (Attribution) | Difference-in-Differences Analysis: We compare the change in housing stability of the treatment group against the change in the comparison group over the same time period. This approach filters out external factors (like improvements in the overall economy or job market) that would affect both groups equally, isolating the program’s specific contribution. | Simply comparing final outcomes isn’t enough because the groups might have started at different levels. By comparing how much each group changed, we account for baseline differences and external trends, providing stronger evidence that our program caused the improvements. |
SUMMARY
Impact assessment matters because social impact work requires more than good intentions. It requires credible evidence that your efforts are truly making a difference. Without understanding whether your intervention is actually causing positive change, it becomes difficult to know whether resources are being used effectively, whether programs should be expanded, or whether strategies need to be revised. Strong impact assessment helps organizations move beyond assumptions and make decisions grounded in evidence rather than perception alone.
This is especially important when decisions affect funding, policy, scaling efforts, and, ultimately, the lives of the individuals and communities being served. By strengthening your ability to determine what is genuinely effective, impact assessment supports more responsible stewardship of resources, improves organizational learning, and increases the likelihood that successful interventions can be adapted and sustained over time.
Understanding evaluation designs such as quasi-experimental and randomized controlled trial methods also equips you to make more informed methodological decisions. No single design is appropriate for every context. Instead, effective evaluation requires balancing rigor with feasibility, taking into account the goals of the assessment, the available resources, ethical considerations, and the practical realities of implementation. Selecting the right design strengthens the credibility of your findings and helps ensure that the conclusions drawn are both meaningful and actionable.
ENDNOTES:
1 - Schochet, Peter Z. 2013. Designing and Conducting Strong Quasi-Experiments for Education Research. Washington, DC: National Center for Education Evaluation and Regional Assistance, U.S. Department of Education.
2 - Brown, C. H., et al. 2019. “Experimental and Quasi-Experimental Designs in Implementation Research.” Psychiatry Research volume(issue): page range.
3 - Chien, A., et al. (2020). “Protocol for a randomized controlled trial evaluating the impact of the Nurse-Family Partnership’s home visiting program in South Carolina on maternal and child health outcomes.” Evaluation and Program Planning.
Resources published and shared by the Ballard Center are not necessarily endorsed by BYU or The Church of Jesus Christ of Latter-day Saints.