Nyumbani/Hifadhi/Do Community Programs Work? How to Evaluate Without Oversimplifying
Makala Evaluation FAQ

Do Community Programs Work? How to Evaluate Without Oversimplifying

Evaluation is most useful when it explains what changed, why it may have changed, and what the program should learn next.

Dr. Petar Ilić/ 28 Juni 2026 /9 dakika kusoma /Programme Evaluation
Do Community Programs Work? How to Evaluate Without Oversimplifying

“Just measure impact” sounds tidy until a small nonprofit asks what that actually means on Monday morning. A neighbourhood mentoring project, a migrant-led advice group, or a youth workshop rarely has the budget, sample size, or control over people’s lives that would make a clean laboratory answer possible. That does not make evaluation fake. It means the evaluation question has to fit the program.

This FAQ is for teams that need evidence they can use. It follows the Institute’s usual sequence: what is known, what is contested, and what we cannot say.

What does it mean to evaluate a social project?

To evaluate a social project is to ask, in a structured way, whether the work happened as intended, whether anything changed for participants or systems, and how plausible it is that the program contributed to that change.

That is broader than asking whether the program “worked.” A project can be well run but too small to shift the outcome it cares about. Another can show promising participant feedback while still excluding people who most needed it. Evaluation should make those distinctions visible.

For a small nonprofit, a useful evaluation often has three layers:

  • Implementation: Did we deliver what we said we would deliver?
  • Experience and quality: Did participants find it accessible, respectful, safe, and relevant?
  • Outcomes: What changed in knowledge, behaviour, confidence, access, relationships, or institutional response?

The German phrase “soziales Projekt” can cover everything from a volunteer initiative to a publicly funded service. The evaluation does not need the same scale in every case. It needs a clear question, credible evidence for that question, and honest language about limits.

What is a theory of change, and why does it matter?

A theory of change is the program’s explanation of how activities are expected to lead to outcomes. It is not a decorative diagram for funders. It is the set of assumptions you are testing.

For example: if a tenants’ information session is meant to reduce panic after a rent increase, the theory might be that people need understandable information, time to ask questions, and referral routes before they can act. If a youth civic program is meant to increase participation, the theory might be that repeated low-pressure contact builds trust, and trust makes public speaking or volunteering more likely.

The value is not that the theory is always right. The value is that it can be inspected. Which step is weak? Did people attend but not understand? Did they understand but still face barriers? Did the program improve confidence but not change the institutional response?

A good theory of change includes context. In Germany, a program may depend on local housing pressure, Jobcenter practice, school cooperation, transport access, language availability, or whether people trust the organizing institution. Evaluation that ignores those conditions may blame the program for problems it cannot solve, or credit it for changes caused elsewhere.

What evidence should a small nonprofit collect?

Collect the least amount of evidence that can answer the real question responsibly. More data is not automatically better. It can burden participants, create privacy risks under GDPR, and leave teams with spreadsheets they never analyze.

Useful evidence for small programs often includes:

  • attendance and reach, including who was not reached where this can be assessed ethically;
  • simple pre/post questions on knowledge, confidence, or intended next steps;
  • short feedback on accessibility, safety, relevance, and respect;
  • staff or volunteer delivery notes, recorded soon after sessions;
  • referral or follow-up information, if participants consent and data protection is clear;
  • qualitative interviews or group reflections with participants;
  • observations of whether the process matched the design;
  • records of institutional changes, such as new meeting routines or revised materials.

The strongest design depends on the claim. If the claim is “participants valued the workshop,” feedback and interviews may be enough. If the claim is “the workshop increased knowledge,” a simple before-and-after measure helps. If the claim is “the program reduced school absence,” you need a much more careful design because many other factors can affect attendance.

Avoid collecting sensitive categories just because they look sophisticated. If migration background, disability, income, or racialization is relevant, explain why, ask in a respectful way, and make clear how the information will be protected and used.

How do we measure outcomes without flattening people?

Outcomes are changes that matter. They can be individual, relational, organizational, or institutional.

Individual outcomes might include knowledge of rights, reduced confusion, increased confidence to contact an advice service, or a completed application. Relational outcomes might include stronger peer networks or improved trust between a community group and a school. Institutional outcomes might include a changed appointment process, a new interpretation policy, or a public agency responding in writing instead of informally.

The contested part is how to measure outcomes that are real but not easily counted. Confidence, belonging, trust, and dignity matter, but one tick-box can miss their shape. That is why mixed evidence is often stronger: a short scale can show direction, while interviews can explain what the change meant and where it did not happen.

Be careful with proxy outcomes. Attendance is not impact. Satisfaction is not learning. A referral is not the same as access. A social media share is not policy change. These measures can still be useful, but they should be named as signals, not final proof.

Do we need a comparison group?

Sometimes. A comparison group helps answer whether change was likely connected to the program rather than to something else. In social research, this is part of causal inference: the attempt to understand cause and effect, not only correlation.

For many nonprofits, a randomized trial is unrealistic or inappropriate. But comparison thinking is still possible. You might compare people before and after the program. You might compare similar groups in different locations. You might compare participants who used different parts of a service. You might compare your findings with stable external data, such as broader survey patterns from sources like SOEP when the question is suitable.

Each option has limits. Before-and-after designs cannot rule out outside events. Comparing participants with non-participants can be distorted by selection: people who join may already be more motivated, less isolated, or more available. Administrative data may not capture the outcome you care about.

The honest sentence is often: “The evidence is consistent with the program contributing to this change, but it cannot prove the program was the only cause.” That is not weak. It is accurate.

What is process evaluation?

Process evaluation asks whether the program operated as planned and how people experienced it. It is especially important when outcomes are delayed, hard to measure, or dependent on other institutions.

Process questions include:

  • Did the intended group know about the offer?
  • Was the venue reachable by public transport and accessible to disabled participants?
  • Were language needs met?
  • Did staff follow the safeguarding or confidentiality plan?
  • Were sessions delivered with enough time for questions?
  • Did referral partners actually have capacity?
  • Did participants drop out at a particular point?

Process evidence prevents a common error: judging a theory before checking whether it was implemented. If a program did not reach shift workers because all sessions were at 15:00, the finding is not simply “the program failed.” The more useful finding is that the delivery model did not match the audience’s schedule.

How should we handle unintended effects?

Every evaluation should ask what else happened, including harms. Programs can create waiting lists, raise expectations that cannot be met, expose participants to stigma, overload volunteers, or shift unpaid work onto community members. A well-liked project can still have unequal effects.

This is where qualitative methods are valuable. Open questions, interviews, and staff reflection can reveal effects that were not in the original indicators. A participant may say the workshop helped them understand their rights but also made them realize how little support was available. A volunteer team may report that informal follow-up became emotionally heavier than expected.

Unintended effects are not automatically reasons to stop. They are reasons to redesign. The evaluation should ask whether the program has the capacity, partnerships, and boundaries its own model requires.

How can we learn without pretending the evaluation is perfect?

Use a claim ladder. At the bottom are descriptive claims: “We delivered six workshops and reached this audience.” Next are experience claims: “Participants reported that the format was understandable and respectful.” Then come outcome claims: “Participants showed improved knowledge on these questions.” Higher still are contribution claims: “The pattern suggests the program contributed to earlier advice-seeking.” At the top are causal claims: “The program caused this outcome.”

Most small nonprofit evaluations belong in the middle of the ladder. That is not a failure. It is often enough for learning, adaptation, and responsible reporting.

Learning requires deciding in advance how findings will be used. What result would make you change the format? What would make you stop an activity? What would justify expansion? What would require a partner, rather than more workshops?

Evaluation should leave a team with better questions, not only better sentences for a report. The most useful final page may say: keep this part, change this part, stop claiming this part, and investigate this uncertainty before scaling.

What can evaluation not tell us?

Evaluation cannot remove judgment. It cannot tell a community what values to hold, whether a funder should prioritize one need over another, or how much uncertainty is politically acceptable. It also cannot prove that a complex social change came from one project when many forces moved at once.

What it can do is narrow the space for wishful thinking. It can show whether the theory of change is plausible, whether the program reached the people it intended to reach, whether participants experienced it as useful, what changed, what did not, and what should be claimed with caution.

For campaign measurement, Civic Futures Lab goes deeper into leading indicators before a policy win. Here, the evaluation task is simpler and harder: describe the program honestly enough that the next decision is better than the last one.

More from the archive