Nyumbani/Hifadhi/When Data Collection Harms the People It Is Meant to Help
Makala Ethics casebook

When Data Collection Harms the People It Is Meant to Help

Data collection can create harm even when the cause is good; ethical research asks who carries the risk and who controls the result.

Dr. Amina Sow/ 28 Juni 2026 /9 dakika kusoma /Social Research
When Data Collection Harms the People It Is Meant to Help

The first warning sign is often gratitude. A team thanks people for “sharing their voices,” stores the recordings, writes a report, and moves on. The participants have done the emotional work. The organization has gained material. Nothing visibly changes in the community that carried the story.

The cases below are composite. Details have been changed and combined to avoid identifying any person or organization. They are not scandals to solve after the fact. They are patterns to recognize before a survey, interview project, monitoring form, or evaluation begins.

Case 1: The same painful story, collected again

A local advice project wanted to understand barriers faced by people who had been treated unfairly in public offices. The intention was serious. Staff had seen enough repeated patterns to know that anecdotes alone would be dismissed. They planned interviews, hoping that detailed accounts would make the issue harder to ignore.

But many potential participants had already told their stories several times: to a volunteer, to a lawyer, to a counsellor, to a public office, to another research team, and sometimes to journalists. Each retelling required chronology, names, documents, translation, and the effort of staying composed. Some people agreed because they wanted the problem to be taken seriously. Others felt that refusing would make them seem ungrateful for support.

The harm did not come from one cruel question. It came from repetition without a clear benefit to the people doing the repeating.

What is known: recounting distressing experiences can be burdensome. Some people find testimony meaningful, especially when it is voluntary, well-supported, and connected to action. Others experience the request as extraction, particularly when they have no control over how the story is later used.

What is contested: research teams disagree about when detailed narratives are necessary. Some argue that institutions ignore abstract categories unless they hear human detail. Others argue that repeated requests for pain reproduce the same unequal relationship: affected people must become evidence before systems respond.

What we cannot say: we cannot assume that a story is healing because it is spoken. We also cannot assume that silence means there is no harm. People may decline because the price of participation is too high.

An ethical redesign would start by checking what evidence already exists. Could prior anonymized case notes answer part of the question? Could participants choose a shorter format? Could the team ask about systems and consequences rather than requiring full personal histories? Could support be available if an interview opens distress? MindForward Collective is the better sibling hand-off for aftercare, emotional safety, and distress planning. A research design should not improvise that support at the moment someone becomes overwhelmed.

Case 2: A demographic table that makes people findable

In another composite case, a small organization collected detailed demographic data to show that a public service was failing several groups at once. The form asked for age, gender, disability, nationality, language, residence status, neighborhood, household type, and the exact office visited. The team planned to publish only tables, not names.

On paper, this looked responsible. In practice, some combinations pointed to one or two people. In a small town, “single parent, wheelchair user, recent arrival from a specific country, using one local office” was not anonymous. Even if the organization never intended harm, a leaked spreadsheet, an over-detailed chart, or a careless quote could expose someone.

Unsafe demographic data can harm people through recognition, stigma, and future misuse. This is especially serious where data concerns migration status, health, disability, religion, gender identity, experiences of violence, debt, housing insecurity, or conflict with public authorities.

What is known: removing names is not the same as anonymizing data. Rare combinations can identify people. Small subgroup tables can reveal more than a team realizes. Open-text responses can contain names, dates, locations, and details that undo anonymity.

What is contested: communities and researchers may disagree about whether to collect sensitive categories at all. Without such data, unequal effects can disappear in averages. With such data, people may face risk. The answer is not a universal ban. It is a proportionality test: collect only what is needed, protect it carefully, and decide in advance what will never be published.

What we cannot say: we cannot promise zero risk once sensitive data exists. We can reduce risk, restrict access, aggregate results, delete unnecessary details, and avoid creating files that no one can safely manage.

Digital Dignity Lab is the appropriate hand-off for detailed data safety practices. At the research-design level, the rule is: if a variable will not be used for a clear analysis purpose, do not collect it. If it will be used, decide whether broader categories would answer the question with less risk.

Case 3: Categories that turn people into problems

A youth-focused project wanted to understand why some young people were not attending activities. Its form asked whether respondents came from “problem families,” whether they were “integration resistant,” and whether their parents “valued education.” The staff did not intend insult. They were using phrases common in local administrative talk.

Participants read the categories differently. The survey seemed to have already decided what was wrong with them. Some stopped responding. Others selected answers defensively. A few wrote angry comments. The dataset became less reliable because the categories were stigmatizing.

Stigmatizing categories harm in two ways. They can wound directly, and they can produce bad evidence. If people feel judged, they may avoid the survey, choose socially acceptable answers, or refuse future cooperation.

What is known: categories shape what can be seen. A survey that asks only about individual deficits will find individual deficits. An interview guide that never asks about school rules, transport, racism, language access, disability access, care work, or money will miss structural barriers.

What is contested: some policy systems use categories that communities dislike but cannot avoid. A project may need to ask about benefit receipt, residence status, school track, or disability recognition because those categories affect access to rights and services. The ethical task is to explain why the category is being used and to avoid turning administrative labels into identities.

What we cannot say: we cannot infer a person’s values, motivation, or culture from non-participation. A missed appointment, unanswered email, or unfinished form may reflect shift work, fear, translation needs, inaccessible design, unstable housing, childcare, or previous bad experiences with institutions.

A better survey would ask about conditions: timing, transport, language, cost, safety, discrimination, digital access, disability access, trust, and whether the activity felt relevant. It would leave room for people to reject the premise.

Consent can fail even when a form has a checkbox.

It fails when people believe services depend on participation. It fails when the project uses academic or legal language that participants cannot reasonably understand. It fails when people are told that data is anonymous but the form collects identifying details. It fails when a person agrees to an interview but not to having a quote used in a public report, and the team treats those as the same permission.

In a German and EU context, GDPR gives a general legal frame for personal data, but ethical consent is broader than compliance. The ethical question is whether people understand the choice, can refuse without penalty, and know what will happen next.

Good consent is specific. “You agree to take part in this interview” is not the same as “you agree that anonymized quotes may appear in a public report.” “We will use this for research” is not the same as “we may share aggregated findings with a city committee, funder, or media partner.” A person can agree to one use and refuse another.

Community ownership is not a thank-you paragraph

Extractive research often has a recognizable path: outsiders define the question, collect local knowledge, interpret it elsewhere, publish it in their own language, and return with a summary after the important decisions have been made.

Community ownership asks different questions:

  • Who helped define the research question?
  • Who reviewed the categories and wording?
  • Who can see raw or summarized data?
  • Who decides which findings are too sensitive to publish?
  • Who benefits from the report?
  • Who can challenge the interpretation?
  • What remains in the community after the project ends?

Ownership does not mean every participant must vote on every sentence. It means the people most affected by the research are not reduced to data sources. Advisory groups, paid community reviewers, shared interpretation workshops, accessible summaries, and agreements about publication limits can all shift control.

A practical harm test before collecting data

Before launching a project, ask:

  1. What decision will this data inform?
  2. What evidence already exists?
  3. What is the smallest amount of new data needed?
  4. Which questions could expose, stigmatize, or distress someone?
  5. Can people refuse without losing services, goodwill, or future access?
  6. Who will hold the data, and for how long?
  7. What will not be published even if it is interesting?
  8. How will participants learn what came from their contribution?
  9. Who can say that the analysis is wrong or incomplete?

These questions answer “Wann kann Forschung Schaden verursachen?” Research can harm when it extracts stories without benefit, repeats trauma, exposes identities, hardens stigma, confuses consent, or removes control from the people whose lives are being described.

Avoiding harm does not mean avoiding evidence. It means collecting evidence with discipline. Strong research asks not only whether a question can be answered, but whether it should be asked in that form, by that team, at that time, with those protections.

More from the archive