The demand for impact measurement has never been higher. The credibility of how it is done has never been more questioned. This is the dilemma the sector is not quite ready to confront.
The Impact Monitoring Dilemma
Here is a scenario most people in development finance will recognise. A fund needs to show its investors that it is “doing good”. It commissions a survey of its borrowers or beneficiaries. A report comes back with scores, rankings, and a few compelling quotes. Everyone is satisfied. The money keeps flowing.
It sounds reasonable. And it is, up to a point. The problem is what happens when you look more closely at how those scores are produced.
A recent article by Mila Agius, published on the Dark Side of Development Substack, does exactly that. She targets a commercial firm – 60 Decibels – that sells social impact measurement services to funds, donors, and microfinance institutions. The article – “False Precision: How Business Consultants Monetise the Measurement of Poverty” – is blunt and worth reading in full. But the questions it raises go well beyond one company.
What 60 Decibels Does
The model is straightforward. 60 Decibels surveys clients of microfinance organisations and development programmes using standardised phone interviews. Fifteen minutes per person. The same questions, across dozens of countries and 58 languages. Answers get aggregated into an Index Score. That score then influences where investment capital goes.
The company describes itself as the world’s leading customer insights company for social impact. It has worked with major donors and development finance institutions. Its reports are professionally presented, full of charts and regional benchmarks.
Agius’s argument is that this appears to be science but isn’t. She calls it scientism: the imitation of scientific method without its actual rigour.
The Methodological Problem – the Impact Monitoring Dilemma
The core methodology, called Lean Data, was introduced in the Stanford Social Innovation Review in 2015. That is a practitioner magazine, not an academic journal. Agius searched for independent peer-reviewed validation of the method and found none in the public record.
This matters. Real scientific methods get stress-tested by other researchers. They get replicated. They get challenged. If they hold up, confidence grows. If they don’t, they get revised or abandoned. Lean Data has not been through that process.
The firm’s own reports acknowledge the limitations. Without control groups, causation cannot be proven. Self-reported data carries inherent bias. Question interpretation varies across cultures. These caveats are there, on page 12. The marketing materials are not written that way.
There is also a framing problem. The researcher Gerd Gigerenzer has repeatedly shown that how a question is worded determines the answer you get. A small design choice can produce dramatically different results across populations. A standardised instrument designed centrally and translated into dozens of languages, without deep local expertise, is especially vulnerable to this effect.
The Structural Conflict – the Impact Monitoring Dilemma
Then there is the business model. A microfinance organisation pays 60 Decibels to assess its own clients. Good ratings help attract investment. Bad ratings can simply go unpublished, because participation in the index is voluntary, and results do not have to be disclosed.
The firm acknowledges this in its own reports: the sample may skew towards better-performing institutions that are confident in a positive outcome. Which means the industry picture the data presents may not reflect reality at all.
The funding chain adds another layer. Some of the same donor organisations that finance microfinance institutions also provided grants to develop the Lean Data methodology in the first place. Donors need evidence of impact to justify their budgets. 60 Decibels provides that evidence, packaged attractively. Everyone in the chain benefits. The people whose phone answers generated the data do not have a seat at the table.
The article was widely shared and discussed among practitioners. That alone says something. The discomfort it triggered is not the discomfort of being attacked from the outside. It is the discomfort of recognition.
My Comment
I want to be clear about something. This is not really about 60 Decibels. They are one supplier responding to a real demand.
The impact and development industry, meaning the client, consistently asks for assessments that are fast, affordable, and easy to communicate. Someone will always provide that.
And here, I want to be honest about the other side of the argument, because Agius’s critique, if taken too literally, leads to an uncomfortable place.
We cannot run a full randomised controlled trial on every impact programme in Uganda or every solar energy project in Bangladesh. I have been there.
Academic-grade impact evaluation is slow, expensive, and often produces results long after the decisions have already been made. Monitoring and evaluation cannot cost as much as the project itself. It cannot last as long as the project either. The sector genuinely needs efficient, practical tools. We must validate the models, improve them, take responsibility for execution, and be honest in assessing results.
So I am not dismissing the idea of measurement. Quite the opposite. Measuring impact matters. It is the only honest way to know whether the money is going where it should. The question is how we do it without fooling ourselves.
Part of the answer is flexibility. A fixed, standardised model applied uniformly across contexts carries its own risks.
There is a name for it: the man with a hammer sees every problem as a nail. If you measure everything through the same instrument, you will find what the instrument is designed to find. You will miss what it was never designed to see.
Good impact monitoring needs to adapt to what it is measuring. A programme supporting women’s financial inclusion in rural Pakistan requires different questions than one financing SMEs in Nairobi. The populations are different. The power dynamics are different. The definition of a meaningful outcome is different. A single standardised phone survey cannot carry all of that weight, and pretending otherwise is where the trouble starts.
The stakes are higher than they might appear.
In development finance, impact is not a footnote to the financial return. It is the reason the transaction exists at all. Donors accept concessional terms because the impact justifies the subsidy.
Commercial investors accept below-market returns because the impact justifies the trade-off. If the impact measurement is not credible, both positions become indefensible. A positive financial return without demonstrated impact raises the question of why public or philanthropic money was involved in the first place. A negative financial return without demonstrated impact is simply a loss.
The entire logic of blended finance rests on impact being real, measurable, and honestly reported. That is what makes the credibility gap so consequential.
My concern, therefore, is not that 60 Decibels exists. My concern is the credibility gap. A product designed around speed and commercial viability cannot be presented to taxpayers and institutional stakeholders as an independent scientific assessment. That framing is where the trust breaks down.
Closing that gap requires fixing both sides. The supply side needs more transparency: full disclosure of question wording and honest presentation of limitations front and centre, rather than in a footnote.
But the demand side needs to change too. Clients need to accept results they cannot control, commission work that might reflect badly on their portfolio, and stop rewarding the most reassuring report rather than the most rigorous one.
That second part is much harder. It is an institutional problem, not a methodological one.
There is a quote from Peter Bernstein’s Against the Gods that I keep coming back to when I think about this:
“The Commanding General is well aware that the forecasts are no good. However, he needs them for planning purposes.” — Peter L. Bernstein, Against the Gods: The Remarkable Story of Risk
Everyone knows the numbers are imperfect. But a number, however constructed, provides the appearance of accountability. It gives donors something to show. It gives funds something to report. And so the demand for it persists, regardless of what it actually measures.
The result is a market for credible impact measurement that is not consistently producing credible impact measurement.
The solution is not to abandon measurement. It is to stop selling approximations as certainties.
Keep it real. Sweat Your Assets.
Thanks for reading. Browse the full article archive, catch past episodes on Spotify and Apple Podcasts, watch more on YouTube, and if you want it delivered straight to your inbox, the monthly newsletter is one click away: subscribe here.