Cumulative Evaluation for Impact

In this blog, Prof Paul Montgomery, from the Third Sector Research Centre at the University of Birmingham, and associate of 3DI, writes about how evaluation is a cumulative process, each stage building on what has gone before.

Paul’s work over the past three decades has involved over 50 systematic reviews, more than a dozen randomised trials, and a range of other studies including qualitative studies, case-controlled studies, and other forms of evaluations. He is probably best known for leading the CONSORT-SPI and GRADE-CI reporting instruments, which are the gold standard ways in which RCTs and the certainty of evidence are assessed in systematic reviews of complex interventions respectively.

Evidence is not a unitary thing. As Louisa Mitchell pointed out talking about All Child in last month’s blog, the type of evaluation and evidence needed hinges very much on the question being asked.

A good parallel can be found in law. A criminal trial, where the penalties can be severe, requires an evidence standard of ‘beyond reasonable doubt’, whereas in civil trials, the standard is ‘on the balance of the probabilities’.

It is worth considering the research question very carefully — in fact, in great detail — when thinking about the research methods you may want to use. The findings you end up with stand or fall in large part on the decisions made at this point. What is it that you really want to know?

Here, I will try to explain how using a linear process known as the ORBIT model can be helpful for many behavioural interventions which are common in the charity sector.

Cumulative Evaluation for Impact 1

In the diagram above, the ORBIT model aims to show how having a question that is important — perhaps because your charity has noticed a problem is increasing in severity — leads into the first phase.

At this point, a process of definition and refining should start. What are the components of the intervention you are providing? What is known from previous studies about it? Might you be able to collaborate with a local university to get them to review existing evidence for you about it to help you learn from previous mistakes?

Could you hold some focus groups and/or interviews to see what your population of service users think about the intervention? What matters to them? How would they know (first) if things went well or badly? Are there particular groups who don’t benefit from it?

In sum, Phase 1 is mostly desk-based review work and then qualitative evidence to find out about the components of the intervention and about the population involved. Don’t forget that in thinking about the components, you might want to know about how much of the intervention people need (dose) — could they use less or more? Might that help you think about a comparison group later?

This work sets you up for the testing stage in Phase 2.

At this point, what this model proposes is that before proceeding, we need to know whether a trial is possible, and whether there are serious risks to consider. To do that, we just need to focus on the people getting the intervention — nobody else. Don’t worry about comparison groups for now.

We need to be sure that a study is possible, that levels of possible harms we may not have considered have been tested out, and that the measures we thought about earlier will actually work.

So, in the preliminary testing phase, you might run short tests to see whether people will complete the questionnaires. Do they get bored and fail to do them? Or not understand them? You might want to run a small study with just a few people in it to see whether it is feasible within your organisation. Is it too disruptive? Then how can you make it possible? What are the challenges to the day-to-day running?

If that goes well, you might try to scale up and run a larger ‘efficacy’ study and see how that goes. Does that work out and still enable you to carry on with normal services?

With all this work in the preliminary studies, did the results indicate positive results? Did you come away thinking that the studies were worth doing?

Could you find independent people to look over the results, provide you with feedback that made you think it was worth continuing? How big was the ‘effect size’?

Did you see a meaningful change in the lives of the people you are working with, and were you measuring it in a way that others can understand and think is sensible?

Phase 1 – Pilot Work
Montgomery, Paul, Ryus, Caitlin R., Dolan, Catherine S., Dopson, S. & Scott, Melinda. (2012)
Sanitary Pad Interventions for Girls’ Education in Ghana: A Pilot Study, PLoS ONE 7(10).
doi: 10.1371/journal.pone.0048274

Phase 2 – Refining/Feasibility
Dolan, Catherine S., Montgomery, Paul, Ryus, Caitlin R., Dopson, Sue & Scott, Linda M. (2013)
A Blind Spot in Girls’ Education: Menarche and its Webs of Exclusion in Ghana, Journal of International Development 26(5): 643–657.
doi: 10.1002/jid.2917

Phase 3 – Efficacy Trial
Montgomery, P., Hennegan, J., Dolan, C., Wu, M. & Scott, L. (2016)
Menstruation and the Cycle of Poverty: A Cluster Quasi-Randomised Control Trial of Sanitary Pad and Puberty Education Provision in Uganda, PLOS ONE 11(12): e0166122.
doi: 10.1371/journal.pone.0166122

Hennegan, J., & Montgomery, P. (2016)
Do Menstrual Hygiene Management Interventions Improve Education and Psychosocial Outcomes for Women and Girls in Low and Middle Income Countries? A Systematic Review, PLOS ONE 11(2).
doi: 10.1371/journal.pone.0146985

Hennegan, J., Dolan, C., Steinfeld, L., & Montgomery, P. (2017)
A qualitative understanding of the effects of reusable sanitary pads and puberty education: implications for future research and practice, Reproductive Health.
doi: 10.1186/s12978-017-0339-9

Hennegan, J., Dolan, C., Wu, M., Scott, L., & Montgomery, P. (2016)
Measuring the prevalence and impact of poor menstrual hygiene management: a quantitative survey of schoolgirls in rural Uganda, BMJ Open 6(12).
doi: 10.1136/bmjopen-2016-012596

Only when you are really prepared are you ready for a larger study! More on that in a later blog.

I hope this was a useful introduction. You can read more about ORBIT here: https://pmc.ncbi.nlm.nih.gov/articles/PMC4522392/

One last thing — in the end, remember that research is not about a black and white choice between whether your intervention is working or not. The reality is that most interventions work for some people but not others, some of the time, in some circumstances.

The research should help you hone it so that you can make it better. Don’t think about the overall result — instead, think about subgroups for whom you might tweak it to improve outcomes.

In this month’s Impact Matters, Professor Paul Montgomery-Marks, Professor of Social Intervention at the Third Sector Research

Cumulative Evaluation for Impact