What is the best way to measure change in relationship education programs?

Traditional Pre-post Research Designs vs. Retrospective-pre-post Designs

The Bottom-line at the Top: Program evaluation researchers looked at two different ways to measure participants’ change as a result of a relationship education (RE) program: (1) the traditional pre-post design that compared targeted outcomes before and after the program, and (2) a retrospective-pre-post design. In this kind of research design, participants are asked at the end of the program to think back to their relationship before the program started and report their levels of a targeted outcome. Then they also report their current levels on these outcomes. How do these two methods of assessing change compare? The researchers found that the traditional pre-test reports could not be compared to post-test reports – that they were no longer “equivalent measures” – so they could not be used to assess average change. On the other hand, retrospective-pre-test reports and post-test reports were “equivalent measures,” so they could be used to assess change in program outcomes. RE evaluation researchers may need to use retrospective-pre-post designs to adequately capture change in certain kinds of outcomes in RE programs. 

Citation: Crapo, J. S., Bradford, K., & Higginbotham, B. J. (2025). Reconsidering the utility of mean comparisons in evaluative work. Family Relations, 74(4), 1578-1590. https://doi.org/10.1111/fare.13186 

Study Set-up
[Note; this research roundup gets a little technical. I’ll explain things as best I can in plain English. It’s an important study to understand. So, read on!]

  • A common way of measuring change in relationship outcomes for participants in relationship education (RE) programs is to compare the group mean from the pre-test of a certain targeted outcome to the group mean of that outcome after participants have completed the program (a post-test or a follow-up post-test). 

  • But there could be a problem. What if the program intervention actually changes how participants understand and think about the outcome that is measured, not just the level of that outcome? 

  • Here’s an example: What if your fatherhood education program targets the quality of father-infant interaction as an important change point. You ask fathers before the intervention begins to rate the quality of their interaction with their infant. But what if they don’t understand some of the key elements of that interaction very well? Maybe they don’t know about sensitive responses, “serve and volley,” and the importance of talking to infants to enhance language development. Dads may think they are doing a pretty good job, and they rate themselves pretty high. Then, they go through the program instruction and learn about these concepts that they weren’t doing so well. At the post-test, they rate themselves again, but this time, they are more aware of the things they should be doing. They are doing better in these areas now, but they know they will need to improve more. As a result, they rate themselves at the post-test a little more realistically. Statistically, this shows up as not much positive change (or even negative change), so the program doesn’t look like it’s making much of an impact on fathers’ quality of father-infant interaction. Depressing! 

  • Evaluation researchers call this possible problem “response-shift bias.” That is, you learn something as a result of the educational intervention that actually shifts your understanding of a concept, like sensitive father-infant interactions. 

  • Can you compare a pre-test score to a post-test score if there is “response-shift bias,” that is, if the way participants think about a target outcome has actually been changed by the intervention? Probably not a good idea. When this shift occurs, you may no longer be measuring the same thing at different points in time. Instead, you are measuring somewhat different things at different times. So, this may not be a fair measure of real change. 

  • Some evaluation researchers believe that response-shift bias is common in educational interventions, especially when participants learn how to think about something differently, like positive communication or partner loyalty or self-care. 

  • One possible solution to this problem is to forgo pre-test measures. Instead, you wait until the end of the program, then ask participants to reflect back to how things were before the program started, say with positive communication skills. Having learned what positive communication skills are (and are not), they can try to remember and report how they were doing on this before the program began. (Yes, some worry that it’s hard to do this accurately.) This is called a “retrospective-pre-post” measure. Then you ask them to report on how things are now – a post-test measure. Evaluators can then compare the retrospective pre-post response to the post-test response to assess change. 

  • Evaluation researchers have debated the pros and cons of the traditional pre-post design and the retrospective-pre-post design. But researchers at Utah State University used some sophisticated statistical analyses to try to assess whether the meaning of a measure – like positive communication or involved fathering – shifts as a result of an educational intervention, making assessments of change more difficult. 

  • If the meaning of something shifts over time, not just the mean or average score, then a retrospective-pre-post design may be a better way to assess change. 

Brief Note on Study Methods

  • The sample was 112 adult individuals who attended an 8-hour PICK (Personal Choices and Knowledge) program, commonly known as “How to avoid falling in love with a jerk (or jerkette),” by Dr. John VanEpp.  

  • Participants completed a traditional pre-test before the 4-week program. Then after the program, they completed both a retrospective-pre-test (“thinking back to before the program . . .”) and a post-test. 

  • Participants reported on two targeted outcomes: relationship confidence and perceived knowledge about healthy relationships. Both of these target outcomes could be vulnerable to “response-shift bias.” That is, the educational intervention could shift the way participants actually think about a certain outcome, not just their level of that outcome. 

  • Researchers tested for “measurement equivalence” of traditional pre-tests to post-tests and retrospective-pre-tests to post-tests using a series of confirmatory factor analysis models. When a measure is equivalent before and after an intervention, it suggests that the meaning of the measure has stayed the same; when the post-test measure is not equivalent to the pre-test, that means the meaning of the measure has changed for them. (Technical note: the researchers tested for configural equivalence, weak equivalence, and strong equivalence. Strong equivalence is needed in order to compare pre-tests and post-tests to assess change in the mean or average level.) 

Key Findings

  • For both outcome measures, the retrospective-pre-test measures demonstrated strong measurement equivalence with the post-test measures, so it was appropriate to compare these measures to assess change.

  • For both outcomes, the traditional pre-test measures did not demonstrate strong measurement equivalence with post-test measures. So, they could not be used to assess change.

Implications for the RE Field

  • Is a traditional pre-test/post-test evaluation the best way to assess whether an educational intervention has created positive change in RE participants? This study suggests that it may not be. If the educational intervention changes how you actually understand or think about a targeted outcome, like father-infant interaction or effective communication, and not just how much of it you have, then you can’t really compare pre-test and post-test average scores. In this situation, evaluators may consider employing a retrospective-pre-post design rather than the traditional pre-test/post-test design. If you are going to expand the concept in people’s minds or shift how they understand it, a retrospective-pre-post design may be a better way to get at change. (Of course, evaluators should test for measurement equivalence between retrospective-pre-post and post-test measures, not just assume it.)

  • Should evaluators use both kinds of measures or just one? Early on in the evaluation process, researchers may want to include both traditional pre-post reports and retrospective-pre-post reports. If statistical analyses show that traditional pre-post reports and post-test reports are not equivalent measures, but that retrospective-pre-post reports and post-test reports are equivalent, then they can drop the traditional pre-tests and just use retrospective-pre-post reports.

  • If evaluators measure outcomes that aren’t really addressed in the intervention, should they use a traditional pre-post design? For example, maybe program developers anticipate that the couple communication skills intervention will spillover and indirectly positively affect parent-child relationships, but they only target couple communication patterns in the intervention, not parenting skills. In these kinds of situations, it’s probably safe to use a traditional pre-test/post-test design. 

  • What about something like marital happiness or co-parenting satisfaction? Does the educational intervention change how people understand or think about these kinds of outcomes? An educational intervention probably doesn’t affect how people understand things like happiness or satisfaction – but we hope that it increases their levels of happiness or satisfaction! So, just using a traditional pre-post design could work. To be safe, however, it’s probably a good idea for evaluators to test for measurement equivalence on all their outcome measures. 

  • Can a retrospective-pre-post design document bigger change than a traditional pre-post design? Possibly. Often RE participants overestimate how well they are doing at something before starting a program because they don’t understand something very well. But then they learn more about relationships in the program and can see that they were missing some things and have a lot of room for improvement. If this is the case, then a traditional pre-post design will likely underestimate change, and a retrospective-pre-post design will likely show a bigger change. (Researchers, however, worry that retrospective-pre-post designs can be biased, too. As always, “more research is needed . . “)

  • What if you have a randomized control group in your evaluation design? Technically, in a randomized controlled trial (RCT), a pre-test is not necessary; differences between treatment and control groups at the end of the program theoretically are a good measure of change. But it’s possible that treatment-group participants’ understanding of the outcome actually changes as a result of the treatment (because they learned more about it) and that control-group participants’ understanding does not change (because they did not receive the intervention). In this case, comparing treatment- and control-groups at the end of the intervention may not be appropriate and may underestimate real change. Evaluators should first establish measurement equivalence on their outcome measures.   

  • Bonus Implication Does a retrospective-pre-post assessment actually contribute to learning? That is, if you have participants think back on how things were before beginning the program, can this contribute to positive effects of the program? There isn’t any research on this question yet, but conceptually it makes sense to me. Asking participants at the end of a program to reflect back on where there were before the program started may heighten their awareness of positive change and give them a greater sense of hope for the future. So, an evaluation need may also enhance learning. 

Last Word: To effectively measure change in many targeted outcomes of RE programs, it may be better to use a retrospective-pre-post research design instead (or in addition to) a traditional pre-post design. And evaluation researchers: always assess measurement equivalence of outcome variables, over time don’t assume it.

Next
Next

Can brief mindfulness instruction and practice embedded in RE programs enhance relationship outcomes?