top of page

Be Notified of New Research Summaries -

It's Free!

Can Algorithms Reduce Racial Disparities in Child Protection?

  • Writer: Greg Thorson
    Greg Thorson
  • 1 day ago
  • 7 min read

Rittenhouse, Putnam-Hornstein, and Vaithianathan (2026) ask whether introducing a predictive risk algorithm changes racial disparities in child protection decisions. They examine administrative data on child maltreatment referrals to Allegheny County, Pennsylvania, from 2010 to 2020, including screening decisions, algorithmic risk scores, child demographics, and foster care removals. They find that the Allegheny Family Screening Tool reduced Black-White disparities in investigation screening by 3.3 percentage points, or 31%, and by 3.1 points, or 48%, after accounting for risk. They also find that home-removal disparities fell by about 1.7 percentage points, roughly 48% of the preexisting gap.


Why This Article Was Selected for The Policy Scientist

This article addresses a consequential policy question: whether algorithm-assisted decisions can reduce disparities while preserving the accuracy of high-stakes public decisions. The issue is increasingly important as governments adopt predictive tools in child welfare, criminal justice, health care, and other settings. The study is timely because concerns about algorithmic bias often shape whether such systems are adopted at all. Published in the Journal of Policy Analysis and Management, a leading applied public policy journal, the article makes an important contribution by comparing algorithm-assisted decisions with human-only decision-making. The administrative data are unusually rich, and the difference-in-differences and event-study approaches strengthen causal interpretation. Generalizability beyond Allegheny County remains uncertain, making replication in other jurisdictions especially valuable.


Full Citation and Link to Article

Rittenhouse, K., Putnam-Hornstein, E., & Vaithianathan, R. (2026). Algorithms, humans, and racial disparities in child protection systems: Evidence from the Allegheny Family Screening Tool. Journal of Policy Analysis and Management, 45(4), e70121. https://onlinelibrary.wiley.com/doi/10.1002/pam.70121 Wiley Online Library


Central Research Question

The study asks whether introducing a predictive risk model into child protection decision-making increases or decreases racial disparities relative to a system based primarily on human judgment. More specifically, it examines the Allegheny Family Screening Tool (AFST), which provides child protection screeners with an algorithmically generated risk score to help determine whether allegations of child maltreatment should be investigated. The central comparison is therefore not between an algorithm and some abstract ideal of unbiased decision-making, but between algorithm-assisted human decisions and the human-only process that existed previously. This distinction is important because both algorithms and human decision-makers may introduce systematic errors or disparities.

The authors focus on two outcomes. First, they examine whether the AFST changed the difference between referrals involving Black and White children in the probability of being screened in for investigation. Second, they examine whether implementation altered racial differences in the probability that a child was removed from the home and placed in foster care following a referral. They also investigate whether reductions in disparities reflected improved decision quality by examining false-positive and false-negative screening decisions.


Previous Literature

The study builds on several related literatures concerning algorithmic decision-making, racial disparities, and child protection. Kleinberg et al. (2018) provide an important foundation by arguing that machine-learning tools can sometimes improve human decisions when predictions are combined with administrative data. At the same time, research has documented circumstances in which algorithms reproduce or amplify disparities. Obermeyer et al. (2019), for example, identified racial bias in a widely used health-care algorithm, while Arnold et al. (2021) examined algorithmic bias in criminal justice. Previous studies comparing algorithm-assisted systems with human decision-making have produced mixed results, including Stevenson and Doleac (2021), Albright (2019), and Howell et al. (2021).


The article also builds directly on work concerning algorithms in child protection. Chouldechova et al. (2018) described the development, validation, fairness assessment, and deployment of the AFST. De-Arteaga et al. (2020) found that screeners changed their decisions in response to algorithmic scores and generally moved their judgments closer to the model’s recommendations. Cheng et al. (2022) compared actual AFST decisions with a hypothetical system in which the algorithm operated without human judgment. Grimon and Mills (2022) provide particularly relevant evidence from a randomized controlled trial. They found that giving child protection workers access to an algorithmic risk score reduced child injury hospitalizations and also reduced racial disparities in child protection contact.


A separate literature documents racial disparities within child welfare systems. Kim et al. (2017) estimate that approximately one-third of U.S. children experience some contact with child protective services before age 18 and report substantial racial differences in investigations. Wildeman and Emanuel (2014) similarly document large differences in foster care placement. Drake et al. (2011) emphasize that observed racial differences may partly reflect differences in underlying exposure to risk factors, while research summarized by Drake et al. (2021) highlights the strong relationship between poverty and maltreatment. Baron et al. (2026) provide especially relevant evidence that referrals involving Black children can receive different treatment even when children have similar potential for future maltreatment.


Data

The analysis uses detailed administrative records from the Allegheny County, Pennsylvania, Office of Children, Youth and Families. The principal data cover child maltreatment referrals between 2010 and 2020, although the exact period varies across analyses because some referral categories and retrospective algorithm scores are unavailable during earlier years. The principal analytical sample contains 93,650 referrals, approximately 85 percent of which are General Protective Services referrals involving allegations such as neglect rather than statutory child abuse.


The data are unusually detailed. Each referral can be linked to children and other household members, and the records contain demographic characteristics, allegations, reporter categories, screening decisions, foster care placements, and identifiers for individual call screeners. The researchers also observe AFST risk scores. For earlier referrals, they use retrospectively generated scores based on a common version of the model, allowing them to compare similarly measured risk across periods before and after implementation.


The main analyses are conducted at the referral level because screening occurs at that level. A referral is classified as involving Black children if at least one child is identified as Black and as involving White children when it includes at least one White child and no Black child. The authors measure removal as whether any child associated with the referral was removed from the home within three months. Referrals involving neither Black nor White children are excluded from the principal racial comparisons.


Methods

The principal identification strategy exploits the August 2016 implementation of the AFST. The authors compare changes in screening outcomes for referrals involving Black children with changes for referrals involving White children before and after implementation. Regression models include year and month-of-year fixed effects as well as controls for allegation type, reporter type, child ages, number of children, and substance-related concerns.


The causal interpretation depends on the assumption that, absent the AFST, Black-White differences would have followed similar trends. The authors examine this assumption using event-study models that estimate racial differences over time before and after implementation. The pre-implementation patterns generally support the required parallel-trends assumption.


For foster care removals, the authors use an additional difference-in-differences design. General Protective Services referrals constitute the treatment group because screeners exercise discretion over whether these referrals are investigated. Child Protective Services referrals provide a comparison group because Pennsylvania law requires them to be investigated regardless of the AFST score. This design helps distinguish changes associated with the AFST from broader changes in racial disparities occurring simultaneously within the county.


The authors also estimate results separately across low-, medium-, and high-risk referrals. This heterogeneity analysis is particularly useful because the AFST includes a high-risk protocol that generally defaults the highest-risk referrals toward investigation unless a supervisor overrides the recommendation.


Findings/Size Effects

Before implementation, substantial racial disparities existed in investigation decisions. Referrals involving Black children were approximately 10.9 percentage points more likely than those involving White children to be screened in. After controlling for referral characteristics and algorithmically estimated risk, the difference remained approximately 6.4 percentage points.


Implementation of the AFST reduced the unconditional Black-White screening disparity by approximately 3.3 percentage points, equivalent to about 31 percent of the preexisting gap. When comparing referrals with similar algorithmic risk scores, the disparity fell by approximately 3.1 percentage points, or about 48 percent. The reductions were especially pronounced among high-risk referrals. In that group, implementation reduced the racial screening gap by approximately 6.0 percentage points, representing roughly 65 percent of the preexisting disparity.


Downstream foster care outcomes followed a similar pattern. The AFST reduced the racial difference in home removals occurring within three months of a referral by approximately 1.7 percentage points. This represented about 48 percent of the preexisting Black-White removal gap. Both the simple pre-post racial comparison and the difference-in-differences analysis using mandatory-investigation referrals as a control produced broadly similar estimates.


The results suggest that the high-risk protocol was an important mechanism. The AFST increased the likelihood that high-risk referrals were investigated, and the strongest reductions in racial disparities occurred within this category. Because investigators and subsequent caseworkers did not observe AFST scores, the algorithm could not directly determine foster care placement. Instead, changes in removals appear to result from changes in which cases were initially selected for investigation.


Decision consistency also improved. Before implementation, individual screeners varied considerably in the size of their Black-White screening differences. After implementation, this variation narrowed, suggesting that the AFST constrained unusually disparate screening patterns.


Finally, the authors examine decision errors. False negatives—cases screened out but followed by a home removal within six months—declined among high-risk referrals for both Black and White children. False positives generally declined among low- and medium-risk referrals, with results depending somewhat on whether false positives were defined by later foster care placement or substantiated maltreatment. Taken together, these patterns indicate that reduced racial disparities were not achieved simply by lowering investigation rates without regard to risk.


Conclusion

The study concludes that algorithm-assisted decision-making can, under some institutional arrangements, reduce racial disparities relative to human-only decision-making while simultaneously improving the accuracy and consistency of decisions. The findings are strongest for high-risk referrals, where the AFST does more than provide information: it establishes a default recommendation favoring investigation unless a supervisor overrides it.

The authors nevertheless emphasize important limitations. Racial disparities remained after implementation, and the AFST affects only one stage of a much larger child protection process. It cannot address disparities in which families are initially reported, nor can it directly determine later decisions involving substantiation, services, case opening, or removal.


The findings are also context dependent. Allegheny County combines a particular algorithm with a specific administrative structure and high-risk protocol. Algorithms that merely present information, or jurisdictions with different baseline disparities, populations, institutional practices, or limits on worker discretion, could produce different results. Consequently, the study does not establish that predictive algorithms will generally reduce racial disparities. Instead, it demonstrates that the design of the algorithm, the decision rules surrounding its use, and the manner in which human professionals interact with it are central to determining its effects.

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
Screenshot of Greg Thorson
  • Facebook
  • Twitter
  • LinkedIn


The Policy Scientist

Offering Concise Summaries*
of the
Most Recent, Impactful 
Public Policy Research

*Summaries Powered by ChatGPT

bottom of page