Transcript for 2026 Summer Training Series Session 3: Linking NCANDS and AFCARS Presenter: Alexander F. Roehrkasse, Ph.D., Butler University National Data Archive on Child Abuse and Neglect (NDACAN) [MUSIC] [VOICEOVER] National Data Archive on Child Abuse and Neglect. [ONSCREEN CONTENT SLIDE 1] Welcome to the 2026 NDACAN Summer training series! The session will begin at 12pm EST. This session is being recorded. Please submit questions to the Q&A box! [Alyssa Lindsey] All right. Hello, everyone. Welcome to our 2026 NDACAN Summer Training Series. My name is Alyssa Lindsey. I'm the Graduate Research Associate here at NDACAN. Before I kick off our time together and pass it along to our presenter, I just want to give a couple of housekeeping items. We are recording this session so that we can post the recording, transcripts, slides, any other related materials on our website several weeks following each presentation. And I'll announce their availability on the Child Maltreatment Research L or CMRL listserv. And I'll put the information on how you can subscribe to that listserv in the chat. Please let me know if you have any questions by using the Q&A box, which is on the lower right-hand side of your Zoom screen. And because we're recording this session, we're using a webinar format, which means that the Q&A box that I mentioned is really the only way that you have throughout the presentation to communicate, ask questions. Please, as the presenter is going, feel free to put your questions throughout. And then at the end of the presentation, we'll have time to go through those questions in the order that they come in. We really encourage you to ask questions throughout and just know we'll get to them at the end of our time together. Next slide. [ONSCREEN CONTENT SLIDE 2] NDACAN Summer Training series. National Data Archive on Child Abuse and Neglect Duke University, Cornell University, UC San Francisco, & Mathematica [Alyssa Lindsey] Great. So welcome again to the NDACAN Summer Training Series. NDACAN stands for the National Data Archive on Child Abuse and Neglect. It's housed at Duke University, Cornell University, UC San Francisco, and Mathematica. Next slide. [ONSCREEN CONTENT SLIDE 3] Laying the groundwork foundational skills for using administrative data in child welfare research. Logo for the Children's Bureau features an image of overlapping blue and white silhouettes of children next to red and white stripes. To the right is the text "Children's Bureau: An Office of the Administration for Children & Families." Logo image which features a semicircle of icons representing people holding hands positioned on top of the acronym NDACAN, and next to the text National Data Archive Child Abuse and Neglect. [Alyssa Lindsey] And NDACAN does two learning offerings every year. We have our monthly Office Hours series during the academic year, and in the summertime we have the Summer Training Series. The theme of this year's 2026 Summer Training Series is "Laying the Foundational Skills for Using Administrative Data and Child Welfare Research." And this series is designed for both the expert and also the beginner or the intermediate folks. So we're using this to be useful for folks who are both new to the area or who are seeking to strengthen their existing skill sets. And throughout these summer presentations, we're aiming for each session to be useful, a useful starting point for conducting research on child welfare using the administrative data sets that we talk about in each presentation. Next slide. [ONSCREEN CONTENT SLIDE 4] NDACAN Summer Training series schedule. July 1st: Overview of NDACAN administrative datasets July 8th: Data cleaning and management July 15th: Linking NCANDS and AFCARS July 22nd: Handling missing data July 29th: Data presentation and visualization [Alyssa Lindsey] And this is just our schedule of the NDACAN Summer Training Series. If you weren't able to attend the first two, no worries. The recordings and transcripts will be available shortly. In fact, the recordings and materials for the first session are currently available on our website, and I'll make sure to put that link in the chat for everyone. We'll be meeting every Wednesday at 12 Eastern for an hour for the month of July. These are the topics that we'll cover. We've already had presentations on an overview of NDACAN administrative data sets. We had a presentation last week on data cleaning and management. This week, we'll focus on linking NCANDS and AFCARS. And then the last two presentations will focus on handling missing data and data presentation and visualization. Next slide. [ONSCREEN CONTENT SLIDE 5] Session Agenda. Linking NCANDS and AFCARS Demonstration in R Q & A [Alyssa Lindsey] Great. So today's session, as I mentioned, is going to start with a brief presentation over linking NCANDS and AFCARS data, following with a demonstration in R using those data. And then we will end with audience Q&A that you submit in that Q&A box that I mentioned earlier. Next slide. [ONSCREEN CONTENT SLIDE 6] Linking NCANDS and AFCARS [Alyssa Lindsey] Great. So now I will pass it off to Alex, who's going to give the presentation today. Thank you. [Alex Roehrkasse] Thanks, Lisa. My name is Alex Roehrkasse. I'm a research associate at the archive and an assistant professor of sociology and criminology at Butler University. Thanks for being here today. I'm really excited to give this presentation. I gave last week's presentation on data, cleaning and management. Today's presentation is going to build on some of those foundational skills, introduce some new skills. I want to acknowledge that data linking or record linkage is a huge topic, a topic that's had significant methodological advances in recent years. We're just going to be skimming the surface here today, introducing some foundational skills. and trying to demonstrate those skills through concrete examples using archive administrative data. We'll be talking about how to link NCANDS to itself across multiple years. Ditto AFCARS. But we'll be building to an example where we link NCANDS to AFCARS as a way of demonstrating some of the most powerful uses of administrative data when combined in this way. A reminder that the presentation might feel fast for some in attendance today, and might feel slow for others. No worries, hang in there. Remember also that the slide deck for today, a recording of the presentation, but also the code and output for the demonstration in R will be available on the NDACAN website in the coming days and weeks. Alright, let's get started. [ONSCREEN CONTENT SLIDE 7] What is record linkage? Linkage combines multiple data sources based on one or more shared variables Internal record linkage NDACAN administrative data files (NCANDS, AFCARS, NYTD) can be linked to each other at the child level using unique (encrypted) child IDs External record linkage Aggregated NDACAN data can be linked at the aggregate level to external sources using common variables: Time: year, month, half-month Place: state, county Demographic groups: sex, race/ethnicity, age [Alex Roehrkasse] Okay, what do we mean when we talk about record linkage or data linking. Linkage essentially combines multiple data sources based on one or more shared variables. We'll be talking about these shared variables today as linking variables. The variables, the information that allows us to combine these multiple data sources in the first place. We're going to be focusing today on internal record linkage. But I want to clarify that it's possible to link records internally within the NDACAN ecosystem, but also link records externally. Internally, NDACAN administrative data files, particularly NCANDS, AFCARS and NYTD can be linked to one another. We can do this at the aggregate level, say, measuring certain outcomes at the state year level or the county month level. But most common, and frankly, most powerful, is to link these data files at the record level, at the child level, using unique, encrypted child identifiers. We can also link NDACAN data to external data, data outside the NDACAN ecosystem but we can't do this external linkage at the child level. Currently, our data infrastructure does not allow for that in a responsible way, in a way that, protects the, privacy, the non-identifiability of the children who are represented in our data sets. So for that reason, data users most often link to external data at an aggregate level. A very common use case here would be if you wanted to, say, measure a population incidence rate, the rate at which children experience maltreatment reports or placement into foster care. This would require information about the child population that would come from outside the NDACAN data ecosystem. [ONSCREEN CONTENT SLIDE 8] Benefits to linking. Combining sources expands range of measures Individual-level record linkage: within NDACAN administrative data Aggregate-level record linkage: external data sources Repeated observations Enhance missing data solutions Help identify and address measurement error Enable longitudinal research designs [Alex Roehrkasse] Why do we link data? What are the upsides to engaging in this research practice? I think they're basically twofold. First, data linkage allows us to combine sources in ways that expand the information at our disposal, increase the number of measures that we have when we're analyzing our data. For example, let's say we want to study children who show up in the AFCARS, but we want to know something about their maltreatment history. That information is not included in the AFCARS, but if we can link children in AFCARS to NCANDS, we can start to recover some of that information. Another reason to link data is to generate repeated observations of the same unit, usually the same child, over time. Because the data we're talking about today are administrative data, insofar as children show up again and again when these data are collected, we can generate longitudinal data, or panel data, that follow children over time. This can be especially useful in dealing with missing data problems which are common in our administrative data. It can help identify and address sources of measurement error. And perhaps most excitingly, it can enable longitudinal research which is of interest, not only if you're interested in change over time in trends, but also, if you're interested in things like causal inference. Having panel data allows us to implement a number of different causal inferential designs that aren't feasible with cross-sectional data. [ONSCREEN CONTENT SLIDE 9] Pitfalls of linking. Myriad errors arise from linking less-than-clean data Non-missing values of shared variables may not agree Linking may result in (systematic) measurement error NDACAN child IDs are state-specific; interstate moves lead to false negative links NDACAN data reflect variation and changes in record-keeping Data linkage can create/amplify missing data problems [Alex Roehrkasse] There are nevertheless some serious downsides to data linkage things about which we need to be very careful and highly aware. When our data are less than clean, this can generate any number of problems. When we try to link less than clean data, these errors can really multiply. So when it comes to data linking, it's especially important that our data are really, really tidy. I'd refer you back to last week's presentation for some foundational skills about how to clean and tidy your data. Sometimes we find that when we link data, values that should agree don't. What does it mean when one observation of a child substantively disagrees with another observation of a child? How do we reconcile these discrepancies? The best choices aren't always obvious. Lastly, linking can result in measurement error, including systematic measurement error. Let me give 2 brief examples, common examples. Throughout today's presentation, we're mostly going to be talking about child IDs as linking variables. That is to say, variables that uniquely identify children. The variables that uniquely identify children in NDACAN administrative data are state-specific. They're specific to each state. What this means is that when a child moves from one state to another we're not able to link them across that interstate move. That child will appear as two distinct children when it is, in fact, a single child in two different places at two different points in time. When we try to link this child, it will result in a false negative link. That is to say, a scenario where there is in actuality a true link between these records, and we fail to observe it. A second case. NDACAN data also reflect variation and changes in record keeping. States sometimes change the way they record children's unique identifiers. Sometimes these changes break our ability to follow children across time. This will have similar consequences to the challenge just described, except not only when children move across states, when children cross that record-keeping threshold within a given state. Lastly, we'll be talking about missing data next week, but I'll say here that data linkage can create and amplify existing missing data problems, particularly when our record linkage rates are low. [ONSCREEN CONTENT SLIDE 10] Data linkage and child welfare system contact. NCANDS: Child protective history AFCARS: Foster care experience NYTD: Transition out of care [Alex Roehrkasse] Okay, to kind of set the stage for some of the linking we'll be doing today, a very, very brief recap of what was presented in our first session this summer, an overview of the administrative data that NDACAN distributes and maintains. First, NCANDS. NCANDS compiles records of investigated child maltreatment reports. So this is often where we go to learn about children's history of contact with CPS. AFCARS, more specifically the AFCARS foster care files track children's contact with the foster care system. So whenever a child passes through that system, they generate a record. NYTD is a slightly different administrative data set. It tracks only children aging out of foster care. So a much smaller population. Furthermore, it's a sample of that population, not a full population data set. But it can be very important, depending on the research questions you're interested in. We're going to be focusing today on linking NCANDS and AFCARS to themselves over multiple years and eventually to linking the NCANDS to the AFCARS. [ONSCREEN CONTENT SLIDE 11] Merging by row: stacking data Multiple years of NCANDS or AFCARS data can be combined simply by “stacking” them, or in R, by row-binding year-specific data frames Stacking is usually smooth because NDACAN administrative data have highly consistent structure over time Take care to distinguish between the occurrence, recording, and submission of information about events [Alex Roehrkasse] I'm going to walk us through today two different ways of linking data. These are different conceptually and different practically. The first way of thinking about linking data isn't always understood as record linkage or data linking, but I want to make the case that it is. And this is what we'll call merging by row, or stacking data. Multiple years of NCANDS or AFCARS data can be combined simply by stacking them. You can even imagine literally setting one data set on top of the other. In R, we call this row binding where we are binding together by row year specific data frames. Stacking is actually usually pretty smooth using our datasets, because they tend to have highly consistent structures over time. Unfortunately, they're not entirely consistent. Again, see last week's presentation for some issues that can arise when variables aren't named exactly the same way in each year. Whenever we're linking by stacking, we want to take special care to distinguish between the occurrence, the recording, and the submission of information about events. Again, see last week's presentation for some thoughts about this and tips and tricks for dealing with it. [ONSCREEN CONTENT SLIDE 12] Stacking as linking By stacking multiple years of data within a given source (NCANDS, AFCARS), users can create longitudinal data, i.e. repeated observations Record-level linking Users can link children by identifying multiple observations with shared unique child identifiers (ideally StFCID, a combination of child identifier and state) Aggregate linking By summarizing data at the state or county level, users can link geographies by identifying multiple observations with shared geographic identifiers (state or county FIPS codes) [Alex Roehrkasse] What does it mean to think about stacking data as record linkage? Well, by stacking multiple years of data within a given source, users can generate longitudinal data. In other words, repeated observations of some unit over time. We can do this at the record level. So imagine linking children simply by identifying multiple observations across years. Using some shared, unique child identifier. For our purposes today, and for most research purposes, this child identifier is going to be the variable STFCID. FCID stands for Foster Care ID. Recall that I said that most child IDs are unique only within states. And so the ST stands for state. The STFCID variable is a concatenation of a state identifier and a unique child identifier. So across all data sets, it uniquely identifies single children in administrative data. We can also link by stacking aggregate data. This just requires us to summarize our stacked data at some aggregate level, say the state or the county, the year or the month. We use some shared identifier, like a state or county FIPS code. [ONSCREEN CONTENT SLIDE 13] Linking by stacking. Two tables are shown side by side. Table 1 on the left is a simplified NCANDS Child File Data showing SubYr, StFCID, RptDt, and RptDsp values for records from 2019 and 2020. The data are stacked by year with 2019 records listed first, and 2020 records next. In between the tables is the text "Rearrange by StFCID and RptDt". Table 2 on the right is a simplified NCANDS Child File Data showing SubYr, StFCID, RptDt, and RptDsp values for records from 2019 and 2020. The data are sorted by RptDt and by StFCID. [Alex Roehrkasse] Here's a trivial example of linking by stacking. Consider a toy example where, on the left hand, we see a sort of simplified version of the NCANDS child file. In the first column we see the submission year. Notice. We only have 2 values here, 2019 and 2020. Essentially, what we've done is stacked the 2019 Child File on top of the 2020 Child File. In the second column, you'll see this unique child identifier. These are simplified versions of a child identifier, but the actual child identifiers are always a two-letter code corresponding to the state from which the record comes, and then a long string of numbers. In the third column you'll see the report date variable, and in the fourth column the report disposition variable. On the right-hand side, you'll see the exact same table all I've done is rearrange the rows. On the left-hand side, things are ordered first by submission year, and then by STFCID. On the right-hand side, I order things first by unique child identifier, and then by report data. What you see from the green highlighted rows is that simply by rearranging after stacking. We have repeated observations of some children over time. The child from Alabama had a report in 2019. and it were in the 2019 submission year, and again in 2019 calendar year, but in the 2020 submission year. All of a sudden we have panel data or longitudinal data tracking children over time. These data are organized long. We could imagine, though, pivoting them or reshaping our data. To generate rows corresponding uniquely to individual children. [ONSCREEN CONTENT SLIDE 14] Cautionary notes about linking by stacking. Child identifiers are state-specific Children experiencing recorded events in multiple states will therefore appear as multiple children States change identification procedures Children experiencing recorded events within a state may therefore appear as multiple children Prior or future events may be unobserved Children experiencing recorded events outside the sampling frame will appear as having too few events [Alex Roehrkasse] A couple cautionary notes about linking by stacking. First, remember that child identifiers are state-specific. What this means is that children experiencing recorded events in multiple states will appear falsely as multiple children. It's important to think about what kind of bias this introduces to your research, depending on your estimand. Generally speaking, this won't bias estimates of incidents. but has pretty more significant consequences for measures of prevalence, or sequence, or life course measures. Secondly, as I've also mentioned already, states change their identification procedures. This can introduce similar biases, but for children who remain in a single state and experience events on either side of a threshold where recording practices change. Lastly, in scenarios like this, it's very important to think carefully about how you define your sample in terms of child age and historical time. Children may experience recorded events that fall outside of our sampling frame. We may think we're experiencing the first and second event experienced by a child, when, in fact, it's the third and fourth experience. So care here is needed when we're making claims about cumulative events or sequencing. [ONSCREEN CONTENT SLIDE 15] Selected NCANDS child file multi-year linkage rates. Note: Numbers are percentages of children in NCANDS Child Files, linked by ChID in adjacent years, for whom DOB and sex match (true positive link). Table showing selected NCANDS child file multi-year linkage rates by state across year pairs from 00-01 to 20-21. States listed include Arizona, California, Florida, Illinois, Indiana, Massachusetts, Michigan, New York, Oregon, Pennsylvania, Tennessee, and Texas. Most cells show linkage rates near 95 to 100 percent, with several zero or missing entries highlighted in yellow, green, gray, and white. [Alex Roehrkasse] On this slide, you see a table that represents internal analysis of archive data conducted by archive staff. Let me walk you through what you're looking at. For a select number of states listed on the left-hand column, what we've done is link the NCANDS child files in adjacent years. So, for example, link 2012 to 2013, 2013 to 2014, 2014 to 2015, etc, etc. We then examine those records which we do successfully link, and ask: what percentage of those linked records have consistent measures of date of birth, and sex? we would expect linked records to have consistent values of date of birth and sex. We would like the percentage to be 100. But, as you can see, that's the case for only a small minority of adjacent years in these states. In other cases, we appear to be matching records which do not have consistent values of time-invariant child characteristics. This begs questions about whether we have false positive links. I want to leave you with 2 thoughts walking away from this slide. First, Data quality, is going to influence link quality in ways that vary by state and by year. It's important that you investigate on your own the quality of the data you use and its implications for record linkage. Second, there's kind of two ways we want to think about linkage going wrong. One is false negative as I've talked about already, cases where we in reality, know that a child has experienced multiple events, but we fail to observe that repeated process. This is what we'd call a false negative. On the other hand, false positives are illustrated here by this table scenarios where we observe links which are not, in fact, repeated experiences of events experienced by distinct children. [ONSCREEN CONTENT SLIDE 16] Merging by column: joining data. The NCANDS and AFCARS (as well as NYTD) can be linked to each other by adding columns to each row, or in R, “joining” the data. Joining is more complex than stacking. [Alex Roehrkasse] Okay, now, I want to move to a second kind of linking. That we'll call merging by column. In the language of R, we talk about this as joining data. This is generally more in line with what we mean when we talk about record linkage. It's both more complex than stacking, but also more powerful than stacking. When we talk about combining different data sets, like linking NCANDS to AFCARS, we'll always be talking about joining data. We join data when we want to add different variables from distinct data sets. Measuring this information for the same unit. [ONSCREEN CONTENT SLIDE 17] Joining data. Four Venn diagrams illustrating four distinct methods of joining two datasets, NCANDS and AFCARS. Inner join includes only those records present in both datasets, left and right joins include all those records present in one dataset, and full join includes all records present in either dataset. [Alex Roehrkasse] I want to introduce some basic sort of language and concepts around joining data. imagine that we're joining NCANDS to AFCARS, and imagine that each circle here represents each of those datasets. Whenever we join the two datasets, they will overlap only partially. Some observations will be observed in both. Some will only be observed in one and in the other. This means that when we join our data, we can decide which records we want to keep. When we inner join data, we retain only those records that are jointly observed in both of the linking datasets. When we left or right join the data, we keep all of the records in one dataset, and only those records from the other dataset that link to them. When we full join our data, we keep all of the data in both linking data sets, and we need to take care to keep track of which observations are linked and which are not. [ONSCREEN CONTENT SLIDE 18] Joining multiple observations. One to one: Linking variable is unique in both datasets. One to many: Linking variable is unique in one dataset. Many to many: Linking variable is not unique in either dataset. [Alex Roehrkasse] Whenever we're joining data, that is to say, merging by column. We also need to think carefully about how rows line up with one another. Most data joins are one-to-one joins. In this scenario a linking variable uniquely identifies rows in both data sets. If this is the case, a row in one data set will always exactly correspond to a row in the other data set. Sometimes we do a one-to-many, or less likely, a many-to-many merge. Let me give an example of a one-to-many merge. Let's say we were linking records from the AFCARS foster care files to external data about, say, unemployment rates at the state year level. Our state year observations of the unemployment rate. would map onto many different AFCARS records for all children living in that state year. Sometimes we intend a one-to-many or a many-to-many match. Sometimes, though, if we've failed to set up our linkage correctly, we will accidentally do a one-to-many or many-to-many match. And in this scenario we may see our data set explode. The number of observations go up by a factor of 10 or 100 or 1,000. This is often an indication that something's gone wrong. [ONSCREEN CONTENT SLIDE 19] Linking by joining: aggregate data. A sample "NCANDS" data table and a sample "AFCARS" data table are shown with a plus sign in between them, and then an equal sign, and to the right of the equal sign is a table "Linked Data". Description of the Sample NCANDS Data Table: The colums are FY, StaTerr, and NRpt. The data are collapsed to the state-year level, with NRpt being the number of reports occurring in each state year. Description of the Sample AFCARS Data Table: The colums are FY, StaTerr, and NRem. The data are collapsed to the state-year level, with NRembeing the number of removals occurring in each state year. Description of the Linked Data Table: The colums are FY, StaTerr, NRpt, and NRem. The data are summarized to the state-year level counting up the number of removals that occurred in each state-year. [Alex Roehrkasse] Let me give briefly a somewhat trivial example of linking by joining aggregate data. On the left hand, we see a toy version of the NCANDS where we have collapsed or summarized our data to the state year level. The leftmost column is the fiscal year. Next, we have a two-letter state identifier. And then the N report variable is a count of the number of reports occurring in each state year. We then see AFCARS data, where, again, we've summarized things by the state year level. Counting up the number of removals that occurred in each state year. Clearly, each of these little tables represents some pre-processing and some cleaning. We've already done some, stacking by linking. We've done some summarizing, and we've renamed our variables so that our linking variables, in this case fiscal year and state identifier, match exactly. We could then join our data using these fiscal year and state ID variables to create a new linked data set where for each unit, namely each state year, we have observations of information coming from both NCANDS and AFCARS. This is a pretty trivial example, though you can almost imagine literally smashing these data sets together left to right. [ONSCREEN CONTENT SLIDE 20] Linking by Joining: microdata Question: What are the maltreatment histories of children aged 0-3 (NCANDS) placed into foster care in FY2023 (AFCARS)? Process: Sample using 2023 AFCARS Foster Care File. Stack and summarize 2020-2024 NCANDS Child Files. Left-join one-to-one using unique child identifier StFCID. Analyze maltreatment histories of placed children. [Alex Roehrkasse] More valuable, more interesting, but also more challenging, is linking by joining microdata. This is where we'll start to transition to our demonstration in R. And as we do so, I want to put forward a research question that will highlight the necessity but also practicality of joining data at the micro level. Let's say we wanted to answer the question, what are the maltreatment histories of children aged 0 to 3 who were placed into foster care in fiscal year 2023? This is a question we can only answer using linked microdata. How would we answer this question? First, we would generate a primary sample using the 2023 AFCARS foster care file. We would pull out only those children who are age 0 to 3, and only those children who entered foster care in that federal fiscal year. Then we'd set that aside. Next, we would take the 2020 through 2024 NCANDS child files. We would stack them and summarize them, grouping by unique child ID, and generating for each unique child account of the number of reports or number of substantiated reports that that child experienced over those years. Once we had those 2 cleaned and tidied data sets, we would have 2 data sets in which each row corresponded to a unique child. We're then prepared to left join our data in a one-to-one merge using our unique child identifier. Once we've done this, we have linked data that we can use to analyze the maltreatment histories of children placed into foster care. As we transition now to a demonstration in R, we're going to be working toward an example where we do exactly this and answer this research question. [ONSCREEN CONTENT SLIDE 21] Demonstration in R [ONSCREEN CONTENT SLIDE 22] Additional resources R for Data Science (2e), Hadley Wickham, Mine Çetinkaya-Rundel, and Garrett Grolemund [Alex Roehrkasse] As we transition to R, I want to leave you briefly with an additional resource. I mentioned this resource last week, but I want to highlight this one again as being especially useful for foundational skills relevant to data linkage. I want to remind you that in last week's presentation, I gave a little bit of information about how to get up and running with R and RStudio. I won't reiterate that today. Instead, we'll start to move to our demonstration in R, and if you'll bear with me as we pull up RStudio. Here we are. [ONSCREEN CONTENT] ######### # NOTES # ######### # This program file demonstrates strategies discussed in # session 3 of the 2026 NDACAN Summer Training Series # "Linking NCANDS and AFCARS." # For questions, contact the presenter # Alex Roehrkasse (aroehrkasse@butler.edu). # Note that because of the process used to anonymize data, # all unique observations include partially fabricated data # that prevent the identification of respondents. # As a result, all descriptive and model-based results are fabricated. # Results from this and all NDACAN presentations are for training purposes only # and should never be understood or cited as analysis of NDACAN data. [Alex Roehrkasse] Hopefully, this interface is somewhat familiar at this point. We're working in RStudio in the upper left-hand quadrant. You'll notice that we're working from an R script. I've left some notes at the top of this script, indicating that this script is for this presentation. Here is how you can reach me if you have any questions about it. A reminder that this script and all the output from today will be posted on our website. And another reminder that just as last week, we're going to be working with data that have realistic properties, but which are not real data. I've taken measures to fabricate the data that we're using today, anonymizing them to prevent the identification of children covered in our data. So we won't understand what we're doing today as an analysis or a citation of NDACAN data. [ONSCREEN CONTENT] ##################### # TABLE OF CONTENTS # ##################### # 0. SETUP # 1. LINKING BY STACKING: AGGREGATE DATA # 2. LINKING BY STACKING: MICRODATA # 3. LINKING BY JOINING: MICRODATA [Alex Roehrkasse] We're gonna do 3 things today. We'll set up the environment. But I'm going to skip over that. If you want to walk through of how to set up the environment. I'll refer you back to last week's presentation. Then we'll do examples of three kinds of linking. Linking by stacking to produce aggregate data. Linking by stacking to produce microdata, and then linking by joining microdata. That last research example we talked about. Okay, we'll go ahead and set up the environment. [ONSCREEN CONTENT] ############ # 0. SETUP # ############ ## SETTING UP THE ENVIRONMENT ## # Let's clear the environment. rm(list=ls()) # Pacman installs packages if necessary, otherwise loading them. if (!requireNamespace("pacman", quietly = TRUE)){ install.packages("pacman") } pacman::p_load(data.table, tidyverse) # Let's define some filepaths # (note the organization of project and data folders). project <- 'C:/Users/aroehrkasse/Box/Presentations/-NDACAN/2026_summer_series/' data <- 'C:/Users/aroehrkasse/Box/NDACAN/2026_summer_series/' # And set one as the working directory. setwd(project) # Always set a seed to allow for reproduction of random processes. set.seed(1013) [Alex Roehrkasse] And we're going to read in today 3 different data files. The first data file we'll read in, we created in last week's presentation. This is a cleaned and tidied stacked version of the NCANDS child files for submission years 2020 through 2024. You'll see that that object has now appeared in our environment over on the upper right hand quadrant. [ONSCREEN CONTENT] ## READING DATA ## # Let's read in our cleaned # anonymized versions of the # NCANDS Child Files for 2020-2024. nc <- read_rds(paste0(data,'ncands_clean.rds')) # Let's also read in cleaned # anonymized versions of the # AFCARS Foster Care AB files for 2023 and 2024. ac23 <- read_rds(paste0(data,'afcars23_clean.rds')) ac24 <- read_rds(paste0(data,'afcars24_clean.rds')) # For the purposes of the presentation, # note again that I have sampled New England # and selected a small number of key variables. [Alex Roehrkasse] We're also going to read in 2 files of AFCARS data corresponding to the AB foster care files for 2023 and 2024. A reminder that for the purposes of demonstration today, just to keep things moving and avoid getting stuck with computational demands, we're analyzing just a sample from New England states, and I've selected a small number of key variables that will be of interest for today's demonstration. [ONSCREEN CONTENT] # Recall from S2 that nc is already "stacked" # after we unlisted our NCANDS submission-year # into a single data frame. nc |> + count(staterr,subyr) staterr subyr n 1: CT 2020 74880 2: CT 2021 79922 3: CT 2022 79203 4: CT 2023 76600 5: CT 2024 70930 6: MA 2020 16093 7: MA 2021 15555 8: MA 2022 17957 9: MA 2023 19654 10: MA 2024 20495 11: ME 2020 16165 12: ME 2021 14249 13: ME 2022 15501 14: ME 2023 15644 15: ME 2024 14980 16: NH 2020 25308 17: NH 2021 23445 18: NH 2022 20607 19: NH 2023 22437 20: NH 2024 20134 21: RI 2020 3622 22: RI 2021 3371 23: RI 2022 4574 24: RI 2023 4744 25: RI 2024 4207 26: VT 2020 9308 27: VT 2021 8078 28: VT 2022 7009 29: VT 2023 7215 30: VT 2024 6151 staterr subyr n [Alex Roehrkasse] Okay, let's first link by stacking, analyzing aggregate data through this method of linking. Recall that from last week's presentation, this NC object is already a stacked version of the NCANDS. We can see this by counting up the number of records. Corresponding to each state and submission year. We have multiple years of data for each state. In each state year, we have tens of thousands of records. We did this by programmatically cleaning our data, creating a list of NCANDS child files, L-applying our cleaning function, and then eventually unlisting our child files into a single data frame. We can more explicitly stack our data by row binding multiple years of data. Let's illustrate this with AFCARS. [ONSCREEN CONTENT] # We can do the same explicitly by row-binding # multiple AFCARS AB files. # Note that this requires consistent variable naming and encoding. ac <- ac23 |> + bind_rows(ac24) > ac |> + count(st, fy) st fy n 1: CT 2023 12359 2: CT 2024 12147 3: MA 2023 4593 4: MA 2024 4687 5: ME 2023 1852 6: ME 2024 1786 7: NH 2023 3396 8: NH 2024 3442 9: RI 2023 1530 10: RI 2024 1428 11: VT 2023 2636 12: VT 2024 2412 [Alex Roehrkasse] We can take our AFCARS 2023 data and use the bind rows function to bind to it the 2024 data. When we do this, and count the number of rows corresponding to each state and year, again, we see that we have some stacked data. [ONSCREEN CONTENT] # Say we wanted to link states across years, # measuring the annual proportion of maltreatment reports # that were substantiated. # This requires using NCANDS to link at the state-year level # by stacking and summarizing. # Recall also that we need Y+1 years of data to avoid # bias from delayed reporting, so let's examine reports # from FY2020-2023 using data from FY2020-2024. state_year <- nc |> + filter(rptdt %between% + c('2019-10-01','2023-9-30')) |> # define sampling frame + mutate(fy = if_else(month(rptdt) >= 10, + year(rptdt) + 1, + year(rptdt)), # create fiscal year variable + sub = case_when(rptdisp == 'Substantiated' ~ 1, + is.na(rptdisp) ~ NA, + T ~ 0)) |> # create substantiation indicator + group_by(fy, staterr) |> # group by year and state + summarize(nsub = sum(sub, na.rm = T), # count substantiated reports + n = n(), # count all reports + .groups = 'drop') |> + mutate(sub_prop = nsub/n) |> # create substantiation proportion variable + arrange(staterr, fy) # organize [Alex Roehrkasse] Say that we want to link states across years, measuring, let's say, the annual proportion of maltreatment reports that were substantiated. This requires using NCANDS to link at the state year level by stacking and then summarizing. Recall from last week's presentation that we need an additional year of data when we're analyzing NCANDS to avoid bias from delayed reporting. So we'll analyze fiscal years 2020 through 2023 by using data from the 2020 through 2024 child files. How will we do this? We'll take our stacked data file. We'll pull out only those records which have report dates following in our period of analysis. We'll define our sampling frame explicitly using report date rather than submission date. We'll generate a fiscal year variable which corresponds to the report date as long as the report is in January through September. For example, if a report comes in November of 2019, it's actually in the 2020 fiscal year, and so we'll add a year to that. We'll create an indicator variable, which is equal to 1 when the report disposition is substantiated, and otherwise equal to 0. We'll group our data by the state year level and generate two variables: a count of the reports that are substantiated, and a count of all reports. We then ungroup our data and generate a proportion variable that captures the proportion of all records that were substantiated. When we run this code, and plot the output, we see that we now have a panel of state level information where, for each state across multiple years, we observe the proportion of reports that were substantiated. [ONSCREEN CONTENT] # Notice that we now have a panel of state-years # with the linking variables (a combination of state and year) # and the outcome of interest over time (substantiation proportion). state_year |> ggplot(aes(x = fy, y = sub_prop, color = staterr)) + geom_point() + geom_line() [ONSCREEN CONTENT R CODE IMAGE 1] Graph of fy on x-axis and sub_prop on y-axis. A panel of state-years 2020-2023 for CT, MA, ME, NH, RI, and VT with the linking variables (a combination of state and year) and the outcome of interest over time (substantiation proportion). [Alex Roehrkasse] We now have longitudinal data, or panel data, resulting from linking by stacking. This is a simple example of linking by stacking at the aggregate level. What about at the record level or micro data? Linking microdata by stacking actually requires less manipulation, but probably requires more care and caution. Let me try to illustrate this using the AFCARS data. [ONSCREEN CONTENT] # Let's try to link individual children across the # 2023 and 2024 fiscal years. In principle, # once our AB files are stacked, this only requires # use of the linking variable, the unique child ID StFCID, # to identify children over time. # Notice that with just a little rearranging, we identify # linked children: rows with matching values of the # linking variable StFCID. ac |> + arrange(stfcid, fy) |> # order descending by state, then year + select(fy, st, stfcid, entered, exited, inatend, inatstart) |> + slice(1:10) # pick out the first 10 rows fy st stfcid entered exited inatend inatstart 1: 2023 CT CT000490093038 1 0 1 0 2: 2024 CT CT000490093038 0 0 1 1 3: 2023 CT CT000490273805 1 0 1 0 4: 2024 CT CT000490273805 0 1 0 1 5: 2023 CT CT000490376258 1 1 0 0 6: 2023 CT CT000490435166 1 0 1 0 7: 2024 CT CT000490435166 0 1 0 1 8: 2023 CT CT000490437443 0 0 1 1 9: 2024 CT CT000490437443 0 1 0 1 10: 2023 CT CT000490850584 0 0 1 1 [Alex Roehrkasse] Let's try to link individual children across the 2023 and 2024 fiscal years. In principle, once our files are stacked, this only requires using a linking variable, namely the unique child ID STFCID, to identify those children appearing in both fiscal years. Notice that with just a little rearranging of this stacked data, we can identify linked children, namely, rows with matching values of the linking variable. So we take our stacked AFCARS data, we arrange things first by unique ID and then fiscal year. This is important whenever we're using commands that rely on the ordering of records. We'll pull out a few variables of interest and look at just the first 10 rows of the resulting data. What do we see? Simply by rearranging and kind of glimpsing this stacked dataset we see that in a number of cases, we have repeated observations of children across multiple years. Here's a child from Connecticut with an ID ending in 38. We see that that child appears multiple times in our stacked dataset. In other words, we have a panel or a longitudinal dataset tracking that child over time. Ditto the child with an id ending in 05. In row 5 of our data, though, you'll notice that we have a singleton. This child only appears in 2023. Have we failed to link this child? Well, no. If we look at the exited variable, this child has a positive value. Meaning that they exited foster care in 2023 and didn't re-enter in 2024. So in all likelihood, we shouldn't observe a linked record for that child. [ONSCREEN CONTENT] # We can examine our raw linkage rates by state and year: # What proportion of children with records in one year # also have a record in the other year? link_success <- function(df) { + df |> + arrange(stfcid, fy) |> # important to do whenever you use lead()/lag() + mutate(linked = if_else(fy == 2023 & stfcid == lead(stfcid) | + fy == 2024 & stfcid == lag(stfcid), + 1, + 0)) |> # create a linkage indicator + group_by(st, fy, linked) |> # group to count + summarize(n = n()) |> # count + group_by(st, fy) |> # regroup + mutate(linked_prop = n/sum(n)) |> # calculate linkage proportion + filter(linked == 1) |> # keep only success rates + ggplot(aes(x = linked_prop, y = st, color = factor(fy))) + # visualize + geom_point() + + scale_x_continuous(limits = c(0,1)) + } ac |> + link_success() `summarise()` has regrouped the output. ℹ Summaries were computed grouped by st, fy, and linked. ℹ Output is grouped by st and fy. ℹ Use `summarise(.groups = "drop_last")` to silence this message. ℹ Use `summarise(.by = c(st, fy, linked))` for per-operation grouping instead. [Alex Roehrkasse] We can examine our raw linkage rates by state and year asking, for example, what proportion of children with records in one year also have a record in the other year. Let's define a function and call it our link success function, where we'll evaluate the success of our links. This function will take only one input, a data frame, and we'll do a number of things to that data frame. We'll first arrange it by unique id. We'll generate a linked indicator indicating whether or not that record was linked to the subsequent or preceding year. We'll then do some grouping and summarizing, generating a variable that tells us the proportion of records that were linked. And then we'll visualize that success. We'll define the function, and then apply, using a pipe, apply that function to the stacked AFCARS object. [ONSCREEN CONTENT R CODE IMAGE 2] Graph of proportion of children with records in one year that also have a record in the other year. A plot with x-axis state and y-axis linked_prop. Two years, 2023 and 2024, are represented with red and blue dots respectively, and for all states the 2023 and 2024 dots all have linked_prop values between 0.6 and 0.75. [Alex Roehrkasse] Looking at the figure in the bottom right hand quadrant, we see that In general, between two-thirds and three-quarters of records were linked. It's interesting to see that this is fairly consistent across states and fairly consistent across years. But truthfully, this doesn't actually tell us very much. Because the truth is that children can fail to have repeated records for two completely different reasons. On the one hand, they might actually have only been in foster care for one of the two years. Or, they might have been in foster care for both years, but we failed to link them. We have to be a little more careful and creative if we want to actually evaluate the successfulness of our link. [ONSCREEN CONTENT] # But this doesn't tell us much, because children can fail # to have repeat records for two completely different reasons: # (1) they were actually only in foster care for 1 of 2 years, or # (2) they were in foster care both years but we failed to link them. # We can use additional variables to evaluate our link quality, # e.g. information about whether a child was in care at the # beginning or end of the reporting period. # If we limit our data to children whose records # we should be able to link, we can get a better sense of our # linkage quality. ac |> + filter((fy == 2023 & inatend == 1) | + (fy == 2024 & inatstart == 1)) |> + link_success() `summarise()` has regrouped the output. ℹ Summaries were computed grouped by st, fy, and linked. ℹ Output is grouped by st and fy. ℹ Use `summarise(.groups = "drop_last")` to silence this message. ℹ Use `summarise(.by = c(st, fy, linked))` for per-operation grouping instead. [Alex Roehrkasse] We can do this using additional variables to evaluate our link quality. Most helpful here is information about whether a child was in care at the beginning or the end of the reporting period. If we limit our data to children whose records we should be able to link, we can get a better sense of how well we've done. Consider this example. Take our stacked data but then pull out only those children who were in foster care at the very end of 2023 or in foster care at the very start of 2024. In principle, unless these children transitioned in or out of care on the very single day when the fiscal year transferred over, we should see these children in both years. This is a strong test of our link success rate. So we'll filter out only these children and then apply our link success function to that data set. [ONSCREEN CONTENT R CODE IMAGE 3] Graph of proportion of children who were in foster care at the very end of 2023 or in foster care at the very start of 2024. A plot with x-axis state and y-axis linked_prop. Two years, 2023 and 2024, are represented with red and blue dots respectively, and for all states the 2023 and 2024 dots all have linked_prop values between 0.9 and 1. [Alex Roehrkasse] What we see now is that our link rate is much higher, much closer to 100%. But in many cases, not quite 100%. This is a stronger indication of false negatives. Scenarios where we should be able to identify a link. But have failed to do so. [ONSCREEN CONTENT] # Another concern, however, is false positives, # i.e. links that don't exist but which we mistakenly observe. # We can use information like date of birth to verify # that linked children have consistent time-invariant attributes. ac |> + arrange(stfcid, fy) |> + mutate(linked = if_else(fy == 2023 & stfcid == lead(stfcid) | + fy == 2024 & stfcid == lag(stfcid), + 1, + 0), # create a linkage indicator + dob_match = if_else(fy == 2023 & dob == lead(dob) | + fy == 2024 & dob == lag(dob), + 1, + 0)) |> # create birthdate match indicator + filter(linked == 1) |> # keep only successful links + group_by(st, fy, dob_match) |> + summarize(n = n()) |> + group_by(st, fy) |> + mutate(false_pos_prop = n/sum(n)) |> # calc. prop. with matched DOB + filter(dob_match == 1) |> # keep only success rates + ggplot(aes(x = dob_match, y = st, color = factor(fy))) + # visualize + geom_point() + + scale_x_continuous(limits = c(.9,1)) `summarise()` has regrouped the output. ℹ Summaries were computed grouped by st, fy, and dob_match. ℹ Output is grouped by st and fy. ℹ Use `summarise(.groups = "drop_last")` to silence this message. ℹ Use `summarise(.by = c(st, fy, dob_match))` for per-operation grouping instead. # Looks good in this toy example, but don't count on it with real data! [ONSCREEN CONTENT R CODE IMAGE 4] Graph of proportion of child records we have a linked and have a matching date of birth. A plot with x-axis state and y-axis linked_prop. Two years, 2023 and 2024, are represented with red and blue dots respectively, and for all states the 2023 and 2024 dots all have linked_prop values at 1. [Alex Roehrkasse] Another concern is false positives. Links that we don't that don't actually exist, but which we mistakenly observe. We can use things, as I described in the slide deck, like date of birth, to verify that children who we link have consistent time-invariant attributes. A strong candidate is their date of birth. When we link two children, do their dates of birth match? This chunk of code again generates a link indicator and a date of birth match indicator. It then does some grouping and summarizing and some visualizing to report: how many of our linked children have matching dates of birth? Very happily, it seems in this case to be 100% or nearly 100%. That's great. But remember, this is a toy example. I wouldn't count on this in real data. I'd strongly encourage you to do these kinds of quality checks, checks for both false negatives and false positives whenever you're linking data. [ONSCREEN CONTENT] ###################################### # 3. LINKING BY JOINING: MICRO DATA # ###################################### # Research question: # What are the maltreatment histories of children # aged 0-3 (NCANDS) placed into foster care in FY2023 (AFCARS)? ## PREPARING DATA FOR ONE-TO-ONE JOIN ## # Most essential for linking by joining one-to-one are three things: # 1) In each linking data set, a row corresponds to the unit of analysis. # 2) Any and all linking variables are identically named and encoded. # 3) Any and all non-linking variables are differently named. # First, let's use linking by stacking using NCANDS to # generate two maltreatment history measures: # total reports and total substantiated reports. start <- Sys.time() nc_link <- nc |> filter(rptdt %between% c('2019-10-01','2023-9-30')) |> # define sampling frame mutate(sub = case_when(rptdisp == 'Substantiated' ~ 1, is.na(rptdisp) ~ NA, T ~ 0)) |> # create substantiation indicator group_by(stfcid) |> # group by unique child ID summarize(nsub = sum(sub, na.rm = T), # count substantiated reports nrep = n(), # count all reports #maxdt = max(rptdt), # computationally intensive .groups = 'drop') end <- Sys.time() end - start [Alex Roehrkasse] Okay, now we start to arrive at our final example, the most powerful but also most challenging example, and that's linking by joining, particularly in the case of microdata. Remember our research question. What are the maltreatment histories of children aged 0 to 3, placed into foster care in fiscal year 2023. This is a question that can only be answered using linked microdata. In order to do this though, we need to prepare our data for a one-to-one join. Most essential for linking by joining, one to one, are three things. In each linking data set a row needs to correspond to the unit of analysis. Any and all linking variables need to be identically named and identically encoded. Lastly, any and all non-linking variables need to be differently named. Let's prepare our data. First, we'll use linking by stacking using the NCANDS to generate two maltreatment history measures, the total number of reports that a child experiences and the total number of substantiated reports. Note here that I'm going to set a system timer to track how long this chunk of code takes to run. I'll take our stacked NCANDS data and create a new object. That we're using to prepare for our link. We'll again explicitly sample on historical time, generate a substantiation indicator, and then most critically, group by our unique child identifier. Tell R: understand that each distinct value of this variable corresponds to a distinct person, and then do calculations over each of those distinct people. Namely, for each distinct child, calculate two variables. Counting the number of substantiated investigations and the number of all investigations experienced by each child. Ungroup our data, and output the result. Note that I've commented out another line here, where we generate another variable, which is the date of the last report that that child experienced. I've commented it out here because it's computationally intensive. When I run this chunk of code on my little laptop, you'll notice that it takes a few seconds to run. [ONSCREEN CONTENT] start <- Sys.time() nc_link <- nc |> + filter(rptdt %between% + c('2019-10-01','2023-9-30')) |> # define sampling frame + mutate(sub = case_when(rptdisp == 'Substantiated' ~ 1, + is.na(rptdisp) ~ NA, + T ~ 0)) |> # create substantiation indicator + group_by(stfcid) |> # group by unique child ID + summarize(nsub = sum(sub, na.rm = T), # count substantiated reports + nrep = n(), # count all reports + #maxdt = max(rptdt), # computationally intensive + .groups = 'drop') end <- Sys.time() end - start Time difference of 7.028203 secs #write_rds(nc_link, paste0(data,'nc_link.rds')) nc_link <- read_rds(paste0(data,'nc_link.rds')) # reads pre-processed data [Alex Roehrkasse] About 7 seconds. When we add this argument, it takes several minutes and we don't have time for that today. So what we'll do instead is read in a pre-processed object that includes this max state variable. I'll explain in a minute why that's of interest. Let's now check the output of this object. Something looks a little bit wrong. [ONSCREEN CONTENT] # Let's check the output. Something looks amiss: # one observation has an implausible number of reports. # This arises from a missing value for the ID variable. # We'll ignore this for the purposes of demonstration, but # a careful analysis would inquire into the # causes and consequences. nc_link |> count(nrep) # A tibble: 18 × 2 nrep n 1 237029 2 66065 3 26751 4 12086 5 5795 6 2892 7 1276 8 666 9 338 10 165 11 74 12 36 13 13 14 12 15 3 16 2 17 1 13545 1 [Alex Roehrkasse] When we count up the number of observations having unique values of the number of report variables. We see that most children have one report. A large number of children have a small number of reports. But one child seems to have had 13,000 reports in their first three years of life. That's obviously implausible. [ONSCREEN CONTENT] nc_link |> count(nsub) # A tibble: 13 × 2 nsub n 1 0 209413 2 1 112666 3 2 22118 4 3 6241 5 4 1911 6 5 609 7 6 167 8 7 44 9 8 22 10 9 10 11 10 2 12 11 1 13 1822 1 [Alex Roehrkasse] Similarly, when we count the number of substantiated reports, we see that one child has 1,800 substantiated reports. What's going on? [ONSCREEN CONTENT] nc_link |> filter(nrep > 20 | nsub > 20) # A tibble: 1 × 4 stfcid nsub nrep maxdt RI 1822 13545 2023-09-23 [Alex Roehrkasse] If we pull out only those rows that have very high values of these counts, we see that there's one offending row. it comes from Rhode Island and the Stfcid variable which should have a two letter state code followed by a long string of numbers, has no numbers. It seems to be the case that records from this state or a small number of records from this state were not assigned unique child identifiers. And so a large number of children appear as a single child. This is obviously a problem, and one we would want to address in any research setting. We'll skip over it for now, just for demonstration purposes. [ONSCREEN CONTENT] head(nc_link) # A tibble: 6 × 4 stfcid nsub nrep maxdt CT000410019186 1 1 2022-04-08 CT000410056939 1 1 2021-01-23 CT000410059634 0 1 2020-01-23 CT000410072065 0 1 2020-02-08 CT000410513558 0 1 2020-12-23 CT000410515973 1 1 2021-02-08 [Alex Roehrkasse] When we look at this object, which we've prepared for linkage, we note that we now have a data set where each row corresponds to a unique child. And we have three variables of interest, the number of substantiated reports, the number of reports, and the date of last report. Let's now prepare our AFCARS data, our main sample. [ONSCREEN CONTENT] # Let's prepare our AFCARS data. # We want to select only those children aged 0-3 # entering foster care in FY2023. ac_link <- ac |> + filter(dob %between% + c('2019-10-01','2023-9-30') & # define sampling frame + entered == 1 & + fy == 2023) head(ac_link) fy st recnumbr dob totalrem rem1dt latremdt ageatlatrem 1: 2023 MA 1109152070 2020-12-15 2 2021-02-23 2023-03-07 2 2: 2023 MA 1109160599 2020-09-15 1 2022-11-28 2022-11-28 2 3: 2023 MA 1109208253 2020-09-15 1 2022-11-10 2022-11-10 2 4: 2023 MA 1109211241 2020-12-15 1 2023-05-05 2023-05-05 2 5: 2023 MA 1109230078 2021-01-15 1 2023-02-01 2023-02-01 2 6: 2023 MA 1109230986 2020-08-15 1 2023-04-10 2023-04-10 2 dodfcdt entered exited inatend inatstart stfcid version 1: 2023-08-22 1 1 0 0 MA001109152070 NA 2: 1 0 1 0 MA001109160599 NA 3: 1 0 1 0 MA001109208253 NA 4: 1 0 1 0 MA001109211241 NA 5: 1 0 1 0 MA001109230078 NA 6: 1 0 1 0 MA001109230986 NA # Note that: # 1) Each row in each data frame is a child. # 2) Our linking variable stfcid is similarly named and encoded. # 3) No other variable names match. [Alex Roehrkasse] This is a little more straightforward. We take our AFCARS stacked data. We pull out only those children who are aged 0 to 3 in 2023 and who entered foster care in that year. When we examine this data object, we see that we have a larger number of variables but three things hold the three things we needed to be the case to do a one-to-one join. First, each row in the data set is a child. Each row has a unique value of this STFCID variable. Second, in both of the linking data sets, our linking variable is similarly or identically named and encoded. Lastly, no other variable names in these two datasets match. With that verified, we're now ready to join. [ONSCREEN CONTENT] # Now lets link by joining! # Although R will try to guess your linking variable, # it's always prudent to specify it explicitly. dlink <- ac_link |> + left_join(nc_link, + join_by(stfcid)) [Alex Roehrkasse] We'll use R's left join command to join the two data sets. R thinks it's pretty smart, and if you don't tell it what your linking or joining variables are, it will try to guess them, but it often guesses wrong. And so I want to strongly encourage you always to specify your linking variables, your joining variables explicitly. We'll tell R to take AFCARS and to it, left join our NCANDS data. And in doing so, join it by our unique child identifier, the single variable that uniquely identifies each row. [ONSCREEN CONTENT] # Note a few things: nrow(dlink) == nrow(ac_link) [1] TRUE ncol(dlink) == ncol(ac_link) + ncol(nc_link) - 1 # one linking variable [1] TRUE [Alex Roehrkasse] Notice that we now have a new linked data object in our environment. And note a few things about that object. First, st the number of rows in the new linked data object is exactly equal to the number of rows in our AFCARS object. That's because we left joined NCANDS to AFCARS. We kept all and only those records which were in AFCARS. Notice also that the number of columns in this object is exactly equal to the sum of the columns in our two joining data sets minus the number of linking variables we used. [ONSCREEN CONTENT] # We have a 7% failed link rate, # or false negative rate (i.e. is.na(nrep)) dlink |> + count(nrep) |> + mutate(pct = n/sum(n)*100) nrep n pct 1 1266 47.08069914 2 562 20.89996281 3 334 12.42097434 4 165 6.13611008 5 86 3.19821495 6 47 1.74786166 7 23 0.85533656 8 11 0.40907401 9 2 0.07437709 10 2 0.07437709 11 2 0.07437709 NA 189 7.02863518 [Alex Roehrkasse] We can examine the failure rate for our data linkage. Notice that 7% of records have a missing value for our maltreatment history variables. This arises from the fact that we failed to link records in those cases. [ONSCREEN CONTENT] # Note that we might also be concerned about report timing. # Because we include all maltreatment reports across FY2023, # for some children, later reports may represent # subsequent instances of maltreatment, and should be # subtracted out from our maltreatment history measures. dlink |> + filter(maxdt > dodfcdt) |> # report follows foster care discharge + select(stfcid, maxdt, dodfcdt) |> + nrow() [1] 51 [Alex Roehrkasse] A slightly more detailed point, we might also be concerned about report timing. This is where that max date variable becomes useful. Because in this case we've included all maltreatment reports across Federal fiscal year 2023, for some children later maltreatment reports may actually represent subsequent instances of maltreatment, rather than preceding instances of maltreatment, and should be subtracted out from our maltreatment history measures whenever we're trying to explain foster care placement. What do I mean by this? Well, take, for example, those records where the last maltreatment report actually followed a child's discharge from foster care. It seems that this is the case for 51 children in our linked data set. It's not obvious what the right thing to do in this case is, but it's important to note that sequencing is of pretty significant importance, warrants a lot of caution whenever we're linking data happening over time. That said, we can now explore our basic research question. What are the maltreatment histories of children aged 0 to 3 placed into foster care in federal fiscal year 2023? [ONSCREEN CONTENT] # Nevertheless, we can now explore our research question # What are the maltreatment histories of children # aged 0-3 placed into foster care in FY2023? dlink |> + pivot_longer(cols = c(nrep, nsub), names_to = 'type', values_to = 'n') |> + ggplot(aes(x = n)) + + geom_histogram(binwidth = 1) + + facet_wrap(~type) + + scale_x_continuous(breaks = 0:10) Warning message: Removed 378 rows containing non-finite outside the scale range (`stat_bin()`). [ONSCREEN CONTENT R CODE IMAGE 5] Histograms for nrep and nsub. Two side-by-side histograms labeled “nrep” and “nsub,” showing right-skewed count distributions with most values near 0–2 and fewer larger values up to about 8–9. [Alex Roehrkasse] With a little bit of re-pivoting and some visualization, we can easily generate a histogram where we see the distribution of maltreatment reports and the distribution of substantiated maltreatment reports experienced by children placed into foster care in 2023. This is the kind of analysis that you can only do using linked data, specifically data linking NCANDS and AFCARS at the record level. That's the end of my demonstration in R. So I'll go ahead and close R studio and return to our slide deck. In a moment, we're going to open things up for the Q&A, and I hope you'll enter questions into the Q&A box. [ONSCREEN CONTENT SLIDE 23] Questions? User support: NDACANsupport@cornell.edu Garrett baker: garrett.baker@duke.edu Alyssa Lindsey: Alyssa.lindsey@ucsf.edu [Alex Roehrkasse] But I want to remind you that we are always here to support your research and to answer any questions you have. If you have questions right now, again, please drop them in the Q&A. We'll use our remaining time to talk through them. But if questions arise an hour, a day, a month, a year later, don't hesitate to email me or to email our general user support line, where a staff member will triage your question and make sure that it gets directed to the person most expert in that topic. Thanks for your attention. I look forward to your questions. [Alyssa Lindsey] Thank you, Alex. What a great presentation. I'm gonna put in the chat for everyone our emails that you see here on the screen. And then we can go ahead and jump into Q&A. Again, the Q&A box is at the lower right-hand side of the screen. Please feel free to put your questions in there. We did have one question come through. "To use the micro data files is an institutional IRB from our university needed?" And the answer to that is yes. As part of applying to receive restricted data such as NCAND's child file, the investigator needs to obtain a notification from their Institutional Review Board or IRB that the proposed research project using that restricted data was submitted to the IRB and that the IRB a) found it was exempt or b) approved it by expedited or full review. And then the investigator will email that PDF of the notification to the NDACAN email that you see listed. And Andres also included the instructions for applying to receive restricted data on the website. So thank you so much for this question, and thank you, Andres, for your answer. And I see a question here. "Just wanted to double check, since I ran into this and it was mentioned this session. If children with open reports move states, there is currently no way to track those cases reliably?" [Alex Roehrkasse] This is a great question. Okay. So, it's not self evident that scenarios like this will be handled in the exact same way in every state in every year. My best intuition is that many of these cases will receive in the original state a final designation of "closed, no determination". So the open report may at some point be closed without the determination that the report will actually receive eventually. Let's say that report is then opened in another state and "closed, receiving a disposition". There's no way to link those two distinct investigations to attach the ultimate designation to the original open report preceding an interstate move. So I think ultimately the case looks quite a lot like any other case of an interstate move, and you can't reliably capture the disposition that that that investigation will ultimately receive, if at all. [Alyssa Lindsey] Thank you, Alex. Alright, we have around four minutes left until the top of the hour. Please feel free to put your questions in the Q&A. I also put a link to our frequently asked questions or FAQ page on our website if questions come up as you explore the data further or if you review this presentation again in the future. Well, it looks like we have answered all of the questions that have come in today. Again, if any other questions come up, like Alex says, please feel free to reach out to us using the email addresses on the screen. Thank you all so much for your attendance. Thank you, Alex, for the presentation, and we will follow up with these materials in the future. And I think, Alex, if you want to give a preview of next week. [ONSCREEN CONTENT SLIDE 24] Next week… Date: July 15th Topic: Linking NCANDS and AFCARS Instructor: Alex Roehrkasse [Alex Roehrkasse] Yeah, thanks again everyone for your attention and your time today. I enjoyed the opportunity to talk about this topic. A shameless plug for next week's presentation, we'll be talking about handling missing data. Linking being closely related to this question. So please, consider joining us next week to learn about missing data, a common, important issue with archive administrative data. Thanks, and have a great rest of your day. [VOICEOVER] The National Data Archive on Child Abuse and Neglect is a joint project of Duke University, Cornell University, University of California, San Francisco, and Mathematica. Funding for NDACAN is provided by the Children's Bureau, an office of the Administration for Children and Families. [Music]