Curiosity output: 9
Summer 2026
Research 4
Agricultural Labor Fragility as Food-System Risk: A Baseline Analysis Using the National Agricultural Workers Survey
Background
This project began in PUBH 6281: Analysis of Complex Survey Data, where I developed the initial research question, prepared the analytic dataset, wrote the SAS code, and conducted preliminary analyses using the National Agricultural Workers Survey. It has since developed into my MPH culminating experience under the guidance of Dr. Cleary.
The research examines labor-protection fragility among U.S. crop farmworkers as both a public health concern and a potential food-system risk. Rather than treating wages, employment benefits, and worker protections as isolated labor issues, the project considers how widespread workforce vulnerability may affect agricultural stability, production continuity, and the resilience of the broader food system.
The project uses public-use National Agricultural Workers Survey data from fiscal years 2020 through 2022. I developed a composite measure of labor-protection fragility and applied survey-weighted descriptive and regression methods to examine patterns across the agricultural workforce.
This work represents an important stage in my development as an epidemiologist. It has required me to move from a broad systems question to a reproducible analytic measure, manage and recode complex survey data, evaluate missingness, construct statistical models, and interpret the limits of the available public-use survey design information.
Partial materials
The project is currently in the analysis and manuscript-development stage. Public materials will be added selectively as the work is reviewed and finalized.



Background
Agricultural workers occupy an essential role in the U.S. food system, yet many farmworkers experience employment conditions that may limit access to wages, benefits, and basic labor protections. These conditions are not only individual labor concerns; they also have system-level implications for agricultural production, workforce stability, and food-supply continuity. When the workforce responsible for planting, harvesting, and sustaining crop production experiences high levels of economic and employment vulnerability, labor fragility can become a food-system risk.
Public and political discussion of agricultural labor often centers on immigration status, enforcement, and border policy. However, a systems-level public health lens can shift the focus from individual immigration debates toward the stability, protection, and sustainability of the agricultural workforce. This framing is important because food production depends on labor systems that are often vulnerable to wage instability, lack of employment protections, limited access to benefits, legal precarity, and regional policy pressures. Establishing a baseline measure of labor-protection fragility can therefore help identify structural vulnerabilities in the agricultural workforce and support future policy discussions grounded in workforce and food-system resilience.
This project is informed by the literature on precarious employment. Kreshpaj et al. describe precarious employment as a multidimensional construct that includes employment insecurity, income inadequacy, and lack of rights and protection. Using available National Agricultural Workers Survey public-use variables, this project operationalizes labor-protection fragility as a binary measure that captures income inadequacy and lack of rights/protection. A worker will be classified as having labor-protection fragility if they report at least one of the following indicators: below-minimum-wage earnings, lack of unemployment insurance, lack of workers’ compensation or related injury support, or lack of employer health support for work-related injury or illness.
The National Agricultural Workers Survey is an appropriate data source for this project because it is a nationally recognized survey of U.S. crop farmworkers and includes demographic, employment, wage, benefit, legal status, migrant status, and regional variables. This CE builds from analytic work completed in PUBH 6281: Analysis of Complex Survey Data, where the initial research question, analytic dataset, SAS code, and preliminary analyses were developed. The project has since been revised based on instructor feedback to improve the labor-protection fragility coding, use numeric SAS reference categories, and clarify the limitations of the available NAWS public-use survey design variables.
Specific Aims and Hypotheses
Aim 1: Develop and document a reproducible labor-protection fragility measure using NAWS public-use data from FY 2020–2022.
Aim 2: Estimate the weighted prevalence of labor-protection fragility and its component indicators among U.S. crop farmworkers interviewed in FY 2020–2022.
Aim 3: Examine whether migrant status is associated with labor-protection fragility among U.S. crop farmworkers in FY 2020–2022.
Hypothesis: Migrant workers will have a higher weighted prevalence and higher odds of labor-protection fragility than non-migrant workers after accounting for age, gender, legal status, education, and region.
These aims will be achieved by creating a labor-protection fragility composite from available NAWS public-use variables, estimating weighted descriptive statistics, comparing fragility prevalence by migrant status, and fitting weighted logistic regression models adjusted for relevant covariates.
Methods
Study Design
This project will use a cross-sectional analytic design with public-use data from the National Agricultural Workers Survey. The analysis will focus on respondents interviewed during fiscal years 2020–2022. Because NAWS is a repeated cross-sectional survey, this analysis will estimate population-level associations within the selected time period rather than individual-level change over time.
Population and Inclusion/Exclusion Criteria
For the NAWS component, the study population will include U.S. crop farmworkers represented in the NAWS public-use data. The analytic sample will be restricted to respondents interviewed in FY 2020, FY 2021, or FY 2022. Observations missing the NAWS sampling weight variable will be excluded. Respondents with insufficient information to classify the labor-protection fragility outcome will be excluded from analyses involving the outcome.
Data Sources
The data source will be the NAWS public-use dataset, specifically the NAWSPAD files for interview cycles 1–103. The public-use files contain questionnaire and created variables, including fiscal year of interview, migrant status, demographic characteristics, legal status, education, region, wage indicators, and benefit-related variables. The A–E and F–Y public-use files will be merged using the respondent identifier.
NAWS is adequate for this project because it provides a national public-use source for examining U.S. crop farmworker characteristics and employment-related conditions. The dataset includes sufficient variables to construct a labor-protection fragility measure and examine its association with migrant status while accounting for key demographic and socioeconomic covariates.
Variables
The dependent variable for the NAWS analysis will be a created binary labor-protection fragility measure. Respondents will be coded as having labor-protection fragility if they have at least one of the selected component indicators: below-minimum-wage earnings, lack of unemployment insurance, lack of workers’ compensation or related injury support, or lack of employer health support for work-related injury or illness. Component variables will be recoded so that a value of 1 consistently indicates the presence of a fragility indicator.
The primary independent variable will be migrant status. Covariates will include age, gender, legal status, education, and region. Education will be collapsed into four categories after reviewing the distribution of the original education variable: less than 6th grade, 6th–8th grade, 9th–12th grade/GED, and some college or higher. Legal status and region will be included because both may be associated with employment opportunities, access to protections, and labor-market conditions.
Statistical Analysis Methods
SAS will be used for the NAWS data management and analysis. Data management steps will include importing and merging the relevant NAWS public-use files, restricting the sample to FY 2020–2022, recoding missing and nonresponse values, creating the labor-protection fragility composite, and collapsing education categories. Initial descriptive checks will be used to confirm variable coding and missingness.
Univariate analyses will estimate the weighted prevalence of migrant status, demographic characteristics, legal status, education, region, individual fragility indicators, and the composite fragility outcome. Bivariate analyses will compare the weighted prevalence of labor-protection fragility between migrant and non-migrant workers. Multivariable survey-weighted logistic regression will be used to estimate the association between migrant status and labor-protection fragility, adjusting for age, gender, legal status, education, and region.
The revised SAS code will use numeric class and event statements to avoid errors related to long formatted category labels. The fragility composite will also use revised missing-data logic: respondents will be coded as having fragility if any observed component indicates fragility; respondents will be coded as not having fragility only when all four components are observed and none indicates fragility; otherwise, the composite outcome will be coded as missing.
The public-use NAWS files available for this analysis include the composite sampling weight, PWTYCRD. However, the downloaded public files do not include stratum, primary sampling unit, or replicate-weight variables. Therefore, weighted point estimates will be produced using PWTYCRD, but standard errors, confidence intervals, and hypothesis tests will be interpreted cautiously because the full complex survey design cannot be incorporated with the available public-use variables.
Human Subjects Protection Issues
This project will use publicly available or secondary datasets. The NAWS public-use data are de-identified and do not include direct respondent identifiers. No direct contact with human subjects will occur, and no identifiable private information will be collected by the student.