Automation and Affordability in U.S. Counties

Is the task composition of local work associated with what residents can afford?

Introduction

According to Gallup, 22% of American workers worried technology would make their job obsolete in 2023, up seven percentage points from 2021 (Saad 2023). Concerns about automation were evident years earlier. In 2017, the Pew Research Center reported that 72% of Americans surveyed worried about a future in which technology could perform many jobs currently done by humans (Smith and Anderson 2017). Underlying these concerns was automation, in which technology assumes tasks that were previously performed by human workers. These concerns have a clear economic basis. Job displacement can result in large and long-lasting reductions in worker earnings (Couch and Placzek 2010).

A study found that displaced workers’ earnings fell by more than 30% from 1993 to 2004, remaining about 15% lower six years later (Cozzi and Fella 2016). The burden of those earnings losses depended in part on local prices. As demonstrated by Moretti (2013), the same income could support different living standards in different cities in the United States. This variation also makes the geographic scale of measurement important. As a result, a single value representing an entire state or metropolitan area can obscure the variations in prices and household resources across different communities. For this reason, measuring these conditions locally provides a more realistic view of where households face greater economic strain (Curran et al. 2006).

Examining whether these concerns corresponded with local economic conditions required a measurable indicator of automation risk. Autor’s task based framework provided the basis for measuring automation risk through the tasks performed in local labor markets. Autor, Levy, and Murnane (2003) first demonstrated that computers replaced routine tasks while complementing non-routine ones, which made it possible to identify work more susceptible to automation. Acemoglu and Autor (2011) later expanded the framework by treating tasks as the production unit and different skill groups as their suppliers. Autor and Dorn (2013) then applied the framework to local labor markets, finding that areas concentrated in routine work experienced employment polarization.

Nevertheless, these studies generally examined wages and employment at geographic scales broader than the county. A wage describes what a job paid, while purchasing power describes what that pay could buy. That distinction left unresolved whether the task groups present in local employment are associated with household purchasing power at the county level.

Building on this task based framework, we investigated the relationship between county automation and purchasing power, which is defined as household income adjusted for variations in local price levels. The analysis used a county-level panel spanning from 2008 to 2023, excluding the year 2020, encompassing 11,983 observations across 840 counties. These counties collectively represented a substantial portion of the U.S. population, reaching approximately 84% (U.S. Census Bureau, Population Division 2023). County employment was organized into four task groups to capture differences in the kinds of work performed across local economies.

We conducted a panel regression analysis to investigate the relationship between purchasing power and county task groups, while controlling for factors such as poverty, unemployment, and population. Year indicators accounted for changes over time, and standard errors were clustered by county for statistical inference. Diagnostic checks favored a log specification, while random forest and neural network models assessed whether nonlinear relationships captured patterns missed by the regression model. The most notable finding was that a one percentage point shift from non-routine cognitive to non-routine manual work was associated with approximately $846 less purchasing power at the panel mean (95% confidence interval, ±$38). Task groups remained associated with purchasing power after accounting for poverty and unemployment. These findings have practical implications for policymakers and economic development agencies by helping identify counties that appear similar on conventional indicators but differ in what residents can afford.

Background

Autor developed a task-based approach to measure the variation in automation risk across different types of work. Initially, Autor, Levy, and Murnane (2003) distinguished between routine and non-routine tasks by demonstrating that technology impacts tasks differently based on their potential to be formalized into explicit procedures. Acemoglu and Autor (2011) further expanded his framework by treating tasks as production units and analyzing how workers with varying skills are distributed among them. This progression laid the conceptual foundation for the four task groups depicted in Figure 1, linking technological advancements to shifts in the demand for different types of work.

Building upon this framework, Autor and Dorn (2013) applied the task-based approach to local labor markets. Areas with higher initial concentrations of routine work experienced more pronounced employment polarization. Employment growth was observed at both the upper and lower ends of the wage distribution, while it declined in the middle. These patterns indicated that the local impacts of technological change varied depending on the types of tasks concentrated within each economy. This provided a foundation for exploring whether differences in county task groups were also linked to variations in purchasing power.

Figure 1: Occupations in the Task Framework. Every detailed occupation in the Standard Occupational Classification crosswalk behind the American Community Survey, placed in its assigned task group. Routine tasks follow explicit rules a machine can be programmed to carry out; cognitive tasks work with information rather than with physical objects and people. The axes name these two distinctions rather than measuring them: assignment follows Autor and Dorn (2013) and is categorical, so a point’s position inside its quadrant is scatter for legibility and carries no measured value. A readable subset of occupations is drawn larger and labeled. Community and social service and military occupations are assigned to no group and are excluded from the measure.

Once occupations are categorized based on their task content, it becomes easier to measure differences across local economies. The occupation taxonomy developed by Autor and Dorn (2013) served as the foundation for assigning employment to the four task groups used in this study. Their own measure scored each occupation continuously on a single routine task intensity index built from ratings of its task content. Acemoglu and Autor (2011) provided economic significance to these differences by treating tasks as units of production supplied by workers possessing varying skills. Technological advancements can shift labor demand across different types of tasks. These differences in county task groups provide a way to examine whether local work structures are associated with purchasing power.

Accurately measuring purchasing power requires accounting for differences in local price levels. Regional Price Parities (RPPs), published by the Bureau of Economic Analysis (BEA), offer a measure of these differences relative to the national average. A score of 100 in the RPP represents the national price level. Relative to this baseline, values above 100 indicate higher local price levels, while values below 100 indicate lower local price levels. The index combines prices for goods and services such as housing, food, transportation, medical care, education, and recreation. RPPs became an official BEA statistic in 2014, with estimates available back to 2008. However, BEA does not publish RPPs directly for individual counties. Instead, estimates are available for states, metropolitan areas, and nonmetropolitan portions of states, which can mask price differences among counties within the same region.

Multiple indicators collectively capture different aspects of local economic conditions. Poverty measures the proportion of residents living below an established income threshold. Unemployment measures the share of the labor force actively seeking work. Median household income represents the midpoint of the household income distribution, and RPPs compare local price levels with the national average. Each captures a different dimension of local conditions, and the four produce distinct geographic patterns across counties (Figure 2).

Figure 2: County Economic Conditions. County patterns in 2023 are shown for A) poverty rate in percent, B) unemployment rate in percent, C) median household income in dollars, and D) RPP as an index on which 100 is the national average price level. Each indicator captures a distinct dimension of local economic conditions, producing geographic patterns that do not fully overlap across measures.

Geographic scale presented an additional measurement challenge because broader geographic units could obscure economic differences between nearby counties. Local prices and household resources varied within the same labor market, while earlier studies often combined several counties into a single regional measure. This aggregation could mask meaningful local variation that remained visible at the county level. Curran et al. (2006) showed that accounting for geographic cost of living differences alters how places compare on economic wellbeing, which makes the choice of geographic unit consequential for any measure of household resources. By measuring purchasing power at the county level, we preserved more of this variation by incorporating both household income and local price levels. Additionally, we evaluated whether the association between purchasing power and county task groups was better represented in dollar terms or proportional terms. This approach addressed a gap in prior research, which examined task groups in relation to employment and wages without evaluating their association with purchasing power at the county level.

Data

Data sources and study period

This study integrated five federal datasets published by three agencies, summarized in Table 1, along with the number of records and variables provided by each source. Most sources reported data at the county level, and they were joined using the county Federal Information Processing Standards (FIPS) code and the year. RPPs were published at the metropolitan area level and required an additional matching step. Counties were first linked to their metropolitan areas using the Census Bureau Core Based Statistical Area (CBSA) delineation file and then matched to the corresponding BEA price measure (U.S. Bureau of Economic Analysis 2024; U.S. Census Bureau, Geography Division 2023).

Table 1: Federal Data Sources. The five federal datasets combined to build the county year panel. Each row names the source, the number of records retrieved, the variables it supplies, and the warehouse table it feeds.
Data Source Abbreviation Records Measures Destination Table
Small Area Income and Poverty Estimates Census SAIPE 50,283 median household income, poverty rate county_baseline
Bureau of Economic Analysis Regional Price Parities BEA RPP 884 county price level relative to the national average cbsa_rpp
Core Based Statistical Area delineation Census CBSA 387 county to metropolitan area crosswalk cbsa
American Community Survey 1 year estimates Census ACS 15,690 population, occupational employment by category county_baseline, county_task_exposure
Bureau of Labor Statistics Local Area Unemployment Statistics BLS LAUS 88,004 unemployment rate county_baseline
Scale: 3,143 counties per year, 15 years (2008 to 2023, excluding 2020), 47,140 county year universe, 11,983 rows in the analytical panel.
Records count rows retrieved from each source rather than analytical observations. The ACS publishes 1 year estimates only for areas above 65,000 residents, which reduces the universe to the 848 county panel.
All sources retrieved June 2026 through agency APIs or bulk file download.

The study spans 2008 through 2023 but excludes 2020, because the American Community Survey (ACS) did not publish data due to disruptions of data collection during the COVID-19 pandemic (U.S. Census Bureau, American Community Survey Office 2024). Those estimates served as the occupational employment counts used to construct the task groups. Over the subsequent fifteen years, the Small Area Income and Poverty Estimates (SAIPE) program provided an initial sample of 47,140 county-year observations (U.S. Census Bureau, Small Area Estimates Branch 2024).1

Analytical sample

The initial sample restriction stemmed from the coverage of the ACS occupational estimates. The Census Bureau published ACS 1 year occupational estimates exclusively for counties with populations of at least 65,000. Consequently, counties below this threshold lacked the employment counts necessary to construct the four task groups. Consequently, applying this restriction reduced the sample to 12,087 county-year observations across 848 counties. A subsequent restriction mandated the inclusion of complete model covariates, resulting in the removal of 104 observations with missing unemployment rates.2 However, population, median household income, and poverty rate data were complete throughout the sample.

The final analytical sample comprised 11,983 county-year observations from 840 counties. All models presented in Section 4 and Section 5 used these same observations. To ensure unbiased machine learning evaluation, counties were exclusively included in either the training or test set, preventing observations from the same county from appearing in both partitions. Section 4 provides a detailed description of the evaluation procedure.

The population threshold imposed a limitation on the scope of the results. While the 840 retained counties accounted for approximately 84% of the United States population, they represented a minority of all counties. Consequently, rural counties were underrepresented because the retained counties were generally more populous and metropolitan. The results were applicable to counties above the ACS publication threshold rather than the entire nation, a limitation discussed in Section 6.

Construction of task group measures

Following the framework introduced in Section 2, we classified broad occupational categories in the ACS 1 year estimates into four task groups, applying the task type distinctions of Autor, Levy, and Murnane (2003) and the occupation taxonomy of Autor and Dorn (2013). For instance, clerical and sales occupations were categorized as routine cognitive tasks, while service occupations were classified as non-routine manual tasks. Subsequently, we summed employment within each task group for every county and year.

\[ T_{g,c,t} = \sum_{o \in g} E_{o,c,t} \tag{1}\]

Here, \(T_{g,c,t}\) represented total employment in task group \(g\) for county \(c\) in year \(t\). The sum ran over the occupational categories \(o\) assigned to that group, and \(E_{o,c,t}\) represented employment in a single category. Since the calculation used employment counts without task intensity weights, each total represented the number of jobs in that group. Scoring each occupation continuously, as Autor and Dorn (2013) did, would have required employment counts for detailed occupations, which the Census Bureau did not publish for counties. Every job within a broad category therefore carried equal weight, and variation in task content inside those categories went unmeasured. Dividing each group’s total by the total across all four groups, as shown in Equation 2, converted these totals into proportions between zero and one that summed to one.

\[ G_{g,c,t} = \frac{T_{g,c,t}}{\sum_{g'} T_{g',c,t}} \tag{2}\]

Here, \(G_{g,c,t}\) represented the proportion of employment in task group \(g\), and \(g'\) indexed the four groups summed in the denominator.

These proportions assigned counties of varying sizes to a common scale. Since the four proportions added up to one, non-routine cognitive work was excluded from the regression as the reference group, and the remaining coefficients were interpreted in relation to it. We maintained the four task groups as separate entities because a single routine intensity index would obscure differences between types of work. For instance, a shift away from clerical work and a shift away from assembly work could otherwise appear identical. Section 4 provides an explanation of how the task groups were incorporated into the models.

Additionally, the Census Bureau revised its occupational classification between 2009 and 2010. While some category boundaries changed, their assignment to the four task groups remained consistent. Section 4 reports the corresponding robustness checks.

Construction of the purchasing power measure

Median household income was used because it represented the income of a typical household better than the mean, which was more sensitive to high incomes in a positively skewed distribution (Chiripanhura 2011). It described what households received rather than what they could buy.

The outcome was measured as purchasing power to account for variations in what household income could purchase across different counties. We calculated it by adjusting the county median household income from SAIPE (U.S. Census Bureau, Small Area Estimates Branch 2024) using RPPs. These provided local price levels relative to the national average (U.S. Bureau of Economic Analysis 2024). Since the index was expressed as a percentage, dividing it by 100 converted it into a price multiplier. Subsequently, we divided the median household income by this multiplier to determine purchasing power, as illustrated in Equation 3.

\[ A_{c,t} = \frac{\text{Median household income}_{c,t}}{\text{RPP}_{c,t} / 100} \tag{3}\]

Here, \(A_{c,t}\) represented purchasing power for county \(c\) in year \(t\), and \(\text{RPP}_{c,t}\) represented the RPP for that same county and year.

For instance, consider a county with a median household income of $60,000 and an RPP of 120. This means that its income could buy approximately what $50,000 would buy at the national average prices. This adjustment made household income more comparable across counties by accounting for variations in local prices. However, since BEA did not publish separate RPPs for counties outside metropolitan areas, those counties were assigned the corresponding state-level parity. Within the analytical panel, 45.5% of county-year observations used a metropolitan parity, while 54.5% used the state measure. Consequently, more than half of the panel relied on a state average that did not account for price differences within the state.

Inflation and year indicators

RPPs were adjusted for variations across regions but not for changes in the national price level over time. Consequently, purchasing power was reported in nominal U.S. dollars. A national deflator would apply the same annual adjustment to every county. In the log specification, this adjustment would be incorporated into the year indicators, leaving the task group and control coefficients unchanged. Additionally, the year indicators captured other changes that occurred simultaneously across counties within a year, making it impossible to interpret their coefficients solely as inflation. Section 4 provides an explanation of how the year indicators were incorporated into the model, while Section 6 addresses this limitation.

Control variables

The analysis incorporated poverty rate, unemployment rate, and population as county-level covariates. Poverty and unemployment captured the economic distress mentioned in Section 2, enabling the analysis to determine if task groups provided information beyond these measures. Median household income directly entered the purchasing power outcome, so it was not included as a control variable. Population accounted for variations in county size, which could otherwise affect the estimates.

Figure 3 displayed the distributions of poverty rate, unemployment rate, population, and median household income across the same task group panel used in Figure 9. All four variables were right skewed, but population was the most extreme and therefore entered the analysis in logarithmic form. Poverty rate, unemployment rate, and median household income remained in their original units. This supported the log respecification in Section 4 as a response to the outcome’s own distribution rather than a transformation applied across all variables.

Figure 3: Distribution of the Control Variables. Across the 12,087 county year task group panel, shown as violin plots with an embedded box marking the median, interquartile range, and outliers. A) Poverty rate and B) unemployment rate are both right skewed, with unemployment the more extreme of the two; both are shown in percent. C) Population is plotted on a logarithmic scale because it spans several orders of magnitude, which is also why the models use its log form. D) Median household income, in dollars, is right skewed with a moderate upper tail.

Data architecture and organization

The data architecture separated raw source records from the tables used for analysis. Raw retrievals and intermediate results were stored in a PostgreSQL data lake containing 41 tables, which remained unnormalized to maintain traceability. Processed data were then written to an analytical warehouse containing eleven tables. In this warehouse, county identifiers were standardized, records were restricted to the study period, jurisdictions outside the panel were removed, and foreign key constraints were enforced.

Figure 4 provided a diagram that summarized the path from federal sources through the data lake and warehouse to the analysis-ready tables.

Figure 4: Data Pipeline Architecture, Source to Analysis Ready Output. Federal sources feed a raw data lake, which feeds the normalized warehouse, which feeds the analysis ready county year table. Node fill marks the publishing agency. Warehouse tables list their columns with primary and foreign keys marked, and the four task group proportions marked as generated columns derived from the group totals. Three supporting warehouse tables not used in the models are omitted.

The warehouse is organized around a county table, which is keyed by FIPS code and linked to the state as a reference table. Three analytical tables join to each county by FIPS code and year. The county_baseline table contains population, median household income, poverty rate, and unemployment rate. The county_task_exposure table contains the four task group totals defined in Equation 1, while the county_affordability table contains the purchasing power outcome.

Analysis

The unadjusted relationship

The analysis commenced by examining the unadjusted relationship between task groups and purchasing power. Across the panel, the two manual groups exhibited the most pronounced negative relationships, while non-routine cognitive tasks showed a positive correlation, and routine cognitive tasks remained relatively flat (Figure 5). None of these relationships were strong enough to stand alone, prompting the subsequent adjustment analysis.

Figure 5: Purchasing Power Against Each Task Group. County year purchasing power, in dollars at national average prices, plotted against each task group’s value as a percent of county employment, before any controls. Faint points are individual county years; solid points are means within twenty equal count bins. Each panel keeps its own x axis range, since the four groups span different proportions of county employment. The manual groups decline most steeply and non-routine cognitive rises, an ordering the adjusted estimates preserve.

Variance decomposition

Purchasing power and task groups exhibit variations both across counties and within a county over time, each carrying distinct implications. Each group’s variation divides into these two components (Figure 6). The manual groups show almost complete variation between counties, which contributes to the durability of the purchasing power differences they represent. In contrast, routine cognitive is the only group with substantial within-county movement.

Figure 6: Between County vs. Within County Variation by Task Group. Computed on the panel restricted to 2010 onward, with each year’s cross county mean removed. Bars give each component as a percent of that task group’s total variance, so the two components sum to 100 percent. Three of the four groups appear, since non-routine cognitive is the reference group dropped from the regression, the four proportions summing to one. The manual groups vary almost entirely between counties; routine cognitive is the one group with substantial within county movement.

The Census occupation coding change between 2009 and 2010, as described in Section 3, inflates the apparent variation in routine cognitive occupations within counties. This inflation arises from the coding of the source data rather than any inherent property of the counties. The figure is computed on a panel restricted to 2010 onward, with each year’s cross-county mean removed.

Model specification

Non-routine cognitive tasks served as the reference group because the four task group proportions summed to one. Each remaining coefficient therefore represented the change in purchasing power associated with a shift from non-routine cognitive work to that group, while holding the other variables constant. This made the coefficients interpretable as comparisons between task groups rather than as independent changes in employment. Task group values ranged from zero to one, and coefficients were divided by 100 when reported as percentage point changes.

Year indicators captured the shared conditions across all counties within a specific year, encompassing events like the 2008 recession, its subsequent recovery, and the national price movements described in Section 3. By including these indicators, we prevented attributing national changes to variations in county task groups. Since the year coefficients absorbed all common changes within a year, they were treated as nuisance parameters rather than evidence of increasing purchasing power. The year coefficients traced this pattern across the panel (Figure 7).

Figure 7: Year Effects Relative to 2008. Estimated year coefficients with 95% intervals, each the difference from the 2008 baseline shared by every county in that year. A) The level specification, in dollars of purchasing power. B) The log specification, as a percent difference. The gray band between 2019 and 2021 marks 2020, which the ACS did not publish, so the line spans a year with no estimate rather than passing through one. The series carries national price drift together with every other movement common to a year, which is why it cannot be read as a measure of inflation.

Standard errors were clustered by county because repeated observations within the same county over fifteen years were not independent. A Durbin Watson statistic of 0.496 indicated this dependence. Clustering allowed observations within a county to be correlated, which prevented the uncertainty around the estimated associations from being underestimated. The estimated coefficients are reported in Section 5.

Collinearity and the reference group

The unnormalized employment totals of Equation 1 are not suitable for regression analysis. A variance inflation factor (VIF) quantifies the extent to which a coefficient is affected by its correlation with other inputs, and a value exceeding 10 is considered a serious concern. The four employment totals exhibit factors ranging from 13 to 27, as they all represent employment counts that are proportional to county size and therefore exhibit a strong correlation. When fitted on these totals, the coefficients were unstable, and their standard errors were notably large compared to their magnitudes, which are typical indicators of multicollinearity. The proposed specification addresses this issue twice (Figure 8). Normalizing the data by the four group total, as described in Section 3, eliminates the shared scale, while using non-routine cognitive as the reference group, as established above, removes the constraint that the four proportions must sum to one. Consequently, every factor in the estimated specification falls within the range of 1.2 to 1.6.

Figure 8: Variance Inflation Factors by Task Specification. VIFs for each task group, on a linear scale, computed two ways. The first uses the raw employment counts, the second the group proportions that actually enter the regression. Dotted and dashed vertical lines mark the conventional thresholds of 5 and 10. Every unnormalized total sits well above 10; every estimated factor sits near one. An asterisk marks non-routine cognitive as the reference group, which has no estimated factor because it is dropped from the regression, since the four proportions sum to one.

Specification diagnostics

Distribution of the analytical variables

The analysis next delved into the distributions of the variables used in the specification. Purchasing power and the four task groups spread across the panel with skewness and outliers that guided the diagnostic checks that followed (Figure 9). Purchasing power exhibited right skewness, with a skewness coefficient of 1.17 and a central tendency near $57,000. A long upper tail extended beyond the center of the distribution, so the extremes carry most of the information about model fit.

Figure 9: Distribution of Purchasing Power and Task Groups. Violin plots across the 12,087 county year panel, before the 104 rows missing an unemployment rate are dropped, each with an embedded box marking the median, interquartile range, and outliers. A) Purchasing power, in dollars at national average prices, right skewed with a long upper tail and a heavy cluster of high end outliers. B) The four task groups on a shared percentage scale, restricted to 5% to 65% so the distribution bodies stay legible. Non-routine cognitive holds the highest concentration of county employment.

Residual and quantile diagnostics

A linear model assumes a constant rate of association across the data range and symmetric, evenly distributed errors around the fit. Three diagnostics test these assumptions: residuals plotted against fitted values, a quantile-quantile (QQ) plot of the residuals, and the variance inflation factors already reported. The raw scale model shows clear systematic curvature in the residuals and a marked upper tail departure in the QQ plot (Figure 12, Panels A and B). The model underestimates both the lowest and highest purchasing power counties while fitting the middle well, and the upper tail departure indicates a residual skew of 1.01.

The level model still provides the interpretable dollar estimates we report. However, its own diagnostics reveal that the data exhibit structure that a straight line in dollars cannot represent. The curvature in the residual smoother and the widening spread with fitted values suggest a specific failure in the fixed slope specification, and each subsequent model relaxes the assumption behind it.

Influential points

A fourth diagnostic, Cook’s distance, evaluates whether any single county year disproportionately influences the level specification’s coefficients. None of the observations exceed the conventional threshold of 1.0, so no point warrants further investigation. The highest value recorded is 0.009. Every county year appears in the plot of leverage against standardized residual, with dashed curves tracing constant Cook’s distance (Figure 10). The conventional 0.5 and 1.0 contours fall well outside the plotted range, since reaching 1.0 at the highest observed leverage would require a standardized residual near 40. The labeled points are the four county years with the largest standardized residuals. Loudoun and Stafford County, Virginia, are both high purchasing power counties located in the Washington D.C. exurbs. Williamson County, Tennessee, is a high purchasing power Nashville exurb. Apache County, Arizona, a low purchasing power county situated on the Navajo Nation, sits among the marked points at a lower residual. These counties are the same high-end outliers already visible in the purchasing power distribution (Figure 9).

Figure 10: Residuals Against Leverage With Cook’s Distance Contours, Level Specification. Every county year in the panel, plotted by leverage (its hat value, unitless) and standardized residual (in standard deviations). The curves trace constant Cook’s distance, mirrored above and below zero, each branch carrying its own value. The conventional 0.5 and 1.0 contours fall far outside the observed data, so the two drawn levels are the 99th percentile cutoff used for the sensitivity refit and the observed maximum. The fifteen most influential county years are marked and the four largest standardized residuals labeled by county and year, through leader lines because the points sit too close together to label in place. Several counties appear in more than one marked year, so the marked points outnumber the labels.

No single point threatened the overall fit, so the more informative check was to refit the model after removing the most influential 1% of the panel, 120 county year observations. Table 2 reported the results. The task group coefficients changed by single digit to low double digit percentages, while the poverty coefficient remained relatively stable. This indicated that the main associations were not driven by a small number of extreme counties.

Unemployment rate was the exception. As reported in Section 5, it was the smallest and least precisely estimated control, and its coefficient changed by approximately one third. A second stability check examined whether the main estimates were consistent over time. The model was refit separately for the pre pandemic period from 2010 to 2019 and the post pandemic period from 2021 to 2023, with Section 5 reporting the stability of the coefficients across the two periods.

Table 2: Coefficient Sensitivity to Influential Points. Level specification coefficients, full sample versus excluding the most influential 1% of county year observations by Cook’s distance.
Variable Full Sample Excluding Top 1% Change (%)
Routine Cognitive −$85,791 −$75,273 −12.3
Routine Manual −$70,097 −$62,441 −10.9
Non-Routine Manual −$104,139 −$98,410 −5.5
Poverty Rate −$1,605 −$1,616 0.7
Unemployment Rate $206 $143 −30.6
Log Population $498 $571 14.7
120 of 11983 county year observations excluded, the top 1% by Cook's distance.

Log respecification

The level model indicated that a constant dollar association was insufficient to accurately describe the relationship. Logging purchasing power was tested to see whether the association was better represented proportionally, enabling a comparable shift in a task group to correspond to a larger dollar difference in counties with higher purchasing power. The transformation reduced the right skew in purchasing power and laid the groundwork for reevaluating the residual pattern (Figure 11).

Figure 11: Distribution of Purchasing Power Before and After Log Transformation. County year purchasing power across the analytical panel, shown as a kernel density. A) The original dollar scale is right skewed with a long upper tail, and the mean sits well above the median. B) On the natural log scale the tail pulls in and the two measures nearly coincide. Solid lines mark the median, dashed lines the mean; each panel reports its own skewness.

Refitting the same specification using log purchasing power resolved most of the problems identified in the level model. Systematic curvature in the residuals lessened, while the QQ points followed the reference line more closely after the transformation (Figure 12). Some tail deviation remained, but residual skew declined from 1.01 to 0.09. The log linear specification therefore provided the baseline for evaluating the more flexible models that followed.

Figure 12: Regression Diagnostics, Level vs. Log Specification. Columns are the diagnostic, rows are the specification. A) and B) show the level specification, in dollars. C) and D) show the log specification, in log dollars. Residuals are shown as binned density with a loess smoother, since individual points would overplot. The level specification shows systematic curvature and a heavy QQ tail, while the log specification is substantially flatter and tracks the reference line more closely.

The same comparison also holds in the units of the outcome (Figure 13). On the level scale, the model closely tracks purchasing power throughout the middle of the distribution, where most counties are located. However, both tail bins deviate substantially from their predictions, coming in approximately $8,000 to $9,000 above the expected values. On the log scale, the points closely follow the line throughout the distribution.

Figure 13: Binned Fit Accuracy, Level vs. Log Specification. Fitted values are cut into twenty equal count bins, and each point compares a bin’s mean prediction with its mean actual value; points on the black one to one line indicate no systematic bias in that range. A) On the level scale both tails sit above the line. B) On the log scale the points track it throughout.

The residual pattern did not disappear entirely after the transformation. The remaining structure could reflect nonlinearity, interactions among the variables, or omitted factors that the diagnostics alone could not distinguish. This motivated the use of a random forest, which could capture nonlinear relationships and interactions without specifying them in advance. The log linear specification served as the benchmark, allowing us to test whether a more flexible model captured additional signal in held out counties.

Predictive modeling

The predictive comparison used the county year panel assembled earlier, incorporating state identifiers for the subsequent models. Model complexity only increased when additional flexibility enhanced performance on the held-out counties. The log respecification addressed the issues identified in the level model, while the remaining residual structure prompted testing a random forest capable of capturing nonlinear relationships and interactions. A three hidden layer feedforward neural network served as an additional test to determine if greater model flexibility further improved performance.

All models were evaluated on the same held-out counties from the grouped split, so any differences in performance reflected model structure rather than variations in the evaluation observations. If the more flexible models failed to improve held-out performance, the log linear specification was retained as the simpler representation of the relationship.

Evaluation design

Feature set and reference groups

Consistent with the reference-group approach established in Section 4.2, non-routine cognitive tasks were excluded from the predictive specification as well, maintaining consistency in interpreting the task group coefficients across both inferential and predictive analyses.

State indicators also required a comparable reference category. Mean purchasing power varied widely across states, and a state situated near the center of that distribution served as the reference point (Figure 14). By selecting a state near the middle, comparisons were avoided to states with exceptionally high or low purchasing power, simplifying the interpretation of the resulting coefficients. While the choice of reference state did not affect the overall model fit, it influenced the expression of the state coefficients.

Grouped train and test split

This is not a forecasting exercise, and a time ordered split is not required for its usual reason. A grouped split is still necessary, because a county’s economic profile barely moves year to year, which makes its 2018 and 2019 rows near duplicates. If two adjacent rows land on opposite sides of the split, the model can recall a county rather than learn a relationship. We group by county and split on row membership before any further feature engineering. The split assigns 9,565 county year observations to training and holds out the remaining 2,418, with no county contributing rows to both. All feature engineering was conducted after the split to maintain this separation. Consequently, every test observation originated from a county that was not represented in the training data.

Year and state indicators

A grouped train and test split was used because observations from the same county were closely related across years. A random split could inadvertently place observations from the same county in both sets, potentially allowing the model to use information about a county it had already encountered. By grouping observations by county, we effectively eliminated this overlap and provided a more robust test of whether the observed relationship could generalize to new locations.

The reference state was Texas, whose mean purchasing power was close to the median of the state distribution (Figure 14). By selecting a state near the center, we avoided anchoring comparisons to an unusually high or low purchasing power state, making the coefficients more interpretable. While most state coefficients remained within a few percent of the reference, the largest differences reached 17%. Once the economic controls and year indicators were included, as demonstrated by the permutation results below, state indicators contributed little to held-out performance.

Figure 14: Mean Purchasing Power by State. Mean county year purchasing power for each state, 2008 to 2023 excluding 2020, in dollars at national average prices. Each segment runs from the median of the state means to that state’s own mean, making its length a measure of distance from the median rather than absolute level, and the axis starts at $30,000. Shading separates states below the median from those above it, and the highlighted point marks the reference state the models hold out. States with fewer than five counties in the panel are pooled into a single bucket, which the ranking does not display but which does enter the median.

Collinearity of the indicator blocks

The reference group logic applies here as well (Figure 8). The additional concern was whether the year and state indicators introduced new collinearity after being added to the predictive specification. They did not, as all reported VIFs remained below the conventional threshold of 5 (Figure 15). This confirmed that the indicator blocks could be retained without materially destabilizing the estimated coefficients.

Figure 15: Variance Inflation Factors for the Indicator Blocks. The ten highest values across the year and state indicator blocks, computed on the full feature matrix after dropping the reference task group. All ten are year indicators, labeled by year. The dotted vertical line marks the conventional threshold of 5, which no indicator approaches.

Linear benchmark

This specification is a predictive benchmark rather than a second inferential model. It reuses the diagnostic insights established by the panel regression above, now applied to the training split alone. This shared split enables direct comparability with the subsequent random forest and neural network models. We present two versions of each linear model, estimated using ordinary least squares (OLS). The baseline model uses only the task groups with poverty and unemployment, while the full model incorporates year and state controls. Reporting both versions highlights the changes introduced by time and geography. Among the two, the full model is the one that is carried forward.

Baseline specification

The baseline specification yielded an R² of 0.766. Relative to non-routine cognitive work, the coefficients for the other three task groups were negative, indicating lower purchasing power as counties shifted toward those groups. The largest negative association was observed for non-routine manual work, where a one percentage point shift was associated with approximately $1,305 lower purchasing power. Table 3 reports these estimates alongside those from the full specification.

Full specification

Adding year and state indicators increased the R² from 0.766 to 0.871, indicating that time and geography contributed to additional variation in purchasing power. The unemployment rate coefficient shifted from approximately −$1,207 in the baseline specification to +$450 in the full specification for each one percentage point increase in unemployment. This change in direction suggested that the baseline estimate accounted for geographic and temporal differences that were not previously considered. In contrast, the task group coefficients maintained their direction and relative ordering.

Residual curvature persisted even after the full level specification (Figure 12). This figure was important because the increase in the coefficient alone suggested that incorporating year and state indicators effectively improved the model. However, the diagnostics revealed that while these variables enhanced the overall fit, they failed to eliminate the systematic pattern in the residuals. This distinction indicated that the underlying issue was the functional form of purchasing power on its original dollar scale, rather than omitted time or geographic differences.

Full specification on the logged outcome

The level specification did not fit purchasing power evenly across its range (Figure 13). The largest departures occurred at the lower and upper ends of the distribution, where observed purchasing power was roughly $8,000 to $9,000 above the fitted values. This systematic pattern suggested that the remaining problem was related to the scale of purchasing power rather than omitted controls, providing a clear reason to test the outcome in logarithmic form.

Logging purchasing power reduced the residual pattern observed in Figure 12 and enhanced the performance of the linear specification. After transforming the predictions back to dollars, the log linear model achieved an in-sample accuracy of 0.913 and a held-out R² of 0.896. This close agreement suggested that the model effectively generalized to counties not included in the training data.

Coefficient comparison

Table 4 compared the three specifications. The log specification was chosen for inference because the diagnostics for the level specification indicated that its functional form was unsuitable, making its coefficient estimates less reliable for interpretation.

Table 3: Ordinary Least Squares Coefficients Across Model Specifications and Outcome Scales. Coefficients from the baseline, full level, and full log OLS fits. Purchasing power is county median household income divided by the local price level, stated in dollars at national average prices.
Purchasing Power Log Transformed Purchasing Power
Baseline Full Coefficient % per 1 pp
Task Groups
Routine Cognitive −$81,085 −$87,421 −1.080 −1.07
Routine Manual −$71,490 −$81,143 −1.089 −1.08
Non-Routine Manual −$130,549 −$91,904 −1.213 −1.21
Economic Controls
Poverty Rate −$1,376 −$1,729 −0.029 −2.87
Unemployment Rate −$1,207 $450 0.001 0.08
Non-routine cognitive is the reference task group. Task group coefficients in the three model columns are per unit of group proportion; the final column converts the log coefficients to the percent change in purchasing power associated with a one percentage point shift.

Table 3 revealed that the task group coefficients maintained the same direction across all three specifications, although their magnitudes changed after incorporating controls. The non-routine manual coefficient shifted from approximately -$130,549 in the baseline specification to -$91,904 in the full level specification. Routine cognitive changed from about -$81,085 to -$87,421, while routine manual changed from -$71,490 to -$81,143. These results demonstrated that adjustments altered the magnitude of the associations without altering their overall order.

The poverty rate coefficient increased in magnitude from approximately -$1,376 to -$1,729 for each additional percentage point of poverty. In contrast, the unemployment rate coefficient changed from approximately -$1,207 in the baseline specification to +$450 in the full level specification. This difference showed that poverty remained strongly associated with purchasing power after adjustment, while the unemployment association was comparatively weak and sensitive to specification.

Table 4: Model Specification Comparison. Fit summary across the baseline, full level, and full log OLS specifications.
Baseline Full Model Log Model
Outcome Purchasing Power Purchasing Power log(Purchasing Power)
Year and State Controls No Yes Yes
0.766 0.871 0.913
Residual Diagnostics Failed Passed
Primary Specification No No Yes
R² is in sample, fit on the 9,565 training rows only.

In the log specification, each additional percentage point of poverty was associated with approximately 2.9% lower purchasing power, assuming task groups, time, and geography remained constant. This proportional association was larger than that of any individual task group. Poverty was therefore the strongest single association among the variables reported in Table 3.

Random forest

For the random forest model, hyperparameters were selected using a grid search with grouped five-fold cross-validation on the training counties. The search explored various configurations, including the number of trees ranging from 600 to 1,000, maximum depth from 10 to 20 or unlimited, minimum leaf size from 1 to 5, and the proportion of features considered at each split. Grouping by county ensured that the separation between places during tuning was preserved, reducing the risk of selecting parameters that performed well only due to observations from the same county appearing across folds. The final model used 1,000 trees, unlimited depth, a minimum leaf size of 1, and half of the available features at each split.

The random forest model achieved an R² value of 0.876 on the held-out counties, which was slightly lower compared to the log linear model’s R² of 0.896 on the same held-out counties. This 0.020 difference suggested that the additional flexibility of the random forest model did not substantially enhance its performance. Additionally, the random forest model achieved an R² value of 0.989 on the training counties, resulting in a 0.113 gap between training and held-out performance. This gap indicated that the random forest model captured additional structure in the training data that did not generalize to counties it had not encountered.

Figure 16: Random Forest Residuals vs. Predicted Values. Residuals from the raw target random forest on the held out test counties, both axes in dollars. No systematic curve or trend is visible across the range of predictions.

No systematic curvature or trend appeared in the random forest residuals across the range of predicted values (Figure 16). This suggested that the random forest effectively captured the nonlinear structure that posed challenges for the level linear model without transforming purchasing power. However, the lower held-out R² value of 0.876 and the large gap between training and test performance indicated that this ability to capture the structure did not translate into improved generalization to unseen counties.

Ablation and permutation importance

The random forest is largely insensitive to whether purchasing power is modeled on its original or logarithmic scale. Fitting the model to the logged outcome and then back-transforming the predictions yields an R² of 0.877, which is slightly higher than the R² of 0.876 obtained when the outcome is modeled directly. Since the log specification is also used for the primary linear model, the log target random forest is used for the subsequent ablations and model comparisons. The interpretability measures, including permutation and partial dependence, are computed on the raw target random forest, which reports directly in dollars.

Ablation analysis revealed the model’s effectiveness in compensating for the removal of an entire group of variables and subsequent refitting from scratch. The removal of task groups resulted in a decline in the held out R² value from 0.876 to 0.834, a decrease of 0.042. Since the random forest was retrained after removing the task groups, the remaining variables had the opportunity to recover any overlapping information. Consequently, the remaining decline indicated that poverty, unemployment, time, and geography could not fully replace the information provided by the task groups.

In comparison, eliminating poverty and unemployment reduced the held out R² from 0.877 to 0.698, resulting in a decline of 0.179. This drop was substantially larger than the 0.042 decline observed when the task groups were removed, suggesting that the economic controls provided more of the model’s explanatory power. Even after incorporating the task groups into the economic controls, performance improved, indicating that they contributed information not fully captured by poverty and unemployment alone. All three models used the same random forest specification, logged outcome, and held-out counties, ensuring direct comparability of their performance differences.

Grouped permutation importance

Joint permutation provided a complementary measure of how strongly the fitted random forest relied on the task groups. Permuting the three modeled task group proportions together reduced the held out R² by 0.153 while leaving the fitted model unchanged. This decline was much larger than the 0.042 reduction from removing the task groups and refitting the random forest.

The difference between the two measures revealed what the model could do after removing task group information. During ablation, the random forest was retrained and could recover some overlapping information from poverty, unemployment, and year. Permutation, on the other hand, disrupted the task group information after the random forest had already learned to use it. The larger permutation decline indicated that the fitted random forest heavily relied on the task groups, even though some of their information overlapped with the other variables.

Year indicators showed a decline of 0.128 when permuted, compared to only 0.008 for state indicators. This difference demonstrated that time contributed considerably to held-out performance compared to state after accounting for economic controls and task groups.

Among the three modeled task groups, non-routine manual produced the largest individual permutation decline. These values were not interpreted as additive or independent contributions because the task groups were mechanically related through their employment proportions. Nevertheless, the ordering was consistent with the regression results, where non-routine manual also had the largest negative task group association with purchasing power.

Neural network

The random forest model did not outperform the log linear specification. Whether that reflected a limitation of the random forest or the amount of structure present in the data remained an open question. To further investigate this, we evaluated a three-hidden-layer feedforward neural network as an additional test to determine if greater model flexibility would lead to improved held-out performance. The network consisted of three hidden layers and a single linear output unit for the continuous purchasing power outcome.

Training loss steadily declined, while validation loss began to rise. This divergence indicated that the network was fitting the training data more closely without improving performance on unseen observations, suggesting overfitting. To address this, we added L2 weight regularization, increased dropout, and shortened the early stopping patience. These changes narrowed the gap between training and validation loss, providing a stronger test of whether the neural network could generalize beyond the training data.

Figure 17: Training and Validation Loss, Initial vs. Regularized Network. Loss is mean squared error on the standardized target; both panels share the same y axis. A) The initial network, where validation loss declines briefly then rises while training loss keeps falling, the signature of overfitting; the dashed line and marked point locate the epoch of minimum validation loss. B) The same network after adding L2 regularization, heavier dropout, and a shorter early stopping patience, where the two curves stay much closer together.

Regularization reduced the disparity between training and validation loss (Figure 17), yet the neural network still failed to surpass the random forest or the log linear model. On the same counties, the random forest achieved a held-out R² of 0.877, while the log linear model achieved 0.896. Consequently, the neural network’s additional flexibility did not offer any improvement over the simpler alternatives. Since the study aimed to elucidate the relationship between task groups and purchasing power, further increases in model complexity were deemed unnecessary. Therefore, the log linear and random forest models were retained for the subsequent comparisons.

Cross validation

Five-fold grouped cross-validation was used to assess the consistency of the observed performance in the single train and test split across various county sets. Counties remained grouped within each fold, ensuring that observations from the same county were not included in both training and validation data.

Table 5: Five Fold Grouped Cross Validation Performance. Mean and standard deviation of held out R² for OLS and the random forest, both fit on the log target and scored on the dollar scale after back transformation.
Model Mean R² SD
Ordinary Least Squares 0.9036 0.0080
Random Forest 0.8982 0.0060
Both models are fit on the log target and scored on the dollar scale after back transformation.

The Table 5 revealed that the log linear model and random forest achieved comparable mean values across the five folds. The difference in their average performance was smaller than the variation observed across folds, suggesting that the earlier comparison was not influenced by a specific train and test split. Consequently, the neural network was excluded as it failed to enhance held-out performance and was not retained.

Summary

The panel regression provided the primary inferential results, with standard errors clustered by county and controlling for poverty, unemployment, population, and year. Diagnostic checks supported a proportional relationship, justifying the log specification. However, the variance decomposition revealed that most task group variation occurred between counties rather than within counties over time. Predictive models were evaluated on held-out counties, where the log linear model achieved an R² of 0.896 compared to 0.877 for the random forest. The neural network model provided no further improvement. Removing the task groups reduced the held-out R² by 0.042, indicating that poverty, unemployment, time, and geography did not fully capture the information they provided. These results collectively support retaining the log linear specification as the primary representation of the relationship reported in Section 5.

Results

The results unveiled four main findings. Task groups maintained their association with purchasing power even after accounting for factors such as poverty, unemployment, population, and year. A one percentage point shift from non-routine cognitive to non-routine manual work was linked to a reduction in purchasing power of approximately $846, with a 95% confidence interval of ±$38 at the panel mean. The log specification provided the closest representation of this relationship, achieving a held out R² of 0.896, compared to 0.877 for the random forest, while the neural network offered no further improvement. Poverty rate demonstrated the strongest association with purchasing power, but task groups provided additional insights beyond conventional economic indicators. Moreover, differences in task groups persisted across counties and over time. The subsequent sections examine the magnitude, geographical distribution, and stability of these relationships.

The association

The log specification of the panel regression was used for inference based on the diagnostic results outlined in Section 4. A one percentage point shift from non-routine cognitive to non-routine manual work was associated with approximately $846 lower purchasing power, with a 95% confidence interval of ±$38 at the panel mean (Figure 18). Routine cognitive and routine manual work also exhibited negative associations with purchasing power, although their estimated differences were smaller. This ordering aligns with the level specification reported in Section 4. For comparison, each additional percentage point of poverty was linked to approximately 2.9% lower purchasing power, a larger proportional association than any individual task group.

Figure 18: Task Group Coefficients, Log Specification. From the panel regression, expressed as dollars of purchasing power at the panel mean per one percentage point shift into each group, with 95% intervals. Non-routine cognitive is the reference group at zero. Standard errors are clustered by county.

All three estimates of the purchasing power difference associated with a one percentage point shift from non-routine cognitive work toward another task group remained below zero, with non-routine manual work showing the largest negative association (Figure 18). The narrow intervals around the estimates indicated that the direction and ordering of these task group relationships were precisely estimated after accounting for the other variables in the model.

The level specification provided an additional perspective on the task group associations and economic controls in dollar terms. Among the controls, the poverty rate carried the largest and most precisely estimated coefficient, while the unemployment rate and log population were smaller but remained distinguishable from zero at the 0.05 level (Table 6). The table also reported the model’s overall F test, sample size, and R², offering a summary of the fit and precision of the complete level specification.

Table 6: Panel Regression Coefficients, Level Specification. Dollar coefficients for the task groups and controls, with standard errors, t-statistics, and p-values. Non-routine cognitive is the reference task group.
Variable Coefficient Std. Error t p-value
Routine Cognitive −$85,791 7,689 −11.16 0.000
Routine Manual −$70,097 4,487 −15.62 0.000
Non-Routine Manual −$104,139 5,526 −18.84 0.000
Poverty Rate −$1,605 56 −28.83 0.000
Unemployment Rate $206 97 2.13 0.033
Log Population $498 246 2.02 0.043
F(20, 11962) = 757.5, p < 0.001; n = 11,983; R² = 0.848. Standard errors clustered by county.

Not poverty in disguise

A natural concern is that the task group associations simply reflect poverty, as counties with higher poverty rates also tend to have more routine work. Mean poverty and unemployment rates across the study period were both highest across much of the rural South, the Mississippi Delta, and parts of the Southwest border region (Figure 19), which overlapped with areas of low purchasing power in Figure 24. This overlap was important because each additional percentage point of poverty was associated with approximately 2.9% lower purchasing power in the log specification. Unemployment also showed a concentration along the California coast and Central Valley that was less apparent for poverty, indicating that the two measures captured different dimensions of local economic conditions.

Figure 19: Two Conventional Distress Measures, by County. A) Mean poverty rate and B) mean unemployment rate, 2008 to 2023, in percent. Each panel uses its own color scale, clipped at the 2nd and 98th percentiles so extreme counties do not compress it. Counties without a value are shown in gray. Continental United States only.

A test of whether economic conditions accounted for the associations between task groups compared two log panel specifications estimated on the same observations (Figure 20). The first specification included task groups and year indicators, while the second added poverty, unemployment, and log population. After adding the controls, the estimates for the non-routine manual and routine manual became smaller, indicating that some of their initial associations overlapped with these economic conditions. However, routine cognitive remained relatively stable, staying close to −1.1% in both specifications. Notably, all three task group estimates remained negative and accurately estimated after adjustment. This comparison demonstrated that poverty, unemployment, and population explained part of the relationship, but they did not fully account for the association between task groups and purchasing power.

Figure 20: Task Group Coefficients Before and After Controls. From the log panel regression, before and after adding poverty, unemployment, and population controls, as percent change in purchasing power per one percentage point shift, with 95% intervals. The two specifications are offset vertically within each task group so their intervals do not overlap. Non-routine manual and routine manual shrink with controls; routine cognitive does not.

What carries the signal

Two complementary tests were conducted to assess the contribution of information from the task groups. The ablation test removed an entire feature block and refit the random forest (Figure 21). Removing poverty and unemployment resulted in a 0.178 reduction in held out R², compared to a smaller decline of 0.042 when the task groups were removed. This indicates that the economic controls provided more explanatory information overall, but the task groups still improved performance beyond what poverty and unemployment alone could capture.

Figure 21: Held Out R² Loss From Feature Block Removal. From the log target random forest, relative to the full model’s 0.876. The economic controls are poverty rate and unemployment rate. Removing them costs more than four times as much as removing the task groups.

A second measure left the fitted model unchanged and perturbed each feature or feature block through permutation (Figure 22). Jointly permuting the task groups led to a 0.153 reduction in held out R², compared to 0.128 for the year indicators and 0.008 for the state indicators. Poverty exhibited the largest individual decline. These results suggest that the fitted random forest heavily relied on the task groups even after incorporating economic controls, time, and geography.

The ablation and permutation results provided further insights into the model’s performance. Removing the task groups and refitting the model resulted in a modest reduction of 0.042 in the held-out R² value, suggesting that the model could still recover some overlapping information from poverty, unemployment, and year. In contrast, permuting the task groups after fitting caused a decline of 0.153, indicating that the model could no longer compensate for the loss of information. These two tests together revealed that the task groups contained information that overlapped with other variables but was not fully replaced by them.

Figure 22: Feature Importance by Grouped Permutation. On the held out test counties, measured as the drop in R² when each feature or block is shuffled. The task groups are permuted jointly, as are the year and state indicator blocks. Larger drops indicate features the model relies on more.

The random forest traced the relationship across the observed range of each non-routine cognitive reference comparison (Figure 23). Purchasing power gradually declined across the three modeled task groups, without any distinct thresholds or reversals. Despite capturing these flexible relationships, the random forest achieved a held out R² of 0.877, compared to 0.896 for the log linear model, a difference of 0.019. This pattern aligned with the proportional relationship identified by the log linear specification and explained why the additional flexibility of the random forest did not enhance held out performance. Since the four task groups collectively accounted for the entire range, these curves represented the model’s response to changes in one group while keeping the remaining values constant. Therefore, these curves should not be interpreted as independent changes in employment.

Figure 23: Partial Dependence on Each Task Group. Random forest partial dependence of predicted purchasing power, in dollars, on each non-reference task group, computed on the held out test counties. A) routine cognitive, B) routine manual, C) non-routine manual. Each curve moves one group across its observed range with all other features held at their values, and a panel’s line covers only the segment of the shared x axis that group occupies. Panels share both axes, which makes the declines comparable across groups. The smooth, near monotonic declines match a proportional association.

Where purchasing power is strained

The association also exhibited a distinct geographic pattern. Purchasing power was highest along the metropolitan Northeast corridor, across parts of the upper Midwest, and in certain regions of the mountain West (Figure 24). Conversely, the lowest values were concentrated across the rural South and the southern border region. The map included all counties with a purchasing power value, rather than only those in the analytical panel, because the outcome required only household income and RPP.

This geographic pattern was also reflected in county task groups. Counties were categorized based on the task group that accounted for the largest average proportion of employment over the study period. In 664 counties, non-routine cognitive work was the dominant type of work, with a median purchasing power of $61,162. In contrast, 97 non-routine manual counties had a median purchasing power of $46,657, while 87 routine manual counties had a median purchasing power of $50,368. Routine cognitive work was the dominant group in 0 counties. These differences demonstrated that the association observed across the continuous task group measures was also evident when counties were grouped by the type of work that constituted the largest portion of local employment.

Figure 24: Where Purchasing Power Is Strained. Mean purchasing power by county, 2008 to 2023, in dollars at national average prices, clipped at the 2nd and 98th percentiles so extreme counties do not compress the scale. Counties without a purchasing power value are gray. Continental United States only.

The relationship is proportional

Across the five grouped cross-validation folds, the log linear model and random forest showed similar variation across folds (Figure 25). The single held-out split also demonstrated the same pattern. The log linear model achieved an R² of 0.896, while the random forest achieved an R² of 0.877. The comparison extended to the neural network, which achieved a held-out R² of 0.869 and a mean absolute error (MAE) of $4,145 (Table 7). The neural network did not provide any improvement over the simpler models. Overall, these results indicated that additional model flexibility did not capture enough new structure to improve performance beyond the log linear specification.

Figure 25: Held Out R² by Fold, Linear Model vs. Random Forest. Held out R² across the five grouped cross validation folds, for the linear model and the random forest, both fit on the log target and scored on the dollar scale after back transformation, with the mean marked. The gap between the model means is small relative to the spread across folds.
Table 7: Model Performance on the Held Out Counties. Held out R² and MAE, in dollars, on the same counties from the single grouped split.
Model Mean Absolute Error
Ordinary Least Squares 0.896 $3,636
Random Forest 0.877 $3,916
Neural Network 0.865 $4,169
OLS and the random forest are fit on the log target and scored on the dollar scale after back transformation.
Neural network results are from this single split only; the five fold grouped cross validation table in the analysis section reports cross validated results for the other two models.

A complementary perspective used purchasing power dollars (Figure 26). Across ten equal-sized bins of routine cognitive tasks, observed purchasing power increased in the lower range, peaked near 20%, and then declined. The random forest closely reproduced this pattern, rather than identifying a substantially different relationship. This observation did not contradict the smoother declines in Figure 23, as the partial dependence curves varied one task group while holding the others constant, whereas the binned comparison reflected how the task groups occurred together in actual counties. Since the four groups summed to one, changes in routine cognitive work in the observed data were accompanied by changes in the other groups. Therefore, the agreement between the flexible model and the simpler log linear specification supported retaining the proportional specification rather than adding further model complexity.

Figure 26: Observed and Predicted Purchasing Power by Routine Cognitive Value. Mean observed and predicted purchasing power across ten equal count bins of routine cognitive work on the held out test counties, with predictions from the log target random forest back transformed into dollars. Each point marks a bin’s mean routine cognitive value. The relationship rises then declines rather than moving monotonically, and the random forest reproduces that shape. The panel reports the forest’s held out R² and mean absolute error, in dollars.

The differences are durable

The differences observed across counties were not confined to a specific time period. The average proportion of employment in each task group remained relatively stable over the fifteen-year panel (Figure 27). The largest discontinuity occurred between 2009 and 2010, coinciding with the Census occupation coding change mentioned in Section 3. Outside this period, the national distribution of the four task groups remained comparatively stable.

Figure 27: Average County Task Group Value by Year. Taken as an unweighted mean across panel counties. Each band’s thickness is that group’s mean percent of the four group total, so the four stack to 100% and thickness rather than height carries the value. The four task groups stay nearly constant over the study period. The bold dashed line marks the Census occupation coding change between 2009 and 2010. The shaded band at 2020 marks the year ACS 1 year estimates were not published; the narrow break in the bands at that point reflects the missing year rather than a data error.

County rankings provided a more robust test of whether the same places maintained distinct characteristics over time. Table 8 compared each county’s task group ranking in 2010 with its ranking in 2023 using Spearman correlations. Routine manual and non-routine cognitive work exhibited the greatest stability, indicating that counties with relatively high values in 2010 generally maintained high values in 2023. Non-routine manual work showed more movement, while routine cognitive work experienced the largest changes. This ordering aligns with the variance decomposition in Section 4, which demonstrated that routine cognitive work exhibited more within-county variation compared to the other task groups.

Table 8: County Ordering Stability, 2010 to 2023. Spearman rank correlation of county orderings between the two years, by task group, across counties present in both.
Task Group Rank
Correlation
Routine Cognitive 0.41
Routine Manual 0.87
Non-Routine Cognitive 0.88
Non-Routine Manual 0.65
Computed from 2010 onward to avoid the Census occupation coding change.

The stability is visible county by county, and in the magnitude of change as well as the ordering, over the same two years used in Table 8 (Figure 28). Routine manual and non-routine manual, the two groups with almost no net movement in Figure 27, are centered near zero and run in both directions across counties. That mixed direction is consistent with variation sitting between counties rather than within them. Routine cognitive and non-routine cognitive, the pair whose national levels moved most, show a change that is both larger and far more uniform in direction, which matches the within county movement identified in the variance decomposition of Section 4. Routine cognitive rose in np.int64(2)% of counties and non-routine cognitive in np.int64(96)%, against np.int64(56)% for routine manual and np.int64(34)% for non-routine manual.

Figure 28: County Task Group Change, 2010 to 2023. Percentage point change in each task group’s value, the same two years used in Table 8, one dot per county on a shared scale so magnitude and direction are directly comparable across groups. Dots are spread vertically only to keep counties at the same change from overlapping, so the width of each swarm shows how many counties sit at that change. The black tick marks each group’s median, labelled above it. The share of counties whose value rose is reported in the text rather than on the figure: routine cognitive fell in almost every county while non-routine cognitive rose in almost every one, and the two manual groups split closer to evenly. Routine manual and non-routine manual are centered near zero with mixed direction; routine cognitive and non-routine cognitive show larger, more uniformly signed change in opposite directions.

These movements followed the same cognitive to manual and routine to non-routine dimensions introduced in Figure 1 (Figure 29). Of the 3,209 counties observed in both periods, 87 percent shifted towards non-routine cognitive work, while the remainder moved away from it towards the three task groups associated with lower purchasing power. The wider coverage was worth the change of source because the counties the annual estimates omit are the small and rural ones. They moved in the same direction as the rest, though far less uniformly: 83 percent of them shifted towards non-routine cognitive work, against 99 percent of the counties the panel already covered. Because counties in the eastern half of the map are small enough that one arrow each would overlap into a smudge, neighboring counties were pooled onto a 65 kilometer grid and drawn as 1,457 averaged arrows, colored by how their counties divided rather than by the average alone. Counties moving away from non-routine cognitive work were scattered rather than concentrated, appearing somewhere in 22 percent of the grid cells. The arrow figure draws a direct connection between the observed county changes and the conceptual framework introduced earlier in the study.

Figure 29: Direction of County Task Group Shift, 2010 to 2014 Versus 2019 to 2023. Counties are pooled onto a 65 kilometer grid, one arrow per cell, so a dense eastern cluster reads as a single direction rather than a knot of overlapping arrows. Each arrow points along its cell’s average movement across the manual to cognitive and routine to non-routine axes, with length and thickness following the size of that average, floored so small shifts still read as arrows and capped at the 95th percentile so a few cells do not dominate. Color counts counties rather than averaging them: purple marks a cell holding counties moving in both directions, keeping a county that moved away from non-routine cognitive work visible even where its neighbors moved toward it. Gray counties lack ACS five year estimates for both periods.

Finally, the estimated associations themselves remained consistent over time. Separate specifications for 2010 to 2019 and 2021 to 2023 ran alongside the full panel (Figure 30). The ordering of the three task group coefficients remained unchanged across all three periods, and their magnitudes shifted only slightly. Reporting the coefficients as percentage changes ensured comparability between the periods, even though nominal purchasing power increased later in the panel. The persistence of both the county patterns and the coefficient ordering suggested that the association was not confined to a single recession, recovery, or pandemic period.

Figure 30: Coefficient Stability Across Time Windows. Log panel regression coefficients, as percent change in purchasing power per one percentage point shift out of non-routine cognitive work, refit on the pre pandemic and post pandemic windows alongside the full panel (diamond). The legend carries each window’s county year count, and the horizontal bar through each point is that window’s 95% confidence interval. Only the full panel estimate is labeled with its value, at the left of each row. Standard errors clustered by county in every fit.

The results revealed that county task groups continued to be linked to purchasing power even after accounting for factors such as poverty, unemployment, population, and the year. The largest difference was observed between non-routine manual work and other tasks. A one percentage point shift from non-routine cognitive work to non-routine manual work was associated with a reduction in purchasing power of approximately $846 at the panel mean, with a 95% confidence interval of ±$38. Routine cognitive and routine manual work were also negatively associated with purchasing power compared to non-routine cognitive work. Among the variables examined, poverty showed the strongest association with purchasing power, with each additional percentage point associated with a reduction of approximately 2.9%. The random forest analysis indicated that task groups provided additional information beyond the economic controls. Removing the task groups and refitting the model reduced the held out R² by 0.042, while removing poverty and unemployment reduced it by 0.178. Jointly permuting the task groups further reduced the R² value by 0.153, compared to 0.128 for year indicators and 0.008 for state indicators. The geographic analysis showed that counties dominated by non-routine cognitive work had higher purchasing power, while counties dominated by the two manual groups had lower purchasing power. The stability analysis demonstrated that these county differences persisted over time. Routine manual and non-routine cognitive work maintained the most consistent county ordering between 2010 and 2023, while routine cognitive work exhibited the greatest movement. The ordering of the task group coefficients remained unchanged in the full panel, 2010 to 2019, and 2021 to 2023.

The level model, log linear model, random forest, and neural network were compared to find the most suitable representation of the relationship. The level specification produced negative task group associations but showed systematic curvature in the residuals, suggesting that a constant dollar difference across the purchasing power distribution was inadequate. Logging purchasing power reduced residual skew from 1.01 to 0.09, resulting in an in-sample R² of 0.913 and a held-out R² of 0.896. The random forest captured the nonlinear structure visible in the level model but did not improve held-out performance. Its log target specification achieved a held-out R² of 0.877, 0.019 below the log linear model, while the raw target random forest achieved a training R² of 0.989 but only 0.876 on held-out counties. The neural network added further flexibility, but regularization did not improve on either of the other models. Five-fold grouped cross-validation yielded similar general comparisons, with the difference between the log linear and random forest models remaining smaller than the variation across folds. Across these four specifications, the log linear model preserved the task group associations, resolved most of the diagnostic problems in the level model, and scored highest on held-out counties. The implications and limitations of these results are discussed in Section 6.

Conclusions

Summary of findings

Across all counties analyzed over a span of fifteen years, county task groups remained linked to purchasing power after accounting for poverty, unemployment, population, and year. The largest difference was observed for non-routine manual work. A one percentage point shift from non-routine cognitive to non-routine manual work was associated with approximately $846 lower purchasing power on average, with a 95% confidence interval of ±$38. Routine cognitive and routine manual work were also negatively associated with purchasing power relative to non-routine cognitive work. Among the variables examined, poverty demonstrated the strongest association with purchasing power, with each additional percentage point associated with approximately 2.9% lower purchasing power. The task groups still carried information beyond these conventional economic measures. Removing them and refitting the random forest reduced the held out R² by 0.042, while jointly permuting them reduced it by 0.153. In comparison, permuting the year indicators reduced R² by 0.128, and permuting the state indicators reduced it by only 0.008.

The sequence of models showed which representation of the relationship the data supported. The level specification identified negative associations but revealed systematic curvature in the residuals, suggesting that a constant dollar relationship was inadequate for the data. Logging purchasing power reduced the residual skew from 1.01 to 0.09, resulting in an in-sample R² of 0.913 and a held-out R² of 0.896. The random forest analysis explored whether nonlinear relationships and interactions improved upon this specification, but its log target version achieved a held-out R² of 0.877. The raw target random forest also demonstrated a closer fit to the training data compared to unseen counties, with R² values of 0.989 and 0.876, respectively. The neural network introduced even greater flexibility but failed to yield further improvements after regularization. Five-fold grouped cross-validation yielded similar general comparisons, with differences between the log linear model and random forest being smaller than the variation across folds. Geographic and stability analyses extended these results beyond model performance. County task group differences persisted from 2010 to 2023, and the ordering of the task group coefficients remained consistent across the entire panel, 2010 to 2019, and 2021 to 2023.

The study also clarified the scope of its findings regarding automation and affordability. The analysis did not measure automation adoption, job displacement, or the likelihood of specific counties losing jobs to new technologies. Instead, it examined the types of tasks already present in county employment and assessed their association with purchasing power. Counties with higher concentrations of routine and manual work, which are often emphasized in the automation literature, tended to have lower purchasing power compared to counties with more non-routine cognitive work. This relationship persisted after adjustment and remained consistent across geography and time. However, it did not establish a causal link between automation and purchasing power differences. Therefore, the study indirectly addressed automation by identifying regions where task structures commonly associated with higher automation potential coincided with lower purchasing power. The consequences of actual technological displacement remain for future research.

Contributions

The first contribution shifted the outcome from nominal income to purchasing power. Median household income was adjusted using BEA RPPs, directly incorporating local price differences into the measure of household resources. This adjustment distinguished counties with similar nominal incomes where the same income supported varying levels of purchasing power. Instead of comparing income alone, the study examined what that income could purchase after accounting for local price variations.

The second contribution measured task groups at the county level and tracked them over time. The final panel comprised 840 counties from 2008 to 2023, excluding 2020, and represented approximately 84% of the U.S. population. Using counties provided a finer geographic scale than broader labor market areas and directly connected task groups with local purchasing power, poverty, and unemployment. The panel structure demonstrated that these differences were not confined to a single point in time. County task group patterns remained relatively consistent throughout the study period, enabling the analysis to distinguish enduring differences between counties from short-term changes within them.

The county-level analysis also showed that task groups contained information beyond conventional measures of local economic conditions. Removing the task groups and refitting the random forest reduced the held-out R² by 0.042, while jointly permuting them reduced it by 0.153. Poverty exhibited a stronger association with purchasing power than any individual task group, but poverty and unemployment did not fully capture the information contained within the task groups. Counties that appeared similar on these conventional measures could still differ in the types of work residents performed and the purchasing power associated with those differences.

The third contribution determined the most appropriate representation of the relationship between task groups and purchasing power. The level specification revealed systematic residual curvature, while logging purchasing power reduced residual skew from 1.01 to 0.09 and achieved a held-out R² of 0.896. The random forest model achieved 0.877 on the same held-out counties, and a three-hidden-layer neural network provided no further improvement. Therefore, greater model flexibility did not enhance held-out performance once purchasing power was represented proportionally. The contribution lay in applying those models to test whether the data supported a more complex representation of the relationship.

These contributions resulted in a practical framework for comparing local economic conditions. Purchasing power indicated the extent to which household income extended after accounting for local prices, while task groups described the types of work comprising county employment. Consequently, combining the two revealed differences that income, poverty, or unemployment alone could not fully capture. Counties with similar values on conventional economic indicators could still differ in both their task groups and purchasing power. The framework provided policymakers and economic development organizations with a reproducible method to examine local economic conditions using publicly available data.

Limitations

Five limitations defined the scope of the results. First, the findings were correlational rather than causal. Counties were not randomly assigned to task groups, and the analysis could not isolate task groups from all other characteristics that varied across places. Unobserved factors related to both county employment and purchasing power could therefore contribute to the estimated relationship. The results showed that task groups remained associated with purchasing power after accounting for the measured controls, but they did not establish that differences in task groups caused differences in purchasing power.

Second, the analytical panel was limited to counties above the population threshold required for the occupational data. The final sample included 840 counties, representing approximately 84% of the U.S. population, but smaller and more rural counties were underrepresented. The estimates described the more populous counties included in the panel rather than all U.S. counties. Relationships in counties below the threshold could differ from those observed in the analytical sample.

Third, the geographic precision of RPPs varied across counties. A majority of panel observations relied on a state-level price parity rather than a more local measure, which made purchasing power less geographically precise for those observations. Some within-state differences in local prices were not captured by the outcome. The analysis applied these measures consistently but did not separately test whether the association differed according to the geographic precision of the price measure.

Fourth, purchasing power was measured in nominal dollars, which were not adjusted for changes in the national price level over time. RPPs corrected for price differences within a given year but did not place dollars from different years on a common scale. The year indicators absorbed any shared national price change across all counties. The task group and control coefficients were unaffected. The year coefficients themselves combined inflation with other national conditions that fluctuated between years. They could not be read as a measure of inflation alone. Consequently, dollar comparisons across years represented nominal amounts rather than constant purchasing power.

Finally, the primary panel regression did not include state indicators. Unmeasured state-level differences could have contributed to the estimated associations. The predictive models provided a partial check by including state indicators alongside task groups, poverty, unemployment, and year. Permuting the state indicators reduced the held-out R² by only 0.008, indicating that they contributed relatively little predictive information in that specification. However, the predictive models and panel regression were not identical, and this result did not eliminate the possibility of state-level confounding in the primary estimates.

Ethical considerations

The study required clear boundaries on how the results were interpreted. Since the county was the unit of analysis, the findings described places rather than individual workers. Applying county-level relationships to individuals would constitute an ecological fallacy. Consequently, the results could not be used to infer a worker’s purchasing power, economic vulnerability, or likelihood of being affected by automation based on their task group.

Representation also necessitated caution. Counties were excluded because occupational data were unavailable below the source publication threshold. The publication rule set the sample, and the analysis inherited it. As a result, smaller and more rural counties were underrepresented in the panel. Extending the findings to those counties would require extrapolation beyond the observed data, which was particularly important when comparing geographic patterns or identifying places that might warrant further attention.

The data also carried a selection bias that the analysis could not eliminate. The publication threshold excluded smaller counties before constructing the analytical sample, resulting in a panel that was not a random representation of all U.S. counties. Consequently, rural and less populous counties were underrepresented, making the results more accurate for the included counties than for those excluded by the threshold. This limitation matters because smaller rural counties are often central to discussions about automation and local economic vulnerability.

The associational nature of the findings also limited their use for policy decisions. Treating the estimated relationships as causal could lead to the allocation of resources to change a county’s task groups even if an unmeasured economic or geographic factor contributed to the observed difference. While the results supported identifying where lower purchasing power and specific task groups coexisted, they did not establish which intervention would alter those conditions. Similarly, the language used to describe counties required similar care. Labeling places based on presumed automation risk or economic vulnerability could stigmatize communities or deter investment based on relationships that the study did not establish causally.

Finally, the study relied entirely on publicly available aggregate data. No individual-level records or personally identifiable information entered the analytical dataset, and all reported results described county-level patterns rather than individual people.

Future directions

Three directions emerge from these limitations. The first is the coverage gap, which could potentially be bridged using satellite imagery. Jean et al. (2016)’s estimate of local economic conditions from daytime and nighttime imagery, where survey data are scarce, could be applied here. This approach could generate purchasing power estimates for counties below the ACS publication threshold and assess whether the geographic patterns observed in this study extend to areas outside the analytical panel. This would not expand the task group association itself, because comparable occupational estimates would still be unavailable for those counties. It would still show whether the spatial patterns reported here reflect the country more broadly rather than the more populous counties in the sample.

A second direction is to extend the study beyond 2008-2023 as additional occupational, income, price, poverty, and unemployment data become available. The current analysis revealed that the ordering of the task group associations remained stable across the entire panel and across 2010-2019 and 2021-2023. Extending the panel would directly test whether these relationships persist as local labor markets continue to evolve. Additionally, it would allow the results reported here to be evaluated out of time using observations that were not available when the original models were developed. Replicating the direction, magnitude, and proportional form of the associations in later years would provide stronger evidence that the patterns identified in this study were enduring rather than specific to the observed period.

The third approach is to investigate the underlying reasons for the correlation between task groups and purchasing power. Establishing a causal relationship would necessitate identifying a source of variation in task groups that was not influenced by local purchasing power. This could be achieved through factors such as plant openings and closures, major changes in production technology, or identifiable technology adoption shocks. Such an analysis would require a separate study with distinct data and a research design capable of isolating plausible exogenous changes in local task groups. Such an analysis would go beyond confirming the association observed in this study. It would show whether changes in the types of work performed within a county are accompanied by shifts in purchasing power.

Conclusion

This study discovered a consistent link between county task groups and purchasing power. Counties with higher concentrations of routine and manual work generally had lower purchasing power compared to those with more non-routine cognitive work, even after accounting for factors like poverty, unemployment, population, and year. The largest difference was observed for non-routine manual work. At the panel mean, a one percentage point shift from non-routine cognitive to non-routine manual work was associated with approximately $846 lower purchasing power, with a 95% confidence interval of ±$38.

The relationship persisted across both time and model specifications. County task group patterns remained relatively stable over the fifteen-year panel, and the ordering of the estimated associations remained consistent across the examined periods. Diagnostic testing revealed that the relationship was better represented proportionally than as a constant dollar difference. The log linear specification also achieved stronger held-out performance compared to the random forest, while the neural network provided no further improvement. These results suggested that additional model complexity was unnecessary to capture the primary relationship observed in the data.

The findings did not establish a causal relationship between task groups and purchasing power, nor did the study directly measure automation adoption or job displacement. Instead, the results indicated that the types of work concentrated within a county provided information about purchasing power beyond poverty, unemployment, population, and year. While poverty remained an important indicator of local economic conditions, it did not fully explain the association between task groups and purchasing power. Counties with higher concentrations of routine and manual work consistently exhibited lower purchasing power across geography and time. Therefore, understanding differences in purchasing power across counties required considering household income, local prices, and the types of work that characterized the local economy.

References

Abadi, Martín, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, et al. 2016. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems.” https://www.tensorflow.org/.
Acemoglu, Daron, and David Autor. 2011. “Skills, Tasks and Technologies: Implications for Employment and Earnings.” In Handbook of Labor Economics, edited by David Card and Orley Ashenfelter, 4:1043–1171. Elsevier. https://doi.org/10.1016/S0169-7218(11)02410-5.
Autor, David H., and David Dorn. 2013. “The Growth of Low-Skill Service Jobs and the Polarization of the US Labor Market.” American Economic Review 103 (5): 1553–97. https://doi.org/10.1257/aer.103.5.1553.
Autor, David H., Frank Levy, and Richard J. Murnane. 2003. “The Skill Content of Recent Technological Change: An Empirical Exploration.” The Quarterly Journal of Economics 118 (4): 1279–1333. https://doi.org/10.1162/003355303322552801.
Chiripanhura, Blessing. 2011. “Median and Mean Income Analyses: Their Implications for Material Living Standards and National Well-Being.” Economic & Labour Market Review 5: 45–63. https://doi.org/10.1057/elmr.2011.17.
Clarke, Erik, Scott Sherrill-Mix, and Charlotte Dawson. 2025. Ggbeeswarm: Categorical Scatter (Violin Point) Plots. https://CRAN.R-project.org/package=ggbeeswarm.
Couch, Kenneth A., and Dana W. Placzek. 2010. “Earnings Losses of Displaced Workers Revisited.” American Economic Review 100 (1): 572–89. https://doi.org/10.1257/aer.100.1.572.
Cozzi, Marco, and Giulio Fella. 2016. “Job Displacement Risk and Severance Pay.” Journal of Monetary Economics 84: 166–81. https://doi.org/10.1016/j.jmoneco.2016.11.001.
Curran, Leah Beth, Harold Wolman, Edward W. Hill, and Kimberly Furdell. 2006. “Economic Wellbeing and Where We Live: Accounting for Geographical Cost-of-Living Differences in the US.” Urban Studies 43 (13): 2443–66. https://doi.org/10.1080/00420980600970698.
Harris, Charles R., K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, et al. 2020. “Array Programming with NumPy.” Nature 585 (7825): 357–62. https://doi.org/10.1038/s41586-020-2649-2.
Jean, Neal, Marshall Burke, Michael Xie, W. Matthew Davis, David B. Lobell, and Stefano Ermon. 2016. “Combining Satellite Imagery and Machine Learning to Predict Poverty.” Science 353 (6301): 790–94. https://doi.org/10.1126/science.aaf7894.
Jordahl, Kelsey, Joris Van den Bossche, Martin Fleischmann, Jacob Wasserman, James McBride, Jeffrey Gerard, Jeff Tratner, et al. 2020. “Geopandas/Geopandas.” https://doi.org/10.5281/zenodo.3946761.
McKinney, Wes. 2010. “Data Structures for Statistical Computing in Python.” In Proceedings of the 9th Python in Science Conference, 56–61. https://doi.org/10.25080/Majora-92bf1922-00a.
Moretti, Enrico. 2013. “Real Wage Inequality.” American Economic Journal: Applied Economics 5 (1): 65–103. https://doi.org/10.1257/app.5.1.65.
Pebesma, Edzer. 2018. “Simple Features for R: Standardized Support for Spatial Vector Data.” The R Journal 10 (1): 439–46. https://doi.org/10.32614/RJ-2018-009.
Pedersen, Thomas Lin. 2024. Patchwork: The Composer of Plots. https://CRAN.R-project.org/package=patchwork.
Pedregosa, Fabian, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, et al. 2011. “Scikit-Learn: Machine Learning in Python.” Journal of Machine Learning Research 12: 2825–30. https://jmlr.org/papers/v12/pedregosa11a.html.
R Core Team. 2024. R: A Language and Environment for Statistical Computing. Vienna, Austria: R Foundation for Statistical Computing. https://www.R-project.org/.
Saad, Lydia. 2023. “More U.S. Workers Fear Technology Making Their Jobs Obsolete.” Gallup. https://news.gallup.com/poll/510551/workers-fear-technology-making-jobs-obsolete.aspx.
Seabold, Skipper, and Josef Perktold. 2010. “Statsmodels: Econometric and Statistical Modeling with Python.” In Proceedings of the 9th Python in Science Conference, 92–96. https://doi.org/10.25080/Majora-92bf1922-011.
Slowikowski, Kamil. 2026. Ggrepel: Automatically Position Non-Overlapping Text Labels with Ggplot2. https://CRAN.R-project.org/package=ggrepel.
Smith, Aaron, and Monica Anderson. 2017. “Automation in Everyday Life.” Pew Research Center. https://www.pewresearch.org/internet/2017/10/04/automation-in-everyday-life/.
U.S. Bureau of Economic Analysis. 2024. “Regional Price Parities by State and Metro Area.” https://www.bea.gov/data/prices-inflation/regional-price-parities-state-and-metro-area.
U.S. Bureau of Labor Statistics. 2024. “Local Area Unemployment Statistics.” https://www.bls.gov/lau/.
U.S. Census Bureau, American Community Survey Office. 2024. “American Community Survey 1-Year Estimates.” https://www.census.gov/programs-surveys/acs.
U.S. Census Bureau, Geography Division. 2023. “Core Based Statistical Area Delineation Files.” https://www.census.gov/geographies/reference-files/time-series/demo/metro-micro/delineation-files.html.
U.S. Census Bureau, Population Division. 2023. “Vintage 2023 Population Estimates.” https://www.census.gov/programs-surveys/popest.html.
U.S. Census Bureau, Small Area Estimates Branch. 2024. “Small Area Income and Poverty Estimates.” https://www.census.gov/programs-surveys/saipe.html.
Virtanen, Pauli, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, et al. 2020. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python.” In Nature Methods, 17:261–72. https://doi.org/10.1038/s41592-019-0686-2.
Wickham, Hadley. 2016. Ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York. https://ggplot2-book.org/.
Wickham, Hadley, Thomas Lin Pedersen, and Dana Seidel. 2023. Scales: Scale Functions for Visualization. https://CRAN.R-project.org/package=scales.

Software

Data processing, statistical modeling, and machine learning were conducted in Python using pandas, NumPy, SciPy, statsmodels, scikit-learn, TensorFlow, and GeoPandas (McKinney 2010; Harris et al. 2020; Virtanen et al. 2020; Seabold and Perktold 2010; Pedregosa et al. 2011; Abadi et al. 2016; Jordahl et al. 2020). Figures were produced in R using ggplot2, patchwork, scales, sf, ggrepel, and ggbeeswarm (R Core Team 2024; Wickham 2016; Pedersen 2024; Wickham, Pedersen, and Seidel 2023; Pebesma 2018; Slowikowski 2026; Clarke, Sherrill-Mix, and Dawson 2025).

Footnotes

  1. The District of Columbia and Kalawao County, Hawaii, are excluded because they are absent from the county reference file. Connecticut counties leave the panel after 2021 because the state replaced its legacy counties with planning regions beginning in 2022.↩︎

  2. All 104 missing observations came from Connecticut’s eight legacy counties, where unemployment rates were unavailable at the required geography (U.S. Bureau of Labor Statistics 2024). Because the missingness reflected geography rather than unemployment itself, it was treated as missing at random.↩︎