Introduction
According to Gallup, 22% of American workers worried technology would make their job obsolete in 2023, up seven percentage points from 2021 (Saad 2023). Concerns about automation were evident years earlier. In 2017, the Pew Research Center reported that 72% of Americans surveyed worried about a future in which technology could perform many jobs currently done by humans (Smith and Anderson 2017). Underlying these concerns was automation, in which technology assumes tasks that were previously performed by human workers. These concerns have a clear economic basis. Job displacement can result in large and long-lasting reductions in worker earnings (Couch and Placzek 2010).
A study found that displaced workers’ earnings fell by more than 30% from 1993 to 2004, remaining about 15% lower six years later (Cozzi and Fella 2016). The burden of those earnings losses depended in part on local prices. As demonstrated by Moretti (2013), the same income could support different living standards in different cities in the United States. This variation also makes the geographic scale of measurement important. As a result, a single value representing an entire state or metropolitan area can obscure the variations in prices and household resources across different communities. For this reason, measuring these conditions locally provides a more realistic view of where households face greater economic strain (Curran et al. 2006).
Examining whether these concerns corresponded with local economic conditions required a measurable indicator of automation risk. Autor’s task based framework provided the basis for measuring automation risk through the tasks performed in local labor markets. Autor, Levy, and Murnane (2003) first demonstrated that computers replaced routine tasks while complementing non-routine ones, which made it possible to identify work more susceptible to automation. Acemoglu and Autor (2011) later expanded the framework by treating tasks as the production unit and different skill groups as their suppliers. Autor and Dorn (2013) then applied the framework to local labor markets, finding that areas concentrated in routine work experienced employment polarization.
Nevertheless, these studies generally examined wages and employment at geographic scales broader than the county. A wage describes what a job paid, while purchasing power describes what that pay could buy. That distinction left unresolved whether the task groups present in local employment are associated with household purchasing power at the county level.
Building on this task based framework, we investigated the relationship between county automation and purchasing power, which is defined as household income adjusted for variations in local price levels. The analysis used a county-level panel spanning from 2008 to 2023, excluding the year 2020, encompassing 11,983 observations across 840 counties. These counties collectively represented a substantial portion of the U.S. population, reaching approximately 84% (U.S. Census Bureau, Population Division 2023). County employment was organized into four task groups to capture differences in the kinds of work performed across local economies.
We conducted a panel regression analysis to investigate the relationship between purchasing power and county task groups, while controlling for factors such as poverty, unemployment, and population. Year indicators accounted for changes over time, and standard errors were clustered by county for statistical inference. Diagnostic checks favored a log specification, while random forest and neural network models assessed whether nonlinear relationships captured patterns missed by the regression model. The most notable finding was that a one percentage point shift from non-routine cognitive to non-routine manual work was associated with approximately $846 less purchasing power at the panel mean (95% confidence interval, ±$38). Task groups remained associated with purchasing power after accounting for poverty and unemployment. These findings have practical implications for policymakers and economic development agencies by helping identify counties that appear similar on conventional indicators but differ in what residents can afford.
Background
Autor developed a task-based approach to measure the variation in automation risk across different types of work. Initially, Autor, Levy, and Murnane (2003) distinguished between routine and non-routine tasks by demonstrating that technology impacts tasks differently based on their potential to be formalized into explicit procedures. Acemoglu and Autor (2011) further expanded his framework by treating tasks as production units and analyzing how workers with varying skills are distributed among them. This progression laid the conceptual foundation for the four task groups depicted in Figure 1, linking technological advancements to shifts in the demand for different types of work.
Building upon this framework, Autor and Dorn (2013) applied the task-based approach to local labor markets. Areas with higher initial concentrations of routine work experienced more pronounced employment polarization. Employment growth was observed at both the upper and lower ends of the wage distribution, while it declined in the middle. These patterns indicated that the local impacts of technological change varied depending on the types of tasks concentrated within each economy. This provided a foundation for exploring whether differences in county task groups were also linked to variations in purchasing power.
Once occupations are categorized based on their task content, it becomes easier to measure differences across local economies. The occupation taxonomy developed by Autor and Dorn (2013) served as the foundation for assigning employment to the four task groups used in this study. Their own measure scored each occupation continuously on a single routine task intensity index built from ratings of its task content. Acemoglu and Autor (2011) provided economic significance to these differences by treating tasks as units of production supplied by workers possessing varying skills. Technological advancements can shift labor demand across different types of tasks. These differences in county task groups provide a way to examine whether local work structures are associated with purchasing power.
Accurately measuring purchasing power requires accounting for differences in local price levels. Regional Price Parities (RPPs), published by the Bureau of Economic Analysis (BEA), offer a measure of these differences relative to the national average. A score of 100 in the RPP represents the national price level. Relative to this baseline, values above 100 indicate higher local price levels, while values below 100 indicate lower local price levels. The index combines prices for goods and services such as housing, food, transportation, medical care, education, and recreation. RPPs became an official BEA statistic in 2014, with estimates available back to 2008. However, BEA does not publish RPPs directly for individual counties. Instead, estimates are available for states, metropolitan areas, and nonmetropolitan portions of states, which can mask price differences among counties within the same region.
Multiple indicators collectively capture different aspects of local economic conditions. Poverty measures the proportion of residents living below an established income threshold. Unemployment measures the share of the labor force actively seeking work. Median household income represents the midpoint of the household income distribution, and RPPs compare local price levels with the national average. Each captures a different dimension of local conditions, and the four produce distinct geographic patterns across counties (Figure 2).
Geographic scale presented an additional measurement challenge because broader geographic units could obscure economic differences between nearby counties. Local prices and household resources varied within the same labor market, while earlier studies often combined several counties into a single regional measure. This aggregation could mask meaningful local variation that remained visible at the county level. Curran et al. (2006) showed that accounting for geographic cost of living differences alters how places compare on economic wellbeing, which makes the choice of geographic unit consequential for any measure of household resources. By measuring purchasing power at the county level, we preserved more of this variation by incorporating both household income and local price levels. Additionally, we evaluated whether the association between purchasing power and county task groups was better represented in dollar terms or proportional terms. This approach addressed a gap in prior research, which examined task groups in relation to employment and wages without evaluating their association with purchasing power at the county level.
Data
Data sources and study period
This study integrated five federal datasets published by three agencies, summarized in Table 1, along with the number of records and variables provided by each source. Most sources reported data at the county level, and they were joined using the county Federal Information Processing Standards (FIPS) code and the year. RPPs were published at the metropolitan area level and required an additional matching step. Counties were first linked to their metropolitan areas using the Census Bureau Core Based Statistical Area (CBSA) delineation file and then matched to the corresponding BEA price measure (U.S. Bureau of Economic Analysis 2024; U.S. Census Bureau, Geography Division 2023).
| Data Source | Abbreviation | Records | Measures | Destination Table |
|---|---|---|---|---|
| Small Area Income and Poverty Estimates | Census SAIPE | 50,283 | median household income, poverty rate | county_baseline |
| Bureau of Economic Analysis Regional Price Parities | BEA RPP | 884 | county price level relative to the national average | cbsa_rpp |
| Core Based Statistical Area delineation | Census CBSA | 387 | county to metropolitan area crosswalk | cbsa |
| American Community Survey 1 year estimates | Census ACS | 15,690 | population, occupational employment by category | county_baseline, county_task_exposure |
| Bureau of Labor Statistics Local Area Unemployment Statistics | BLS LAUS | 88,004 | unemployment rate | county_baseline |
| Scale: 3,143 counties per year, 15 years (2008 to 2023, excluding 2020), 47,140 county year universe, 11,983 rows in the analytical panel. | ||||
| Records count rows retrieved from each source rather than analytical observations. The ACS publishes 1 year estimates only for areas above 65,000 residents, which reduces the universe to the 848 county panel. | ||||
| All sources retrieved June 2026 through agency APIs or bulk file download. | ||||
The study spans 2008 through 2023 but excludes 2020, because the American Community Survey (ACS) did not publish data due to disruptions of data collection during the COVID-19 pandemic (U.S. Census Bureau, American Community Survey Office 2024). Those estimates served as the occupational employment counts used to construct the task groups. Over the subsequent fifteen years, the Small Area Income and Poverty Estimates (SAIPE) program provided an initial sample of 47,140 county-year observations (U.S. Census Bureau, Small Area Estimates Branch 2024).1
Analytical sample
The initial sample restriction stemmed from the coverage of the ACS occupational estimates. The Census Bureau published ACS 1 year occupational estimates exclusively for counties with populations of at least 65,000. Consequently, counties below this threshold lacked the employment counts necessary to construct the four task groups. Consequently, applying this restriction reduced the sample to 12,087 county-year observations across 848 counties. A subsequent restriction mandated the inclusion of complete model covariates, resulting in the removal of 104 observations with missing unemployment rates.2 However, population, median household income, and poverty rate data were complete throughout the sample.
The final analytical sample comprised 11,983 county-year observations from 840 counties. All models presented in Section 4 and Section 5 used these same observations. To ensure unbiased machine learning evaluation, counties were exclusively included in either the training or test set, preventing observations from the same county from appearing in both partitions. Section 4 provides a detailed description of the evaluation procedure.
The population threshold imposed a limitation on the scope of the results. While the 840 retained counties accounted for approximately 84% of the United States population, they represented a minority of all counties. Consequently, rural counties were underrepresented because the retained counties were generally more populous and metropolitan. The results were applicable to counties above the ACS publication threshold rather than the entire nation, a limitation discussed in Section 6.
Construction of task group measures
Following the framework introduced in Section 2, we classified broad occupational categories in the ACS 1 year estimates into four task groups, applying the task type distinctions of Autor, Levy, and Murnane (2003) and the occupation taxonomy of Autor and Dorn (2013). For instance, clerical and sales occupations were categorized as routine cognitive tasks, while service occupations were classified as non-routine manual tasks. Subsequently, we summed employment within each task group for every county and year.
\[ T_{g,c,t} = \sum_{o \in g} E_{o,c,t} \tag{1}\]
Here, \(T_{g,c,t}\) represented total employment in task group \(g\) for county \(c\) in year \(t\). The sum ran over the occupational categories \(o\) assigned to that group, and \(E_{o,c,t}\) represented employment in a single category. Since the calculation used employment counts without task intensity weights, each total represented the number of jobs in that group. Scoring each occupation continuously, as Autor and Dorn (2013) did, would have required employment counts for detailed occupations, which the Census Bureau did not publish for counties. Every job within a broad category therefore carried equal weight, and variation in task content inside those categories went unmeasured. Dividing each group’s total by the total across all four groups, as shown in Equation 2, converted these totals into proportions between zero and one that summed to one.
\[ G_{g,c,t} = \frac{T_{g,c,t}}{\sum_{g'} T_{g',c,t}} \tag{2}\]
Here, \(G_{g,c,t}\) represented the proportion of employment in task group \(g\), and \(g'\) indexed the four groups summed in the denominator.
These proportions assigned counties of varying sizes to a common scale. Since the four proportions added up to one, non-routine cognitive work was excluded from the regression as the reference group, and the remaining coefficients were interpreted in relation to it. We maintained the four task groups as separate entities because a single routine intensity index would obscure differences between types of work. For instance, a shift away from clerical work and a shift away from assembly work could otherwise appear identical. Section 4 provides an explanation of how the task groups were incorporated into the models.
Additionally, the Census Bureau revised its occupational classification between 2009 and 2010. While some category boundaries changed, their assignment to the four task groups remained consistent. Section 4 reports the corresponding robustness checks.
Construction of the purchasing power measure
Median household income was used because it represented the income of a typical household better than the mean, which was more sensitive to high incomes in a positively skewed distribution (Chiripanhura 2011). It described what households received rather than what they could buy.
The outcome was measured as purchasing power to account for variations in what household income could purchase across different counties. We calculated it by adjusting the county median household income from SAIPE (U.S. Census Bureau, Small Area Estimates Branch 2024) using RPPs. These provided local price levels relative to the national average (U.S. Bureau of Economic Analysis 2024). Since the index was expressed as a percentage, dividing it by 100 converted it into a price multiplier. Subsequently, we divided the median household income by this multiplier to determine purchasing power, as illustrated in Equation 3.
\[ A_{c,t} = \frac{\text{Median household income}_{c,t}}{\text{RPP}_{c,t} / 100} \tag{3}\]
Here, \(A_{c,t}\) represented purchasing power for county \(c\) in year \(t\), and \(\text{RPP}_{c,t}\) represented the RPP for that same county and year.
For instance, consider a county with a median household income of $60,000 and an RPP of 120. This means that its income could buy approximately what $50,000 would buy at the national average prices. This adjustment made household income more comparable across counties by accounting for variations in local prices. However, since BEA did not publish separate RPPs for counties outside metropolitan areas, those counties were assigned the corresponding state-level parity. Within the analytical panel, 45.5% of county-year observations used a metropolitan parity, while 54.5% used the state measure. Consequently, more than half of the panel relied on a state average that did not account for price differences within the state.
Inflation and year indicators
RPPs were adjusted for variations across regions but not for changes in the national price level over time. Consequently, purchasing power was reported in nominal U.S. dollars. A national deflator would apply the same annual adjustment to every county. In the log specification, this adjustment would be incorporated into the year indicators, leaving the task group and control coefficients unchanged. Additionally, the year indicators captured other changes that occurred simultaneously across counties within a year, making it impossible to interpret their coefficients solely as inflation. Section 4 provides an explanation of how the year indicators were incorporated into the model, while Section 6 addresses this limitation.
Control variables
The analysis incorporated poverty rate, unemployment rate, and population as county-level covariates. Poverty and unemployment captured the economic distress mentioned in Section 2, enabling the analysis to determine if task groups provided information beyond these measures. Median household income directly entered the purchasing power outcome, so it was not included as a control variable. Population accounted for variations in county size, which could otherwise affect the estimates.
Figure 3 displayed the distributions of poverty rate, unemployment rate, population, and median household income across the same task group panel used in Figure 9. All four variables were right skewed, but population was the most extreme and therefore entered the analysis in logarithmic form. Poverty rate, unemployment rate, and median household income remained in their original units. This supported the log respecification in Section 4 as a response to the outcome’s own distribution rather than a transformation applied across all variables.
Data architecture and organization
The data architecture separated raw source records from the tables used for analysis. Raw retrievals and intermediate results were stored in a PostgreSQL data lake containing 41 tables, which remained unnormalized to maintain traceability. Processed data were then written to an analytical warehouse containing eleven tables. In this warehouse, county identifiers were standardized, records were restricted to the study period, jurisdictions outside the panel were removed, and foreign key constraints were enforced.
Figure 4 provided a diagram that summarized the path from federal sources through the data lake and warehouse to the analysis-ready tables.
The warehouse is organized around a county table, which is keyed by FIPS code and linked to the state as a reference table. Three analytical tables join to each county by FIPS code and year. The county_baseline table contains population, median household income, poverty rate, and unemployment rate. The county_task_exposure table contains the four task group totals defined in Equation 1, while the county_affordability table contains the purchasing power outcome.
Analysis
The unadjusted relationship
The analysis commenced by examining the unadjusted relationship between task groups and purchasing power. Across the panel, the two manual groups exhibited the most pronounced negative relationships, while non-routine cognitive tasks showed a positive correlation, and routine cognitive tasks remained relatively flat (Figure 5). None of these relationships were strong enough to stand alone, prompting the subsequent adjustment analysis.
Variance decomposition
Purchasing power and task groups exhibit variations both across counties and within a county over time, each carrying distinct implications. Each group’s variation divides into these two components (Figure 6). The manual groups show almost complete variation between counties, which contributes to the durability of the purchasing power differences they represent. In contrast, routine cognitive is the only group with substantial within-county movement.
The Census occupation coding change between 2009 and 2010, as described in Section 3, inflates the apparent variation in routine cognitive occupations within counties. This inflation arises from the coding of the source data rather than any inherent property of the counties. The figure is computed on a panel restricted to 2010 onward, with each year’s cross-county mean removed.
Model specification
Non-routine cognitive tasks served as the reference group because the four task group proportions summed to one. Each remaining coefficient therefore represented the change in purchasing power associated with a shift from non-routine cognitive work to that group, while holding the other variables constant. This made the coefficients interpretable as comparisons between task groups rather than as independent changes in employment. Task group values ranged from zero to one, and coefficients were divided by 100 when reported as percentage point changes.
Year indicators captured the shared conditions across all counties within a specific year, encompassing events like the 2008 recession, its subsequent recovery, and the national price movements described in Section 3. By including these indicators, we prevented attributing national changes to variations in county task groups. Since the year coefficients absorbed all common changes within a year, they were treated as nuisance parameters rather than evidence of increasing purchasing power. The year coefficients traced this pattern across the panel (Figure 7).
Standard errors were clustered by county because repeated observations within the same county over fifteen years were not independent. A Durbin Watson statistic of 0.496 indicated this dependence. Clustering allowed observations within a county to be correlated, which prevented the uncertainty around the estimated associations from being underestimated. The estimated coefficients are reported in Section 5.
Collinearity and the reference group
The unnormalized employment totals of Equation 1 are not suitable for regression analysis. A variance inflation factor (VIF) quantifies the extent to which a coefficient is affected by its correlation with other inputs, and a value exceeding 10 is considered a serious concern. The four employment totals exhibit factors ranging from 13 to 27, as they all represent employment counts that are proportional to county size and therefore exhibit a strong correlation. When fitted on these totals, the coefficients were unstable, and their standard errors were notably large compared to their magnitudes, which are typical indicators of multicollinearity. The proposed specification addresses this issue twice (Figure 8). Normalizing the data by the four group total, as described in Section 3, eliminates the shared scale, while using non-routine cognitive as the reference group, as established above, removes the constraint that the four proportions must sum to one. Consequently, every factor in the estimated specification falls within the range of 1.2 to 1.6.
Specification diagnostics
Distribution of the analytical variables
The analysis next delved into the distributions of the variables used in the specification. Purchasing power and the four task groups spread across the panel with skewness and outliers that guided the diagnostic checks that followed (Figure 9). Purchasing power exhibited right skewness, with a skewness coefficient of 1.17 and a central tendency near $57,000. A long upper tail extended beyond the center of the distribution, so the extremes carry most of the information about model fit.
Residual and quantile diagnostics
A linear model assumes a constant rate of association across the data range and symmetric, evenly distributed errors around the fit. Three diagnostics test these assumptions: residuals plotted against fitted values, a quantile-quantile (QQ) plot of the residuals, and the variance inflation factors already reported. The raw scale model shows clear systematic curvature in the residuals and a marked upper tail departure in the QQ plot (Figure 12, Panels A and B). The model underestimates both the lowest and highest purchasing power counties while fitting the middle well, and the upper tail departure indicates a residual skew of 1.01.
The level model still provides the interpretable dollar estimates we report. However, its own diagnostics reveal that the data exhibit structure that a straight line in dollars cannot represent. The curvature in the residual smoother and the widening spread with fitted values suggest a specific failure in the fixed slope specification, and each subsequent model relaxes the assumption behind it.
Influential points
A fourth diagnostic, Cook’s distance, evaluates whether any single county year disproportionately influences the level specification’s coefficients. None of the observations exceed the conventional threshold of 1.0, so no point warrants further investigation. The highest value recorded is 0.009. Every county year appears in the plot of leverage against standardized residual, with dashed curves tracing constant Cook’s distance (Figure 10). The conventional 0.5 and 1.0 contours fall well outside the plotted range, since reaching 1.0 at the highest observed leverage would require a standardized residual near 40. The labeled points are the four county years with the largest standardized residuals. Loudoun and Stafford County, Virginia, are both high purchasing power counties located in the Washington D.C. exurbs. Williamson County, Tennessee, is a high purchasing power Nashville exurb. Apache County, Arizona, a low purchasing power county situated on the Navajo Nation, sits among the marked points at a lower residual. These counties are the same high-end outliers already visible in the purchasing power distribution (Figure 9).
No single point threatened the overall fit, so the more informative check was to refit the model after removing the most influential 1% of the panel, 120 county year observations. Table 2 reported the results. The task group coefficients changed by single digit to low double digit percentages, while the poverty coefficient remained relatively stable. This indicated that the main associations were not driven by a small number of extreme counties.
Unemployment rate was the exception. As reported in Section 5, it was the smallest and least precisely estimated control, and its coefficient changed by approximately one third. A second stability check examined whether the main estimates were consistent over time. The model was refit separately for the pre pandemic period from 2010 to 2019 and the post pandemic period from 2021 to 2023, with Section 5 reporting the stability of the coefficients across the two periods.
| Variable | Full Sample | Excluding Top 1% | Change (%) |
|---|---|---|---|
| Routine Cognitive | −$85,791 | −$75,273 | −12.3 |
| Routine Manual | −$70,097 | −$62,441 | −10.9 |
| Non-Routine Manual | −$104,139 | −$98,410 | −5.5 |
| Poverty Rate | −$1,605 | −$1,616 | 0.7 |
| Unemployment Rate | $206 | $143 | −30.6 |
| Log Population | $498 | $571 | 14.7 |
| 120 of 11983 county year observations excluded, the top 1% by Cook's distance. | |||
Log respecification
The level model indicated that a constant dollar association was insufficient to accurately describe the relationship. Logging purchasing power was tested to see whether the association was better represented proportionally, enabling a comparable shift in a task group to correspond to a larger dollar difference in counties with higher purchasing power. The transformation reduced the right skew in purchasing power and laid the groundwork for reevaluating the residual pattern (Figure 11).
Refitting the same specification using log purchasing power resolved most of the problems identified in the level model. Systematic curvature in the residuals lessened, while the QQ points followed the reference line more closely after the transformation (Figure 12). Some tail deviation remained, but residual skew declined from 1.01 to 0.09. The log linear specification therefore provided the baseline for evaluating the more flexible models that followed.
The same comparison also holds in the units of the outcome (Figure 13). On the level scale, the model closely tracks purchasing power throughout the middle of the distribution, where most counties are located. However, both tail bins deviate substantially from their predictions, coming in approximately $8,000 to $9,000 above the expected values. On the log scale, the points closely follow the line throughout the distribution.
The residual pattern did not disappear entirely after the transformation. The remaining structure could reflect nonlinearity, interactions among the variables, or omitted factors that the diagnostics alone could not distinguish. This motivated the use of a random forest, which could capture nonlinear relationships and interactions without specifying them in advance. The log linear specification served as the benchmark, allowing us to test whether a more flexible model captured additional signal in held out counties.
Predictive modeling
The predictive comparison used the county year panel assembled earlier, incorporating state identifiers for the subsequent models. Model complexity only increased when additional flexibility enhanced performance on the held-out counties. The log respecification addressed the issues identified in the level model, while the remaining residual structure prompted testing a random forest capable of capturing nonlinear relationships and interactions. A three hidden layer feedforward neural network served as an additional test to determine if greater model flexibility further improved performance.
All models were evaluated on the same held-out counties from the grouped split, so any differences in performance reflected model structure rather than variations in the evaluation observations. If the more flexible models failed to improve held-out performance, the log linear specification was retained as the simpler representation of the relationship.
Evaluation design
Feature set and reference groups
Consistent with the reference-group approach established in Section 4.2, non-routine cognitive tasks were excluded from the predictive specification as well, maintaining consistency in interpreting the task group coefficients across both inferential and predictive analyses.
State indicators also required a comparable reference category. Mean purchasing power varied widely across states, and a state situated near the center of that distribution served as the reference point (Figure 14). By selecting a state near the middle, comparisons were avoided to states with exceptionally high or low purchasing power, simplifying the interpretation of the resulting coefficients. While the choice of reference state did not affect the overall model fit, it influenced the expression of the state coefficients.
Grouped train and test split
This is not a forecasting exercise, and a time ordered split is not required for its usual reason. A grouped split is still necessary, because a county’s economic profile barely moves year to year, which makes its 2018 and 2019 rows near duplicates. If two adjacent rows land on opposite sides of the split, the model can recall a county rather than learn a relationship. We group by county and split on row membership before any further feature engineering. The split assigns 9,565 county year observations to training and holds out the remaining 2,418, with no county contributing rows to both. All feature engineering was conducted after the split to maintain this separation. Consequently, every test observation originated from a county that was not represented in the training data.
Year and state indicators
A grouped train and test split was used because observations from the same county were closely related across years. A random split could inadvertently place observations from the same county in both sets, potentially allowing the model to use information about a county it had already encountered. By grouping observations by county, we effectively eliminated this overlap and provided a more robust test of whether the observed relationship could generalize to new locations.
The reference state was Texas, whose mean purchasing power was close to the median of the state distribution (Figure 14). By selecting a state near the center, we avoided anchoring comparisons to an unusually high or low purchasing power state, making the coefficients more interpretable. While most state coefficients remained within a few percent of the reference, the largest differences reached 17%. Once the economic controls and year indicators were included, as demonstrated by the permutation results below, state indicators contributed little to held-out performance.
Collinearity of the indicator blocks
The reference group logic applies here as well (Figure 8). The additional concern was whether the year and state indicators introduced new collinearity after being added to the predictive specification. They did not, as all reported VIFs remained below the conventional threshold of 5 (Figure 15). This confirmed that the indicator blocks could be retained without materially destabilizing the estimated coefficients.
Linear benchmark
This specification is a predictive benchmark rather than a second inferential model. It reuses the diagnostic insights established by the panel regression above, now applied to the training split alone. This shared split enables direct comparability with the subsequent random forest and neural network models. We present two versions of each linear model, estimated using ordinary least squares (OLS). The baseline model uses only the task groups with poverty and unemployment, while the full model incorporates year and state controls. Reporting both versions highlights the changes introduced by time and geography. Among the two, the full model is the one that is carried forward.
Baseline specification
The baseline specification yielded an R² of 0.766. Relative to non-routine cognitive work, the coefficients for the other three task groups were negative, indicating lower purchasing power as counties shifted toward those groups. The largest negative association was observed for non-routine manual work, where a one percentage point shift was associated with approximately $1,305 lower purchasing power. Table 3 reports these estimates alongside those from the full specification.
Full specification
Adding year and state indicators increased the R² from 0.766 to 0.871, indicating that time and geography contributed to additional variation in purchasing power. The unemployment rate coefficient shifted from approximately −$1,207 in the baseline specification to +$450 in the full specification for each one percentage point increase in unemployment. This change in direction suggested that the baseline estimate accounted for geographic and temporal differences that were not previously considered. In contrast, the task group coefficients maintained their direction and relative ordering.
Residual curvature persisted even after the full level specification (Figure 12). This figure was important because the increase in the coefficient alone suggested that incorporating year and state indicators effectively improved the model. However, the diagnostics revealed that while these variables enhanced the overall fit, they failed to eliminate the systematic pattern in the residuals. This distinction indicated that the underlying issue was the functional form of purchasing power on its original dollar scale, rather than omitted time or geographic differences.
Full specification on the logged outcome
The level specification did not fit purchasing power evenly across its range (Figure 13). The largest departures occurred at the lower and upper ends of the distribution, where observed purchasing power was roughly $8,000 to $9,000 above the fitted values. This systematic pattern suggested that the remaining problem was related to the scale of purchasing power rather than omitted controls, providing a clear reason to test the outcome in logarithmic form.
Logging purchasing power reduced the residual pattern observed in Figure 12 and enhanced the performance of the linear specification. After transforming the predictions back to dollars, the log linear model achieved an in-sample accuracy of 0.913 and a held-out R² of 0.896. This close agreement suggested that the model effectively generalized to counties not included in the training data.
Coefficient comparison
Table 4 compared the three specifications. The log specification was chosen for inference because the diagnostics for the level specification indicated that its functional form was unsuitable, making its coefficient estimates less reliable for interpretation.
| Purchasing Power | Log Transformed Purchasing Power | |||
|---|---|---|---|---|
| Baseline | Full | Coefficient | % per 1 pp | |
| Task Groups | ||||
| Routine Cognitive | −$81,085 | −$87,421 | −1.080 | −1.07 |
| Routine Manual | −$71,490 | −$81,143 | −1.089 | −1.08 |
| Non-Routine Manual | −$130,549 | −$91,904 | −1.213 | −1.21 |
| Economic Controls | ||||
| Poverty Rate | −$1,376 | −$1,729 | −0.029 | −2.87 |
| Unemployment Rate | −$1,207 | $450 | 0.001 | 0.08 |
| Non-routine cognitive is the reference task group. Task group coefficients in the three model columns are per unit of group proportion; the final column converts the log coefficients to the percent change in purchasing power associated with a one percentage point shift. | ||||
Table 3 revealed that the task group coefficients maintained the same direction across all three specifications, although their magnitudes changed after incorporating controls. The non-routine manual coefficient shifted from approximately -$130,549 in the baseline specification to -$91,904 in the full level specification. Routine cognitive changed from about -$81,085 to -$87,421, while routine manual changed from -$71,490 to -$81,143. These results demonstrated that adjustments altered the magnitude of the associations without altering their overall order.
The poverty rate coefficient increased in magnitude from approximately -$1,376 to -$1,729 for each additional percentage point of poverty. In contrast, the unemployment rate coefficient changed from approximately -$1,207 in the baseline specification to +$450 in the full level specification. This difference showed that poverty remained strongly associated with purchasing power after adjustment, while the unemployment association was comparatively weak and sensitive to specification.
| Baseline | Full Model | Log Model | |
|---|---|---|---|
| Outcome | Purchasing Power | Purchasing Power | log(Purchasing Power) |
| Year and State Controls | No | Yes | Yes |
| R² | 0.766 | 0.871 | 0.913 |
| Residual Diagnostics | — | Failed | Passed |
| Primary Specification | No | No | Yes |
| R² is in sample, fit on the 9,565 training rows only. | |||
In the log specification, each additional percentage point of poverty was associated with approximately 2.9% lower purchasing power, assuming task groups, time, and geography remained constant. This proportional association was larger than that of any individual task group. Poverty was therefore the strongest single association among the variables reported in Table 3.
Random forest
For the random forest model, hyperparameters were selected using a grid search with grouped five-fold cross-validation on the training counties. The search explored various configurations, including the number of trees ranging from 600 to 1,000, maximum depth from 10 to 20 or unlimited, minimum leaf size from 1 to 5, and the proportion of features considered at each split. Grouping by county ensured that the separation between places during tuning was preserved, reducing the risk of selecting parameters that performed well only due to observations from the same county appearing across folds. The final model used 1,000 trees, unlimited depth, a minimum leaf size of 1, and half of the available features at each split.
The random forest model achieved an R² value of 0.876 on the held-out counties, which was slightly lower compared to the log linear model’s R² of 0.896 on the same held-out counties. This 0.020 difference suggested that the additional flexibility of the random forest model did not substantially enhance its performance. Additionally, the random forest model achieved an R² value of 0.989 on the training counties, resulting in a 0.113 gap between training and held-out performance. This gap indicated that the random forest model captured additional structure in the training data that did not generalize to counties it had not encountered.
No systematic curvature or trend appeared in the random forest residuals across the range of predicted values (Figure 16). This suggested that the random forest effectively captured the nonlinear structure that posed challenges for the level linear model without transforming purchasing power. However, the lower held-out R² value of 0.876 and the large gap between training and test performance indicated that this ability to capture the structure did not translate into improved generalization to unseen counties.
Ablation and permutation importance
The random forest is largely insensitive to whether purchasing power is modeled on its original or logarithmic scale. Fitting the model to the logged outcome and then back-transforming the predictions yields an R² of 0.877, which is slightly higher than the R² of 0.876 obtained when the outcome is modeled directly. Since the log specification is also used for the primary linear model, the log target random forest is used for the subsequent ablations and model comparisons. The interpretability measures, including permutation and partial dependence, are computed on the raw target random forest, which reports directly in dollars.
Ablation analysis revealed the model’s effectiveness in compensating for the removal of an entire group of variables and subsequent refitting from scratch. The removal of task groups resulted in a decline in the held out R² value from 0.876 to 0.834, a decrease of 0.042. Since the random forest was retrained after removing the task groups, the remaining variables had the opportunity to recover any overlapping information. Consequently, the remaining decline indicated that poverty, unemployment, time, and geography could not fully replace the information provided by the task groups.
In comparison, eliminating poverty and unemployment reduced the held out R² from 0.877 to 0.698, resulting in a decline of 0.179. This drop was substantially larger than the 0.042 decline observed when the task groups were removed, suggesting that the economic controls provided more of the model’s explanatory power. Even after incorporating the task groups into the economic controls, performance improved, indicating that they contributed information not fully captured by poverty and unemployment alone. All three models used the same random forest specification, logged outcome, and held-out counties, ensuring direct comparability of their performance differences.
Grouped permutation importance
Joint permutation provided a complementary measure of how strongly the fitted random forest relied on the task groups. Permuting the three modeled task group proportions together reduced the held out R² by 0.153 while leaving the fitted model unchanged. This decline was much larger than the 0.042 reduction from removing the task groups and refitting the random forest.
The difference between the two measures revealed what the model could do after removing task group information. During ablation, the random forest was retrained and could recover some overlapping information from poverty, unemployment, and year. Permutation, on the other hand, disrupted the task group information after the random forest had already learned to use it. The larger permutation decline indicated that the fitted random forest heavily relied on the task groups, even though some of their information overlapped with the other variables.
Year indicators showed a decline of 0.128 when permuted, compared to only 0.008 for state indicators. This difference demonstrated that time contributed considerably to held-out performance compared to state after accounting for economic controls and task groups.
Among the three modeled task groups, non-routine manual produced the largest individual permutation decline. These values were not interpreted as additive or independent contributions because the task groups were mechanically related through their employment proportions. Nevertheless, the ordering was consistent with the regression results, where non-routine manual also had the largest negative task group association with purchasing power.
Neural network
The random forest model did not outperform the log linear specification. Whether that reflected a limitation of the random forest or the amount of structure present in the data remained an open question. To further investigate this, we evaluated a three-hidden-layer feedforward neural network as an additional test to determine if greater model flexibility would lead to improved held-out performance. The network consisted of three hidden layers and a single linear output unit for the continuous purchasing power outcome.
Training loss steadily declined, while validation loss began to rise. This divergence indicated that the network was fitting the training data more closely without improving performance on unseen observations, suggesting overfitting. To address this, we added L2 weight regularization, increased dropout, and shortened the early stopping patience. These changes narrowed the gap between training and validation loss, providing a stronger test of whether the neural network could generalize beyond the training data.
Regularization reduced the disparity between training and validation loss (Figure 17), yet the neural network still failed to surpass the random forest or the log linear model. On the same counties, the random forest achieved a held-out R² of 0.877, while the log linear model achieved 0.896. Consequently, the neural network’s additional flexibility did not offer any improvement over the simpler alternatives. Since the study aimed to elucidate the relationship between task groups and purchasing power, further increases in model complexity were deemed unnecessary. Therefore, the log linear and random forest models were retained for the subsequent comparisons.
Cross validation
Five-fold grouped cross-validation was used to assess the consistency of the observed performance in the single train and test split across various county sets. Counties remained grouped within each fold, ensuring that observations from the same county were not included in both training and validation data.
| Model | Mean R² | SD |
|---|---|---|
| Ordinary Least Squares | 0.9036 | 0.0080 |
| Random Forest | 0.8982 | 0.0060 |
| Both models are fit on the log target and scored on the dollar scale after back transformation. | ||
The Table 5 revealed that the log linear model and random forest achieved comparable mean values across the five folds. The difference in their average performance was smaller than the variation observed across folds, suggesting that the earlier comparison was not influenced by a specific train and test split. Consequently, the neural network was excluded as it failed to enhance held-out performance and was not retained.
Summary
The panel regression provided the primary inferential results, with standard errors clustered by county and controlling for poverty, unemployment, population, and year. Diagnostic checks supported a proportional relationship, justifying the log specification. However, the variance decomposition revealed that most task group variation occurred between counties rather than within counties over time. Predictive models were evaluated on held-out counties, where the log linear model achieved an R² of 0.896 compared to 0.877 for the random forest. The neural network model provided no further improvement. Removing the task groups reduced the held-out R² by 0.042, indicating that poverty, unemployment, time, and geography did not fully capture the information they provided. These results collectively support retaining the log linear specification as the primary representation of the relationship reported in Section 5.
Results
The results unveiled four main findings. Task groups maintained their association with purchasing power even after accounting for factors such as poverty, unemployment, population, and year. A one percentage point shift from non-routine cognitive to non-routine manual work was linked to a reduction in purchasing power of approximately $846, with a 95% confidence interval of ±$38 at the panel mean. The log specification provided the closest representation of this relationship, achieving a held out R² of 0.896, compared to 0.877 for the random forest, while the neural network offered no further improvement. Poverty rate demonstrated the strongest association with purchasing power, but task groups provided additional insights beyond conventional economic indicators. Moreover, differences in task groups persisted across counties and over time. The subsequent sections examine the magnitude, geographical distribution, and stability of these relationships.
The association
The log specification of the panel regression was used for inference based on the diagnostic results outlined in Section 4. A one percentage point shift from non-routine cognitive to non-routine manual work was associated with approximately $846 lower purchasing power, with a 95% confidence interval of ±$38 at the panel mean (Figure 18). Routine cognitive and routine manual work also exhibited negative associations with purchasing power, although their estimated differences were smaller. This ordering aligns with the level specification reported in Section 4. For comparison, each additional percentage point of poverty was linked to approximately 2.9% lower purchasing power, a larger proportional association than any individual task group.
All three estimates of the purchasing power difference associated with a one percentage point shift from non-routine cognitive work toward another task group remained below zero, with non-routine manual work showing the largest negative association (Figure 18). The narrow intervals around the estimates indicated that the direction and ordering of these task group relationships were precisely estimated after accounting for the other variables in the model.
The level specification provided an additional perspective on the task group associations and economic controls in dollar terms. Among the controls, the poverty rate carried the largest and most precisely estimated coefficient, while the unemployment rate and log population were smaller but remained distinguishable from zero at the 0.05 level (Table 6). The table also reported the model’s overall F test, sample size, and R², offering a summary of the fit and precision of the complete level specification.
| Variable | Coefficient | Std. Error | t | p-value |
|---|---|---|---|---|
| Routine Cognitive | −$85,791 | 7,689 | −11.16 | 0.000 |
| Routine Manual | −$70,097 | 4,487 | −15.62 | 0.000 |
| Non-Routine Manual | −$104,139 | 5,526 | −18.84 | 0.000 |
| Poverty Rate | −$1,605 | 56 | −28.83 | 0.000 |
| Unemployment Rate | $206 | 97 | 2.13 | 0.033 |
| Log Population | $498 | 246 | 2.02 | 0.043 |
| F(20, 11962) = 757.5, p < 0.001; n = 11,983; R² = 0.848. Standard errors clustered by county. | ||||
Not poverty in disguise
A natural concern is that the task group associations simply reflect poverty, as counties with higher poverty rates also tend to have more routine work. Mean poverty and unemployment rates across the study period were both highest across much of the rural South, the Mississippi Delta, and parts of the Southwest border region (Figure 19), which overlapped with areas of low purchasing power in Figure 24. This overlap was important because each additional percentage point of poverty was associated with approximately 2.9% lower purchasing power in the log specification. Unemployment also showed a concentration along the California coast and Central Valley that was less apparent for poverty, indicating that the two measures captured different dimensions of local economic conditions.
A test of whether economic conditions accounted for the associations between task groups compared two log panel specifications estimated on the same observations (Figure 20). The first specification included task groups and year indicators, while the second added poverty, unemployment, and log population. After adding the controls, the estimates for the non-routine manual and routine manual became smaller, indicating that some of their initial associations overlapped with these economic conditions. However, routine cognitive remained relatively stable, staying close to −1.1% in both specifications. Notably, all three task group estimates remained negative and accurately estimated after adjustment. This comparison demonstrated that poverty, unemployment, and population explained part of the relationship, but they did not fully account for the association between task groups and purchasing power.
What carries the signal
Two complementary tests were conducted to assess the contribution of information from the task groups. The ablation test removed an entire feature block and refit the random forest (Figure 21). Removing poverty and unemployment resulted in a 0.178 reduction in held out R², compared to a smaller decline of 0.042 when the task groups were removed. This indicates that the economic controls provided more explanatory information overall, but the task groups still improved performance beyond what poverty and unemployment alone could capture.
A second measure left the fitted model unchanged and perturbed each feature or feature block through permutation (Figure 22). Jointly permuting the task groups led to a 0.153 reduction in held out R², compared to 0.128 for the year indicators and 0.008 for the state indicators. Poverty exhibited the largest individual decline. These results suggest that the fitted random forest heavily relied on the task groups even after incorporating economic controls, time, and geography.
The ablation and permutation results provided further insights into the model’s performance. Removing the task groups and refitting the model resulted in a modest reduction of 0.042 in the held-out R² value, suggesting that the model could still recover some overlapping information from poverty, unemployment, and year. In contrast, permuting the task groups after fitting caused a decline of 0.153, indicating that the model could no longer compensate for the loss of information. These two tests together revealed that the task groups contained information that overlapped with other variables but was not fully replaced by them.
The random forest traced the relationship across the observed range of each non-routine cognitive reference comparison (Figure 23). Purchasing power gradually declined across the three modeled task groups, without any distinct thresholds or reversals. Despite capturing these flexible relationships, the random forest achieved a held out R² of 0.877, compared to 0.896 for the log linear model, a difference of 0.019. This pattern aligned with the proportional relationship identified by the log linear specification and explained why the additional flexibility of the random forest did not enhance held out performance. Since the four task groups collectively accounted for the entire range, these curves represented the model’s response to changes in one group while keeping the remaining values constant. Therefore, these curves should not be interpreted as independent changes in employment.
Where purchasing power is strained
The association also exhibited a distinct geographic pattern. Purchasing power was highest along the metropolitan Northeast corridor, across parts of the upper Midwest, and in certain regions of the mountain West (Figure 24). Conversely, the lowest values were concentrated across the rural South and the southern border region. The map included all counties with a purchasing power value, rather than only those in the analytical panel, because the outcome required only household income and RPP.
This geographic pattern was also reflected in county task groups. Counties were categorized based on the task group that accounted for the largest average proportion of employment over the study period. In 664 counties, non-routine cognitive work was the dominant type of work, with a median purchasing power of $61,162. In contrast, 97 non-routine manual counties had a median purchasing power of $46,657, while 87 routine manual counties had a median purchasing power of $50,368. Routine cognitive work was the dominant group in 0 counties. These differences demonstrated that the association observed across the continuous task group measures was also evident when counties were grouped by the type of work that constituted the largest portion of local employment.
The relationship is proportional
Across the five grouped cross-validation folds, the log linear model and random forest showed similar variation across folds (Figure 25). The single held-out split also demonstrated the same pattern. The log linear model achieved an R² of 0.896, while the random forest achieved an R² of 0.877. The comparison extended to the neural network, which achieved a held-out R² of 0.869 and a mean absolute error (MAE) of $4,145 (Table 7). The neural network did not provide any improvement over the simpler models. Overall, these results indicated that additional model flexibility did not capture enough new structure to improve performance beyond the log linear specification.
| Model | R² | Mean Absolute Error |
|---|---|---|
| Ordinary Least Squares | 0.896 | $3,636 |
| Random Forest | 0.877 | $3,916 |
| Neural Network | 0.865 | $4,169 |
| OLS and the random forest are fit on the log target and scored on the dollar scale after back transformation. | ||
| Neural network results are from this single split only; the five fold grouped cross validation table in the analysis section reports cross validated results for the other two models. | ||
A complementary perspective used purchasing power dollars (Figure 26). Across ten equal-sized bins of routine cognitive tasks, observed purchasing power increased in the lower range, peaked near 20%, and then declined. The random forest closely reproduced this pattern, rather than identifying a substantially different relationship. This observation did not contradict the smoother declines in Figure 23, as the partial dependence curves varied one task group while holding the others constant, whereas the binned comparison reflected how the task groups occurred together in actual counties. Since the four groups summed to one, changes in routine cognitive work in the observed data were accompanied by changes in the other groups. Therefore, the agreement between the flexible model and the simpler log linear specification supported retaining the proportional specification rather than adding further model complexity.
The differences are durable
The differences observed across counties were not confined to a specific time period. The average proportion of employment in each task group remained relatively stable over the fifteen-year panel (Figure 27). The largest discontinuity occurred between 2009 and 2010, coinciding with the Census occupation coding change mentioned in Section 3. Outside this period, the national distribution of the four task groups remained comparatively stable.
County rankings provided a more robust test of whether the same places maintained distinct characteristics over time. Table 8 compared each county’s task group ranking in 2010 with its ranking in 2023 using Spearman correlations. Routine manual and non-routine cognitive work exhibited the greatest stability, indicating that counties with relatively high values in 2010 generally maintained high values in 2023. Non-routine manual work showed more movement, while routine cognitive work experienced the largest changes. This ordering aligns with the variance decomposition in Section 4, which demonstrated that routine cognitive work exhibited more within-county variation compared to the other task groups.
| Task Group | Rank Correlation |
|---|---|
| Routine Cognitive | 0.41 |
| Routine Manual | 0.87 |
| Non-Routine Cognitive | 0.88 |
| Non-Routine Manual | 0.65 |
| Computed from 2010 onward to avoid the Census occupation coding change. | |
The stability is visible county by county, and in the magnitude of change as well as the ordering, over the same two years used in Table 8 (Figure 28). Routine manual and non-routine manual, the two groups with almost no net movement in Figure 27, are centered near zero and run in both directions across counties. That mixed direction is consistent with variation sitting between counties rather than within them. Routine cognitive and non-routine cognitive, the pair whose national levels moved most, show a change that is both larger and far more uniform in direction, which matches the within county movement identified in the variance decomposition of Section 4. Routine cognitive rose in np.int64(2)% of counties and non-routine cognitive in np.int64(96)%, against np.int64(56)% for routine manual and np.int64(34)% for non-routine manual.
These movements followed the same cognitive to manual and routine to non-routine dimensions introduced in Figure 1 (Figure 29). Of the 3,209 counties observed in both periods, 87 percent shifted towards non-routine cognitive work, while the remainder moved away from it towards the three task groups associated with lower purchasing power. The wider coverage was worth the change of source because the counties the annual estimates omit are the small and rural ones. They moved in the same direction as the rest, though far less uniformly: 83 percent of them shifted towards non-routine cognitive work, against 99 percent of the counties the panel already covered. Because counties in the eastern half of the map are small enough that one arrow each would overlap into a smudge, neighboring counties were pooled onto a 65 kilometer grid and drawn as 1,457 averaged arrows, colored by how their counties divided rather than by the average alone. Counties moving away from non-routine cognitive work were scattered rather than concentrated, appearing somewhere in 22 percent of the grid cells. The arrow figure draws a direct connection between the observed county changes and the conceptual framework introduced earlier in the study.
Finally, the estimated associations themselves remained consistent over time. Separate specifications for 2010 to 2019 and 2021 to 2023 ran alongside the full panel (Figure 30). The ordering of the three task group coefficients remained unchanged across all three periods, and their magnitudes shifted only slightly. Reporting the coefficients as percentage changes ensured comparability between the periods, even though nominal purchasing power increased later in the panel. The persistence of both the county patterns and the coefficient ordering suggested that the association was not confined to a single recession, recovery, or pandemic period.
The results revealed that county task groups continued to be linked to purchasing power even after accounting for factors such as poverty, unemployment, population, and the year. The largest difference was observed between non-routine manual work and other tasks. A one percentage point shift from non-routine cognitive work to non-routine manual work was associated with a reduction in purchasing power of approximately $846 at the panel mean, with a 95% confidence interval of ±$38. Routine cognitive and routine manual work were also negatively associated with purchasing power compared to non-routine cognitive work. Among the variables examined, poverty showed the strongest association with purchasing power, with each additional percentage point associated with a reduction of approximately 2.9%. The random forest analysis indicated that task groups provided additional information beyond the economic controls. Removing the task groups and refitting the model reduced the held out R² by 0.042, while removing poverty and unemployment reduced it by 0.178. Jointly permuting the task groups further reduced the R² value by 0.153, compared to 0.128 for year indicators and 0.008 for state indicators. The geographic analysis showed that counties dominated by non-routine cognitive work had higher purchasing power, while counties dominated by the two manual groups had lower purchasing power. The stability analysis demonstrated that these county differences persisted over time. Routine manual and non-routine cognitive work maintained the most consistent county ordering between 2010 and 2023, while routine cognitive work exhibited the greatest movement. The ordering of the task group coefficients remained unchanged in the full panel, 2010 to 2019, and 2021 to 2023.
The level model, log linear model, random forest, and neural network were compared to find the most suitable representation of the relationship. The level specification produced negative task group associations but showed systematic curvature in the residuals, suggesting that a constant dollar difference across the purchasing power distribution was inadequate. Logging purchasing power reduced residual skew from 1.01 to 0.09, resulting in an in-sample R² of 0.913 and a held-out R² of 0.896. The random forest captured the nonlinear structure visible in the level model but did not improve held-out performance. Its log target specification achieved a held-out R² of 0.877, 0.019 below the log linear model, while the raw target random forest achieved a training R² of 0.989 but only 0.876 on held-out counties. The neural network added further flexibility, but regularization did not improve on either of the other models. Five-fold grouped cross-validation yielded similar general comparisons, with the difference between the log linear and random forest models remaining smaller than the variation across folds. Across these four specifications, the log linear model preserved the task group associations, resolved most of the diagnostic problems in the level model, and scored highest on held-out counties. The implications and limitations of these results are discussed in Section 6.
Conclusions
Summary of findings
Across all counties analyzed over a span of fifteen years, county task groups remained linked to purchasing power after accounting for poverty, unemployment, population, and year. The largest difference was observed for non-routine manual work. A one percentage point shift from non-routine cognitive to non-routine manual work was associated with approximately $846 lower purchasing power on average, with a 95% confidence interval of ±$38. Routine cognitive and routine manual work were also negatively associated with purchasing power relative to non-routine cognitive work. Among the variables examined, poverty demonstrated the strongest association with purchasing power, with each additional percentage point associated with approximately 2.9% lower purchasing power. The task groups still carried information beyond these conventional economic measures. Removing them and refitting the random forest reduced the held out R² by 0.042, while jointly permuting them reduced it by 0.153. In comparison, permuting the year indicators reduced R² by 0.128, and permuting the state indicators reduced it by only 0.008.
The sequence of models showed which representation of the relationship the data supported. The level specification identified negative associations but revealed systematic curvature in the residuals, suggesting that a constant dollar relationship was inadequate for the data. Logging purchasing power reduced the residual skew from 1.01 to 0.09, resulting in an in-sample R² of 0.913 and a held-out R² of 0.896. The random forest analysis explored whether nonlinear relationships and interactions improved upon this specification, but its log target version achieved a held-out R² of 0.877. The raw target random forest also demonstrated a closer fit to the training data compared to unseen counties, with R² values of 0.989 and 0.876, respectively. The neural network introduced even greater flexibility but failed to yield further improvements after regularization. Five-fold grouped cross-validation yielded similar general comparisons, with differences between the log linear model and random forest being smaller than the variation across folds. Geographic and stability analyses extended these results beyond model performance. County task group differences persisted from 2010 to 2023, and the ordering of the task group coefficients remained consistent across the entire panel, 2010 to 2019, and 2021 to 2023.
The study also clarified the scope of its findings regarding automation and affordability. The analysis did not measure automation adoption, job displacement, or the likelihood of specific counties losing jobs to new technologies. Instead, it examined the types of tasks already present in county employment and assessed their association with purchasing power. Counties with higher concentrations of routine and manual work, which are often emphasized in the automation literature, tended to have lower purchasing power compared to counties with more non-routine cognitive work. This relationship persisted after adjustment and remained consistent across geography and time. However, it did not establish a causal link between automation and purchasing power differences. Therefore, the study indirectly addressed automation by identifying regions where task structures commonly associated with higher automation potential coincided with lower purchasing power. The consequences of actual technological displacement remain for future research.
Contributions
The first contribution shifted the outcome from nominal income to purchasing power. Median household income was adjusted using BEA RPPs, directly incorporating local price differences into the measure of household resources. This adjustment distinguished counties with similar nominal incomes where the same income supported varying levels of purchasing power. Instead of comparing income alone, the study examined what that income could purchase after accounting for local price variations.
The second contribution measured task groups at the county level and tracked them over time. The final panel comprised 840 counties from 2008 to 2023, excluding 2020, and represented approximately 84% of the U.S. population. Using counties provided a finer geographic scale than broader labor market areas and directly connected task groups with local purchasing power, poverty, and unemployment. The panel structure demonstrated that these differences were not confined to a single point in time. County task group patterns remained relatively consistent throughout the study period, enabling the analysis to distinguish enduring differences between counties from short-term changes within them.
The county-level analysis also showed that task groups contained information beyond conventional measures of local economic conditions. Removing the task groups and refitting the random forest reduced the held-out R² by 0.042, while jointly permuting them reduced it by 0.153. Poverty exhibited a stronger association with purchasing power than any individual task group, but poverty and unemployment did not fully capture the information contained within the task groups. Counties that appeared similar on these conventional measures could still differ in the types of work residents performed and the purchasing power associated with those differences.
The third contribution determined the most appropriate representation of the relationship between task groups and purchasing power. The level specification revealed systematic residual curvature, while logging purchasing power reduced residual skew from 1.01 to 0.09 and achieved a held-out R² of 0.896. The random forest model achieved 0.877 on the same held-out counties, and a three-hidden-layer neural network provided no further improvement. Therefore, greater model flexibility did not enhance held-out performance once purchasing power was represented proportionally. The contribution lay in applying those models to test whether the data supported a more complex representation of the relationship.
These contributions resulted in a practical framework for comparing local economic conditions. Purchasing power indicated the extent to which household income extended after accounting for local prices, while task groups described the types of work comprising county employment. Consequently, combining the two revealed differences that income, poverty, or unemployment alone could not fully capture. Counties with similar values on conventional economic indicators could still differ in both their task groups and purchasing power. The framework provided policymakers and economic development organizations with a reproducible method to examine local economic conditions using publicly available data.
Limitations
Five limitations defined the scope of the results. First, the findings were correlational rather than causal. Counties were not randomly assigned to task groups, and the analysis could not isolate task groups from all other characteristics that varied across places. Unobserved factors related to both county employment and purchasing power could therefore contribute to the estimated relationship. The results showed that task groups remained associated with purchasing power after accounting for the measured controls, but they did not establish that differences in task groups caused differences in purchasing power.
Second, the analytical panel was limited to counties above the population threshold required for the occupational data. The final sample included 840 counties, representing approximately 84% of the U.S. population, but smaller and more rural counties were underrepresented. The estimates described the more populous counties included in the panel rather than all U.S. counties. Relationships in counties below the threshold could differ from those observed in the analytical sample.
Third, the geographic precision of RPPs varied across counties. A majority of panel observations relied on a state-level price parity rather than a more local measure, which made purchasing power less geographically precise for those observations. Some within-state differences in local prices were not captured by the outcome. The analysis applied these measures consistently but did not separately test whether the association differed according to the geographic precision of the price measure.
Fourth, purchasing power was measured in nominal dollars, which were not adjusted for changes in the national price level over time. RPPs corrected for price differences within a given year but did not place dollars from different years on a common scale. The year indicators absorbed any shared national price change across all counties. The task group and control coefficients were unaffected. The year coefficients themselves combined inflation with other national conditions that fluctuated between years. They could not be read as a measure of inflation alone. Consequently, dollar comparisons across years represented nominal amounts rather than constant purchasing power.
Finally, the primary panel regression did not include state indicators. Unmeasured state-level differences could have contributed to the estimated associations. The predictive models provided a partial check by including state indicators alongside task groups, poverty, unemployment, and year. Permuting the state indicators reduced the held-out R² by only 0.008, indicating that they contributed relatively little predictive information in that specification. However, the predictive models and panel regression were not identical, and this result did not eliminate the possibility of state-level confounding in the primary estimates.
Ethical considerations
The study required clear boundaries on how the results were interpreted. Since the county was the unit of analysis, the findings described places rather than individual workers. Applying county-level relationships to individuals would constitute an ecological fallacy. Consequently, the results could not be used to infer a worker’s purchasing power, economic vulnerability, or likelihood of being affected by automation based on their task group.
Representation also necessitated caution. Counties were excluded because occupational data were unavailable below the source publication threshold. The publication rule set the sample, and the analysis inherited it. As a result, smaller and more rural counties were underrepresented in the panel. Extending the findings to those counties would require extrapolation beyond the observed data, which was particularly important when comparing geographic patterns or identifying places that might warrant further attention.
The data also carried a selection bias that the analysis could not eliminate. The publication threshold excluded smaller counties before constructing the analytical sample, resulting in a panel that was not a random representation of all U.S. counties. Consequently, rural and less populous counties were underrepresented, making the results more accurate for the included counties than for those excluded by the threshold. This limitation matters because smaller rural counties are often central to discussions about automation and local economic vulnerability.
The associational nature of the findings also limited their use for policy decisions. Treating the estimated relationships as causal could lead to the allocation of resources to change a county’s task groups even if an unmeasured economic or geographic factor contributed to the observed difference. While the results supported identifying where lower purchasing power and specific task groups coexisted, they did not establish which intervention would alter those conditions. Similarly, the language used to describe counties required similar care. Labeling places based on presumed automation risk or economic vulnerability could stigmatize communities or deter investment based on relationships that the study did not establish causally.
Finally, the study relied entirely on publicly available aggregate data. No individual-level records or personally identifiable information entered the analytical dataset, and all reported results described county-level patterns rather than individual people.
Future directions
Three directions emerge from these limitations. The first is the coverage gap, which could potentially be bridged using satellite imagery. Jean et al. (2016)’s estimate of local economic conditions from daytime and nighttime imagery, where survey data are scarce, could be applied here. This approach could generate purchasing power estimates for counties below the ACS publication threshold and assess whether the geographic patterns observed in this study extend to areas outside the analytical panel. This would not expand the task group association itself, because comparable occupational estimates would still be unavailable for those counties. It would still show whether the spatial patterns reported here reflect the country more broadly rather than the more populous counties in the sample.
A second direction is to extend the study beyond 2008-2023 as additional occupational, income, price, poverty, and unemployment data become available. The current analysis revealed that the ordering of the task group associations remained stable across the entire panel and across 2010-2019 and 2021-2023. Extending the panel would directly test whether these relationships persist as local labor markets continue to evolve. Additionally, it would allow the results reported here to be evaluated out of time using observations that were not available when the original models were developed. Replicating the direction, magnitude, and proportional form of the associations in later years would provide stronger evidence that the patterns identified in this study were enduring rather than specific to the observed period.
The third approach is to investigate the underlying reasons for the correlation between task groups and purchasing power. Establishing a causal relationship would necessitate identifying a source of variation in task groups that was not influenced by local purchasing power. This could be achieved through factors such as plant openings and closures, major changes in production technology, or identifiable technology adoption shocks. Such an analysis would require a separate study with distinct data and a research design capable of isolating plausible exogenous changes in local task groups. Such an analysis would go beyond confirming the association observed in this study. It would show whether changes in the types of work performed within a county are accompanied by shifts in purchasing power.
Conclusion
This study discovered a consistent link between county task groups and purchasing power. Counties with higher concentrations of routine and manual work generally had lower purchasing power compared to those with more non-routine cognitive work, even after accounting for factors like poverty, unemployment, population, and year. The largest difference was observed for non-routine manual work. At the panel mean, a one percentage point shift from non-routine cognitive to non-routine manual work was associated with approximately $846 lower purchasing power, with a 95% confidence interval of ±$38.
The relationship persisted across both time and model specifications. County task group patterns remained relatively stable over the fifteen-year panel, and the ordering of the estimated associations remained consistent across the examined periods. Diagnostic testing revealed that the relationship was better represented proportionally than as a constant dollar difference. The log linear specification also achieved stronger held-out performance compared to the random forest, while the neural network provided no further improvement. These results suggested that additional model complexity was unnecessary to capture the primary relationship observed in the data.
The findings did not establish a causal relationship between task groups and purchasing power, nor did the study directly measure automation adoption or job displacement. Instead, the results indicated that the types of work concentrated within a county provided information about purchasing power beyond poverty, unemployment, population, and year. While poverty remained an important indicator of local economic conditions, it did not fully explain the association between task groups and purchasing power. Counties with higher concentrations of routine and manual work consistently exhibited lower purchasing power across geography and time. Therefore, understanding differences in purchasing power across counties required considering household income, local prices, and the types of work that characterized the local economy.
References
Software
Data processing, statistical modeling, and machine learning were conducted in Python using pandas, NumPy, SciPy, statsmodels, scikit-learn, TensorFlow, and GeoPandas (McKinney 2010; Harris et al. 2020; Virtanen et al. 2020; Seabold and Perktold 2010; Pedregosa et al. 2011; Abadi et al. 2016; Jordahl et al. 2020). Figures were produced in R using ggplot2, patchwork, scales, sf, ggrepel, and ggbeeswarm (R Core Team 2024; Wickham 2016; Pedersen 2024; Wickham, Pedersen, and Seidel 2023; Pebesma 2018; Slowikowski 2026; Clarke, Sherrill-Mix, and Dawson 2025).
Footnotes
The District of Columbia and Kalawao County, Hawaii, are excluded because they are absent from the county reference file. Connecticut counties leave the panel after 2021 because the state replaced its legacy counties with planning regions beginning in 2022.↩︎
All 104 missing observations came from Connecticut’s eight legacy counties, where unemployment rates were unavailable at the required geography (U.S. Bureau of Labor Statistics 2024). Because the missingness reflected geography rather than unemployment itself, it was treated as missing at random.↩︎