Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Partially Identified

Replication code for Partially Identified: privacy suppression and the automation–augmentation measure in the Anthropic Economic Index (Saporito).

The Anthropic Economic Index (AEI) reports, for each O*NET task, the share of observed AI usage that is automation rather than augmentation. That measure is now used in occupation-level wage regressions, an NBER working paper on task chaining, and regional exposure indices marketed for workforce-transition targeting.

This repository shows the measure is partially identified. AEI suppresses any cell holding fewer than 15 conversations from fewer than five accounts. That censors the released data on two margins — interaction modes within observed tasks, and whole tasks within occupations — and downstream users read the resulting absences as measured zeros.

Headline results

Six-digit SOC occupations in the AEI task file 775
With no observed task at all 225 (29.0%)
Usably identified (joint bound < 0.10, ≥ 5 observed tasks) 42 (5.4%)
Major groups containing zero usably-identified occupations 16 of 22
Identified set spanning [0.05, 0.95] 338 (43.6%)
Median joint bound width, all 775 0.883
Share of classified conversation volume suppressed 6.38%

The identified set is narrow only for text-mediated cognitive work. In Farming, Construction, Installation and Repair, Production, and Transportation — 272 occupations, 35% of the universe — the median occupation has zero observed tasks.

Identification is a property of data density, not of the work: log classified conversation volume alone explains 76% of the variation in joint bound width, and 21 occupational dummies add 1.2 percentage points on top of volume and coverage.

Reproducing

Requires R ≥ 4.1. Both scripts are self-contained, install what they need, and fetch the AEI files at runtime from Hugging Face. No manual downloads.

source("R/03-suppression-bounds.R")   # main analysis, ~2 min
source("R/04-mode-decomposition.R")   # section 7 mechanism check, ~1 min

Run 03 first if you want the full set of saved objects; 04 is standalone and can be run on its own.

Both scripts end with a drift table comparing 12–15 metrics against a reference run. A clean run prints All metrics match the reference run. If anything is flagged, check the join rate first — that is the failure mode most likely to cascade.

Tested on Posit Cloud and local RStudio, R 4.6.1, aieconindex 0.2.0.

What's here

R/03-suppression-bounds.R    Main analysis: mode coverage, intensive-margin
                             bounds, occupation crosswalk, extensive-margin and
                             joint bounds, mode composition, drift check
R/04-mode-decomposition.R    Decomposes suppressed volume by interaction mode
                             to test the mechanism claimed in section 7
paper/                       LaTeX source
data/                        CSV outputs — these are the paper's tables
figures/                     Figures 1 and 2

.rds intermediates are gitignored; they regenerate in about two minutes. The CSVs are committed so the paper's numbers can be checked without running anything.

Method in brief

Intensive margin. A task reporting k of five interaction modes has (5 − k) suppressed cells, each holding somewhere in [0, 14] conversations. Assigning all censored mass to automation gives an upper bound and all to augmentation a lower bound. This is an identified set, not a confidence interval: the endpoints are known given the threshold, so the uncertainty is deterministic rather than sampling variability.

Extensive margin. A task enters the released data only if it clears the same floor, so an occupation's unobserved tasks are not zero-usage either. Joint bounds censor both margins at once and are defined for all 775 SOC codes, including the 225 whose identified set is exactly [0, 1].

Identifying assumption. The filter has two criteria and both bounds treat the conversation criterion as binding, so a suppressed cell holds at most 14 conversations. If the account criterion binds anywhere — high volume, few accounts — the true hidden mass is larger and these bounds are too narrow. They are conservative in the direction that matters.

Scope

One release (release_2025_09_15), one week (4–11 August 2025), Claude.ai only, global aggregate. Not replicated across the January, March, or June 2026 releases. Bounds are unweighted across occupations; no employment weighting.

Data and attribution

Anthropic Economic Index, released under CC-BY 4.0 at huggingface.co/datasets/Anthropic/EconomicIndex. Task statements are O*NET data from the U.S. Department of Labor.

Handa, K., Tamkin, A., McCain, M., et al. (2025). Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations. arXiv:2503.04761.

This work is not affiliated with or endorsed by Anthropic. The privacy filter documented here is appropriate disclosure control, clearly described in Anthropic's own materials; the paper is a user's guide to what the released data can support, not a criticism of the design.

Licence

Code MIT (see LICENSE). Paper text CC-BY 4.0. AEI data CC-BY 4.0, © Anthropic.

About

Privacy suppression makes the Anthropic Economic Index automation–augmentation measure partially identified. Only 42 of 775 occupations are usably identified.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages