Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,12 @@ Small area estimation bridges the gap between nationally representative surveys
for reliable estimates at district, county, or municipality level. This tool:

1. Lets you describe your data (variable types, what auxiliary data you have, Stata version).
2. Recommends the most appropriate SAE method from a catalogue of 16 methods.
2. Recommends the most appropriate SAE method from a catalogue of 17 methods.
3. Generates a complete, commented R script and Stata `.do` file — ready to run with your
real variable names filled in.
4. Supports auxiliary variables that come from a sample (an agricultural census or other large
survey) rather than a full census or register, via the measurement-error Fay–Herriot model
(Ybarra & Lohr, 2008), which corrects for the sampling error those auxiliaries carry.

Target users: statisticians in national statistical offices and development organisations,
including those working in countries where Stata 14 is the installed standard.
Expand Down
128 changes: 128 additions & 0 deletions docs/SAE-CATALOGUE.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,8 @@ Use these tokens consistently across all R and Stata templates:
| `{{AUX_VARS_R}}` | Auxiliary variables as R formula terms, e.g. `x1 + x2` |
| `{{AUX_VARS_R_VEC}}` | Auxiliary variable names as a quoted R character vector, e.g. `"x1", "x2"` — for column-selection contexts such as `c({{AUX_VARS_R_VEC}})` |
| `{{AUX_VARS_STATA}}` | Auxiliary variables space-separated, e.g. `x1 x2` |
| `{{AUX_VAR_VARIANCES_R}}` | Auxiliary sampling-variance column names as a quoted R character vector, e.g. `"var_x1", "var_x2"` |
| `{{CI_ARRAY_BUILDER_R}}` | Generated R code assembling the per-domain measurement-error variance–covariance array `Ci` from the variance columns |
| `{{WEIGHT_VAR}}` | Sampling weight variable name |
| `{{DIRECT_EST_VAR}}` | Pre-computed direct estimate column |
| `{{DIRECT_VAR_VAR}}` | Sampling variance of the direct estimate |
Expand Down Expand Up @@ -1311,3 +1313,129 @@ Rao, J.N.K. & Molina, I. (2015). *Small Area Estimation*, 2nd ed. Wiley.
Sinha, S.K. & Rao, J.N.K. (2009). Robust small area estimation. *Canadian Journal of Statistics* 37, 381–399.

Asian Development Bank (2020). *Introduction to Small Area Estimation Techniques: A Practical Guide for National Statistical Offices.* https://www.adb.org/publications/introduction-small-area-estimation-techniques

---

## §7 Survey-based auxiliary variables (added v1.1.0)

Some users have auxiliary variables that come from a large *sample* (such as an agricultural
census run as a sample) rather than a full census or register. Those auxiliaries carry sampling
error. The standard Fay–Herriot model assumes auxiliaries are known exactly; using it here biases
the model, understates uncertainty, and can perform worse than the direct estimate. This section
adds the measurement-error Fay–Herriot model (Ybarra & Lohr, 2008).

When the user declares sample-based auxiliaries, the recommender ranks `fh-me` first for
area-level continuous or proportion targets and attaches a prominent caveat to every standard
known-covariate area-level method (`fh-eblup`, `spatial-fh`, `robust-fh`). If the sampling
variances of the auxiliary estimates are not available, `fh-me` carries a blocking caveat because
the correction cannot be applied without them.

### 7.01 Fay–Herriot with Measurement Error (Ybarra–Lohr)

```
id: fh-me
displayName: Fay–Herriot with Measurement Error (Area-Level, Sample-Based Auxiliaries)
level: area
inferenceType: frequentist
targetTypes: [continuous, proportion]
requiredInputs:
microdata: false
areaAggregates: true
censusAuxiliaries: area
weights: false
contiguityMatrix: false
coordinates: false
requiresAuxiliaryVariances: true
spatial: false
robust: false
mseMethod: jackknife
rPackage: emdi (or saeME)
rFunction: fh(method = "me") / saeME::eblupME
stataPackage: base
stataCommand: (no standard Stata command; use R)
stataMinVersion: 14
stataV14Fallback: |
* There is no standard Stata command for the measurement-error Fay-Herriot model.
* Run the R script for this method (emdi or saeME package).
* If you must stay in Stata, the standard fhsae command ignores auxiliary
* sampling error and may bias the estimates — interpret with caution:
* fhsae {{DIRECT_EST_VAR}} {{AUX_VARS_STATA}}, vardir({{DIRECT_VAR_VAR}}) method(reml)
plainDescription: |
An extension of the Fay–Herriot model for when your auxiliary variables come from a
sample (for example, an agricultural census run as a large sample) rather than a full
census, so the auxiliaries carry their own sampling error. The model adjusts for that
error, leaning more on the direct survey estimate in areas where the auxiliaries are
noisier. It needs the sampling variance of each auxiliary estimate for each area.
whyChooseThis: |
Choose this when your area-level auxiliary totals or means are themselves survey
estimates rather than known population values. Using a standard Fay–Herriot model in
that situation can bias the results and overstate their precision, and can even do worse
than the direct estimate. This method is designed to avoid that.
assumptions:
- Sampling variances of the auxiliary estimates are available for each area.
- The auxiliary measurement error is of the classical type (observed = true + noise).
- All other Fay–Herriot assumptions hold (known direct-estimate variances; correct
linking model; enough areas for variance estimation).
references:
- Ybarra, L.M.R. & Lohr, S.L. (2008). Biometrika 95(4), 919–931.
https://doi.org/10.1093/biomet/asn048
- Harmening, S. et al. (2023). The R Journal 15(1), RJ-2023-039.
https://journal.r-project.org/articles/RJ-2023-039/
- saeME package: https://cran.r-project.org/package=saeME
caveats:
- Requires the sampling variances of the auxiliary estimates; without them the
correction cannot be applied.
- MSE is estimated by jackknife only.
- No Stata equivalent; R is required.
```

**R template:**
```r
# ============================================================
# Fay–Herriot with Measurement Error (Ybarra–Lohr)
# For auxiliary variables that come from a sample (e.g. an agricultural
# census run as a large sample), and therefore carry sampling error.
# Generated by SAE Syntax Generator on {{DATE}}
# Reference: Ybarra & Lohr (2008); Harmening et al. (2023)
# R package: emdi
# Area-level data: {{AREA_DATA}}
# ============================================================

if (!requireNamespace("emdi", quietly = TRUE)) install.packages("emdi")
library(emdi)

area_data <- read.csv("{{AREA_DATA}}")
# Required columns:
# {{DIRECT_EST_VAR}} direct estimate of the target, per area
# {{DIRECT_VAR_VAR}} sampling variance of the direct estimate, per area
# {{AUX_VARS_R}} auxiliary estimates (from the large-sample census)
# {{AUX_VAR_VARIANCES_R}} sampling variance of each auxiliary estimate, per area
# {{AREA_ID}} area identifier

# Build the per-domain measurement-error variance–covariance array Ci.
# emdi expects an array of (p+1) x (p+1) x m, where p = number of auxiliaries,
# the leading row/column is the intercept (zero variance), and off-diagonal
# covariances between auxiliaries are assumed zero unless supplied.
{{CI_ARRAY_BUILDER_R}}

# Fit the measurement-error Fay–Herriot model
fh_me <- fh(
fixed = {{DIRECT_EST_VAR}} ~ {{AUX_VARS_R}},
vardir = "{{DIRECT_VAR_VAR}}",
combined_data = area_data,
domains = "{{AREA_ID}}",
method = "me",
Ci = Ci,
MSE = TRUE,
mse_type = "jackknife"
)

summary(fh_me)
estimators(fh_me, MSE = TRUE, CV = TRUE)

# NOTE: the modified shrinkage factor leans more on the direct estimate in
# areas where the auxiliary variances are large. Compare against the direct
# estimate (the benchmark) to confirm the model is helping rather than harming.
```

**Stata template:** use the `stataV14Fallback` note above (R-only method).
5 changes: 4 additions & 1 deletion docs/adding-a-method.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,8 @@ Open `src/catalogue/my-new-method.ts` and edit each field. The full schema is de
|-------|------|-------------|
| `spatial` | `boolean` | Does the method explicitly model spatial dependence? |
| `robust` | `boolean` | Is the method robust to outliers (M-estimation or similar)? |
| `mseMethod` | `'prasad-rao' \| 'bootstrap' \| 'both' \| 'posterior'` | How mean squared error is estimated |
| `requiresAuxiliaryVariances` | `boolean` (optional) | Set `true` for measurement-error methods that need the sampling variances of the auxiliary estimates (e.g. `fh-me`). Treated as `false` when absent. When `true`, the recommender only offers the method if the user has declared sample-based auxiliaries, and it expects the variance columns to be supplied. |
| `mseMethod` | `'prasad-rao' \| 'bootstrap' \| 'both' \| 'posterior' \| 'jackknife'` | How mean squared error is estimated. Use `'jackknife'` for the measurement-error Fay–Herriot model, whose MSE is estimated by jackknife only. |

### Software

Expand Down Expand Up @@ -112,6 +113,8 @@ be substituted from user input.
| `{{AUX_VARS_R}}` | Auxiliary variables as `var1 + var2 + ...` (R formula syntax) |
| `{{AUX_VARS_STATA}}` | Auxiliary variables as `var1 var2 ...` (space-separated) |
| `{{AUX_VARS_R_VEC}}` | Auxiliary variables as `c("var1", "var2", ...)` (R vector) |
| `{{AUX_VAR_VARIANCES_R}}` | Auxiliary sampling-variance column names as a quoted R character vector, e.g. `"var_x1", "var_x2"` (measurement-error methods) |
| `{{CI_ARRAY_BUILDER_R}}` | Generated R code that assembles the per-domain measurement-error variance–covariance array `Ci` from the variance columns (used by `fh-me`) |
| `{{SURVEY_DATA}}` | Path to the survey CSV file |
| `{{AREA_DATA}}` | Path to the area-level CSV file |
| `{{CENSUS_DATA}}` | Path to the census CSV file |
Expand Down
47 changes: 47 additions & 0 deletions e2e/wizard.spec.ts
Original file line number Diff line number Diff line change
Expand Up @@ -93,4 +93,51 @@ test.describe('Wizard happy path', () => {
const dl = await downloadPromise
expect(dl.suggestedFilename()).toContain('.R')
})

test('sample-based auxiliaries: FH-ME appears and standard FH shows the sampling-error caveat', async ({ page }) => {
// ── Step 1: Upload codebook ─────────────────────────────────────────────
const fileInput = page.locator('[data-testid="file-input"]')
const fixturePath = path.resolve(__dirname, 'fixtures/test-codebook.csv')
await fileInput.setInputFiles(fixturePath)
await page.getByRole('button', { name: /next/i }).click()

// ── Step 2: Assign roles ────────────────────────────────────────────────
await expect(page.getByText('Variable roles')).toBeVisible()
await setRole(page, 'income', 'Target')
await setRole(page, 'edu_rate', 'Auxiliary')
await setRole(page, 'urban_pct', 'Auxiliary')
await setRole(page, 'dir_est', 'Direct estimate')
await setRole(page, 'dir_var', 'Sampling variance')
await page.getByRole('button', { name: /next/i }).click()

// ── Step 3: Data availability ───────────────────────────────────────────
await expect(page.getByText('Data availability')).toBeVisible()

// Continuous target
await page.getByText('Continuous (e.g. income', { exact: false }).first().click()

// Area aggregates + census auxiliaries so FH-EBLUP and FH-ME are both eligible
await page.getByText('area-level direct estimates', { exact: false }).first().click()
await page.getByText('population-level auxiliary variables', { exact: false }).first().click()

// Declare sample-based auxiliaries, with variances available
await page.locator('[data-testid="aux-source-sample"]').click()
await page.locator('[data-testid="aux-has-variances"]').check()

await page.getByRole('button', { name: /next/i }).click()

// ── Step 4: Methods ─────────────────────────────────────────────────────
await expect(page.getByText('Recommended methods')).toBeVisible()

// FH-ME card is present and selectable
const fhMeRadio = page.locator('[data-testid="select-fh-me"]')
await expect(fhMeRadio).toBeVisible()

// Standard FH-EBLUP card shows the prominent sampling-error caveat
await expect(page.locator('[data-testid="sampling-error-note-fh-eblup"]')).toBeVisible()

// FH-ME is selectable
await fhMeRadio.click()
await expect(fhMeRadio).toBeChecked()
})
})
7 changes: 4 additions & 3 deletions src/catalogue/catalogue.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -19,11 +19,12 @@ const EXPECTED_IDS = [
'glmm-count',
'two-part-zinfl',
'hb-unit',
'fh-me',
]

describe('Catalogue index', () => {
it('exports exactly 16 entries', () => {
expect(catalogue).toHaveLength(16)
it('exports exactly 17 entries', () => {
expect(catalogue).toHaveLength(17)
})

it('contains all expected method IDs', () => {
Expand Down Expand Up @@ -128,7 +129,7 @@ describe('Catalogue schema validation', () => {
})

it('has valid mseMethod', () => {
expect(['prasad-rao', 'bootstrap', 'both', 'posterior']).toContain(
expect(['prasad-rao', 'bootstrap', 'both', 'posterior', 'jackknife']).toContain(
entry.mseMethod,
)
})
Expand Down
126 changes: 126 additions & 0 deletions src/catalogue/fh-me.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,126 @@
import type { CatalogueEntry } from '../types/index.js'

const entry: CatalogueEntry = {
id: 'fh-me',
displayName: 'Fay–Herriot with Measurement Error (Area-Level, Sample-Based Auxiliaries)',
level: 'area',
inferenceType: 'frequentist',
targetTypes: ['continuous', 'proportion'],
requiredInputs: {
microdata: false,
areaAggregates: true,
censusAuxiliaries: 'area',
weights: false,
contiguityMatrix: false,
coordinates: false,
},
requiresAuxiliaryVariances: true,
spatial: false,
robust: false,
mseMethod: 'jackknife',
rPackage: 'emdi (or saeME)',
rFunction: 'fh(method = "me") / saeME::eblupME',
stataPackage: 'base',
stataCommand: '(no standard Stata command; use R)',
stataMinVersion: 14,
stataV14Fallback:
'* There is no standard Stata command for the measurement-error Fay-Herriot model.\n' +
'* Run the R script for this method (emdi or saeME package).\n' +
'* If you must stay in Stata, the standard fhsae command ignores auxiliary\n' +
'* sampling error and may bias the estimates — interpret with caution:\n' +
'* fhsae {{DIRECT_EST_VAR}} {{AUX_VARS_STATA}}, vardir({{DIRECT_VAR_VAR}}) method(reml)',
plainDescription:
'An extension of the Fay–Herriot model for when your auxiliary variables come from a ' +
'sample (for example, an agricultural census run as a large sample) rather than a full ' +
'census, so the auxiliaries carry their own sampling error. The model adjusts for that ' +
'error, leaning more on the direct survey estimate in areas where the auxiliaries are ' +
'noisier. It needs the sampling variance of each auxiliary estimate for each area.',
whyChooseThis:
'Choose this when your area-level auxiliary totals or means are themselves survey ' +
'estimates rather than known population values. Using a standard Fay–Herriot model in ' +
'that situation can bias the results and overstate their precision, and can even do worse ' +
'than the direct estimate. This method is designed to avoid that.',
assumptions: [
'Sampling variances of the auxiliary estimates are available for each area.',
'The auxiliary measurement error is of the classical type (observed = true + noise).',
'All other Fay–Herriot assumptions hold (known direct-estimate variances; correct ' +
'linking model; enough areas for variance estimation).',
],
references: [
'Ybarra, L.M.R. & Lohr, S.L. (2008). Biometrika 95(4), 919–931. https://doi.org/10.1093/biomet/asn048',
'Harmening, S. et al. (2023). The R Journal 15(1), RJ-2023-039. https://journal.r-project.org/articles/RJ-2023-039/',
'saeME package: https://cran.r-project.org/package=saeME',
],
caveats: [
'Requires the sampling variances of the auxiliary estimates; without them the ' +
'correction cannot be applied.',
'MSE is estimated by jackknife only.',
'No Stata equivalent; R is required.',
],
rTemplate: `# ============================================================
# Fay–Herriot with Measurement Error (Ybarra–Lohr)
# For auxiliary variables that come from a sample (e.g. an agricultural
# census run as a large sample), and therefore carry sampling error.
# Generated by SAE Syntax Generator on {{DATE}}
# Reference: Ybarra & Lohr (2008); Harmening et al. (2023)
# R package: emdi
# Area-level data: {{AREA_DATA}}
# ============================================================

if (!requireNamespace("emdi", quietly = TRUE)) install.packages("emdi")
library(emdi)

area_data <- read.csv("{{AREA_DATA}}")
# Required columns:
# {{DIRECT_EST_VAR}} direct estimate of the target, per area
# {{DIRECT_VAR_VAR}} sampling variance of the direct estimate, per area
# {{AUX_VARS_R}} auxiliary estimates (from the large-sample census)
# {{AUX_VAR_VARIANCES_R}} sampling variance of each auxiliary estimate, per area
# {{AREA_ID}} area identifier

# Build the per-domain measurement-error variance–covariance array Ci.
# emdi expects an array of (p+1) x (p+1) x m, where p = number of auxiliaries,
# the leading row/column is the intercept (zero variance), and off-diagonal
# covariances between auxiliaries are assumed zero unless supplied.
{{CI_ARRAY_BUILDER_R}}

# Fit the measurement-error Fay–Herriot model
fh_me <- fh(
fixed = {{DIRECT_EST_VAR}} ~ {{AUX_VARS_R}},
vardir = "{{DIRECT_VAR_VAR}}",
combined_data = area_data,
domains = "{{AREA_ID}}",
method = "me",
Ci = Ci,
MSE = TRUE,
mse_type = "jackknife"
)

summary(fh_me)
estimators(fh_me, MSE = TRUE, CV = TRUE)

# NOTE: the modified shrinkage factor leans more on the direct estimate in
# areas where the auxiliary variances are large. Compare against the direct
# estimate (the benchmark) to confirm the model is helping rather than harming.
`,
stataTemplate: `* ============================================================
* Fay–Herriot with Measurement Error — No Standard Stata Equivalent
* Generated by SAE Syntax Generator on {{DATE}}
* Please use the R script (emdi or saeME package) for this method.
* Area-level data: {{AREA_DATA}}
* ============================================================

* There is no standard Stata command for the measurement-error Fay–Herriot model.
* The correction for sampling error in the auxiliary variables requires the R
* implementation (emdi::fh(method = "me") or saeME::eblupME).
*
* If you must stay in Stata, the standard fhsae command ignores auxiliary
* sampling error and may bias the estimates — interpret with caution:
*
* net install fhsae, from("https://raw.github.com/jpazvd/fhsae/master/")
* use "{{AREA_DATA}}", clear
* fhsae {{DIRECT_EST_VAR}} {{AUX_VARS_STATA}}, vardir({{DIRECT_VAR_VAR}}) method(reml)
`,
}

export default entry
2 changes: 2 additions & 0 deletions src/catalogue/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@ import glmmBinary from './glmm-binary.js'
import glmmCount from './glmm-count.js'
import twoPartZinfl from './two-part-zinfl.js'
import hbUnit from './hb-unit.js'
import fhMe from './fh-me.js'

export const catalogue: CatalogueEntry[] = [
direct,
Expand All @@ -33,6 +34,7 @@ export const catalogue: CatalogueEntry[] = [
glmmCount,
twoPartZinfl,
hbUnit,
fhMe,
]

export default catalogue
Loading
Loading