feat(bls): OEWS occupation wages + CES health-care employment — bls.oews / bls.ces, stack bls oews|ces, notebook 7f (refs #695)

OEWS: the BLS API serves only the latest OEWS year, so history comes
from the annual oesm{yy}nat.zip files (2013–2024), which bls.gov only
serves to a full browser header set; each zip's national_M{year}_dl
sheet (xls through 2013, xlsx after; OCC_GROUP through 2018, O_GROUP
plus area/industry columns from 2019) is normalised to one row per SOC
occupation with employment and the hourly/annual wage distribution,
suppressed cells (* # **) as NULL, and the BLS series-id prefix kept
per row. One bib Source per release (module:bls, table:bls.oews,
source:bls-website, year:N) and a cms.ingest_log row with the zip's
sha256.

CES: a closed registry of employment (01) and average hourly earnings
(03) series for health care, ambulatory care, offices of physicians,
home health, hospitals and nursing/residential care (CES65<industry>
<type>), pulled through the public API in 10-year × 25-series chunks
(20 × 50 with BLS_API_KEY), monthly rows with the series id on every
row, one bib Source per series.

stack bls oews [--year]... [--write [--offline]] [--occ]...;
stack bls ces [--write] [--series]... [--annual]. Notebook 7f overlays
clinical-staff wages (31-9092, 29-2061, 29-1141, 29-1171) and
physician-office / ambulatory employment on the family's exposure
dates. CLI docs regenerated.
This commit is contained in:
kert
2026-09-22 16:25:46 -04:00
parent 592ad3cd24
commit b1e06b1b9e
87 changed files with 1953 additions and 73 deletions

View File

@@ -1,6 +1,6 @@
---
title: stack api serve
sidebar_position: 58
sidebar_position: 61
---
# `stack api serve`

View File

@@ -1,6 +1,6 @@
---
title: stack api
sidebar_position: 57
sidebar_position: 60
---
# `stack api`

26
docs/docs/cli/bls-ces.md Normal file
View File

@@ -0,0 +1,26 @@
---
title: stack bls ces
sidebar_position: 35
---
# `stack bls ces`
```
Usage: stack bls ces [OPTIONS]
CES health-care employment and earnings (bls.ces, #695): monthly, seasonally
adjusted, for health care / ambulatory care / offices of physicians / home
health / hospitals / nursing care, via the BLS public API (BLS_API_KEY widens
the per-request limits), one bib Source per series.
╭─ Options ────────────────────────────────────────────────────────────────────╮
│ --series TEXT CES series id(s) to print; default all in the │
│ registry. │
│ --start INTEGER First year to pull with --write. [default: 2006] │
│ --end INTEGER Last year to pull with --write. [default: 2025] │
│ --annual Print annual means instead of months. │
│ --write Pull the registry's series, load bls.ces, cite, │
│ republish. │
│ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯
```

23
docs/docs/cli/bls-oews.md Normal file
View File

@@ -0,0 +1,23 @@
---
title: stack bls oews
sidebar_position: 34
---
# `stack bls oews`
```
Usage: stack bls oews [OPTIONS]
OEWS national occupation wages (bls.oews, #695): one May release per year from
the bls.gov oesm{yy}nat.zip files, every SOC occupation, employment and
hourly/annual wage distribution, cited in bib per release and logged in
cms.ingest_log. --occ prints the series.
╭─ Options ────────────────────────────────────────────────────────────────────╮
│ --year INTEGER Release year(s); default 2013–. │
│ --occ TEXT SOC code(s) to print, e.g. 31-9092. │
│ --write Download, load bls.oews, cite, republish. │
│ --offline With --write: use the cached zips. │
│ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯
```

27
docs/docs/cli/bls.md Normal file
View File

@@ -0,0 +1,27 @@
---
title: stack bls
sidebar_position: 33
---
# `stack bls`
```
Usage: stack bls [OPTIONS] COMMAND [ARGS]...
BLS occupation wages (OEWS) and health-care employment (CES).
╭─ Options ────────────────────────────────────────────────────────────────────╮
│ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯
╭─ Commands ───────────────────────────────────────────────────────────────────╮
│ oews OEWS national occupation wages (bls.oews, #695): one May release │
│ per year from the bls.gov oesm{yy}nat.zip files, every SOC occupation, │
│ employment and hourly/annual wage distribution, cited in bib per │
│ release and logged in cms.ingest_log. --occ prints the series. │
│ ces CES health-care employment and earnings (bls.ces, #695): monthly, │
│ seasonally adjusted, for health care / ambulatory care / offices of │
│ physicians / home health / hospitals / nursing care, via the BLS │
│ public API (BLS_API_KEY widens the per-request limits), one bib │
│ Source per series. │
╰──────────────────────────────────────────────────────────────────────────────╯
```

View File

@@ -1,6 +1,6 @@
---
title: stack comments dockets
sidebar_position: 37
sidebar_position: 40
---
# `stack comments dockets`

View File

@@ -1,6 +1,6 @@
---
title: stack comments extract-ocr
sidebar_position: 35
sidebar_position: 38
---
# `stack comments extract-ocr`

View File

@@ -1,6 +1,6 @@
---
title: stack comments extract
sidebar_position: 34
sidebar_position: 37
---
# `stack comments extract`

View File

@@ -1,6 +1,6 @@
---
title: stack comments seal
sidebar_position: 38
sidebar_position: 41
---
# `stack comments seal`

View File

@@ -1,6 +1,6 @@
---
title: stack comments stats
sidebar_position: 36
sidebar_position: 39
---
# `stack comments stats`

View File

@@ -1,6 +1,6 @@
---
title: stack comments unseal
sidebar_position: 39
sidebar_position: 42
---
# `stack comments unseal`

View File

@@ -1,6 +1,6 @@
---
title: stack comments
sidebar_position: 33
sidebar_position: 36
---
# `stack comments`

View File

@@ -1,6 +1,6 @@
---
title: stack db comment
sidebar_position: 51
sidebar_position: 54
---
# `stack db comment`

View File

@@ -1,6 +1,6 @@
---
title: stack db inspect
sidebar_position: 52
sidebar_position: 55
---
# `stack db inspect`

View File

@@ -1,6 +1,6 @@
---
title: stack db
sidebar_position: 50
sidebar_position: 53
---
# `stack db`

View File

@@ -1,6 +1,6 @@
---
title: stack docs build
sidebar_position: 54
sidebar_position: 57
---
# `stack docs build`

View File

@@ -1,6 +1,6 @@
---
title: stack docs generate
sidebar_position: 56
sidebar_position: 59
---
# `stack docs generate`

View File

@@ -1,6 +1,6 @@
---
title: stack docs serve
sidebar_position: 55
sidebar_position: 58
---
# `stack docs serve`

View File

@@ -1,6 +1,6 @@
---
title: stack docs
sidebar_position: 53
sidebar_position: 56
---
# `stack docs`

View File

@@ -1,6 +1,6 @@
---
title: stack lake deploy
sidebar_position: 41
sidebar_position: 44
---
# `stack lake deploy`

View File

@@ -1,6 +1,6 @@
---
title: stack lake load
sidebar_position: 42
sidebar_position: 45
---
# `stack lake load`

View File

@@ -1,6 +1,6 @@
---
title: stack lake parity
sidebar_position: 44
sidebar_position: 47
---
# `stack lake parity`

View File

@@ -1,6 +1,6 @@
---
title: stack lake validate
sidebar_position: 43
sidebar_position: 46
---
# `stack lake validate`

View File

@@ -1,6 +1,6 @@
---
title: stack lake
sidebar_position: 40
sidebar_position: 43
---
# `stack lake`

View File

@@ -1,6 +1,6 @@
---
title: stack llm hosts
sidebar_position: 48
sidebar_position: 51
---
# `stack llm hosts`

View File

@@ -1,6 +1,6 @@
---
title: stack llm index
sidebar_position: 46
sidebar_position: 49
---
# `stack llm index`

View File

@@ -1,6 +1,6 @@
---
title: stack llm restamp
sidebar_position: 47
sidebar_position: 50
---
# `stack llm restamp`

View File

@@ -1,6 +1,6 @@
---
title: stack llm serve
sidebar_position: 49
sidebar_position: 52
---
# `stack llm serve`

View File

@@ -1,6 +1,6 @@
---
title: stack llm
sidebar_position: 45
sidebar_position: 48
---
# `stack llm`

View File

@@ -1,6 +1,6 @@
---
title: stack mail attach-smarthost
sidebar_position: 99
sidebar_position: 102
---
# `stack mail attach-smarthost`

View File

@@ -1,6 +1,6 @@
---
title: stack mail dkim-export
sidebar_position: 98
sidebar_position: 101
---
# `stack mail dkim-export`

View File

@@ -1,6 +1,6 @@
---
title: stack mail dns
sidebar_position: 97
sidebar_position: 100
---
# `stack mail dns`

View File

@@ -1,6 +1,6 @@
---
title: stack mail down
sidebar_position: 95
sidebar_position: 98
---
# `stack mail down`

View File

@@ -1,6 +1,6 @@
---
title: stack mail provision
sidebar_position: 93
sidebar_position: 96
---
# `stack mail provision`

View File

@@ -1,6 +1,6 @@
---
title: stack mail rotate-creds
sidebar_position: 100
sidebar_position: 103
---
# `stack mail rotate-creds`

View File

@@ -1,6 +1,6 @@
---
title: stack mail seed-mailboxes
sidebar_position: 101
sidebar_position: 104
---
# `stack mail seed-mailboxes`

View File

@@ -1,6 +1,6 @@
---
title: stack mail status
sidebar_position: 96
sidebar_position: 99
---
# `stack mail status`

View File

@@ -1,6 +1,6 @@
---
title: stack mail up
sidebar_position: 94
sidebar_position: 97
---
# `stack mail up`

View File

@@ -1,6 +1,6 @@
---
title: stack mail wire-git
sidebar_position: 102
sidebar_position: 105
---
# `stack mail wire-git`

View File

@@ -1,6 +1,6 @@
---
title: stack mail
sidebar_position: 92
sidebar_position: 95
---
# `stack mail`

View File

@@ -1,6 +1,6 @@
---
title: stack perf show
sidebar_position: 60
sidebar_position: 63
---
# `stack perf show`

View File

@@ -1,6 +1,6 @@
---
title: stack perf
sidebar_position: 59
sidebar_position: 62
---
# `stack perf`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs cpt-ingest
sidebar_position: 74
sidebar_position: 77
---
# `stack pfs cpt-ingest`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs elements
sidebar_position: 66
sidebar_position: 69
---
# `stack pfs elements`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs exposure
sidebar_position: 71
sidebar_position: 74
---
# `stack pfs exposure`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs families
sidebar_position: 68
sidebar_position: 71
---
# `stack pfs families`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs guidance
sidebar_position: 69
sidebar_position: 72
---
# `stack pfs guidance`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs lineage
sidebar_position: 67
sidebar_position: 70
---
# `stack pfs lineage`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs reaction
sidebar_position: 70
sidebar_position: 73
---
# `stack pfs reaction`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs review
sidebar_position: 73
sidebar_position: 76
---
# `stack pfs review`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs utilization
sidebar_position: 72
sidebar_position: 75
---
# `stack pfs utilization`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs
sidebar_position: 65
sidebar_position: 68
---
# `stack pfs`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma eligible
sidebar_position: 86
sidebar_position: 89
---
# `stack prisma eligible`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma export
sidebar_position: 83
sidebar_position: 86
---
# `stack prisma export`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma extract
sidebar_position: 87
sidebar_position: 90
---
# `stack prisma extract`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma fetch
sidebar_position: 89
sidebar_position: 92
---
# `stack prisma fetch`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma flow
sidebar_position: 88
sidebar_position: 91
---
# `stack prisma flow`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma init
sidebar_position: 82
sidebar_position: 85
---
# `stack prisma init`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma ping-llm
sidebar_position: 84
sidebar_position: 87
---
# `stack prisma ping-llm`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma run
sidebar_position: 90
sidebar_position: 93
---
# `stack prisma run`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma screen
sidebar_position: 85
sidebar_position: 88
---
# `stack prisma screen`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma vpn
sidebar_position: 91
sidebar_position: 94
---
# `stack prisma vpn`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma
sidebar_position: 81
sidebar_position: 84
---
# `stack prisma`

View File

@@ -1,6 +1,6 @@
---
title: stack rec list
sidebar_position: 62
sidebar_position: 65
---
# `stack rec list`

View File

@@ -1,6 +1,6 @@
---
title: stack rec opps
sidebar_position: 64
sidebar_position: 67
---
# `stack rec opps`

View File

@@ -1,6 +1,6 @@
---
title: stack rec pfs
sidebar_position: 63
sidebar_position: 66
---
# `stack rec pfs`

View File

@@ -1,6 +1,6 @@
---
title: stack rec
sidebar_position: 61
sidebar_position: 64
---
# `stack rec`

View File

@@ -23,6 +23,7 @@ Usage: stack [OPTIONS] COMMAND [ARGS]...
│ load Ingest data (CCLF, BCDA, seeds). │
│ generate Regenerate code artefacts. │
│ bib Bibliography operations. │
│ bls BLS occupation wages (OEWS) and health-care employment (CES). │
│ comments CMS rulemaking comment text extraction & analysis. │
│ lake Lakehouse schema and data. │
│ llm Local RAG — index, search, tag. │

View File

@@ -1,6 +1,6 @@
---
title: stack zot dump-schema
sidebar_position: 76
sidebar_position: 79
---
# `stack zot dump-schema`

View File

@@ -1,6 +1,6 @@
---
title: stack zot fix-dates
sidebar_position: 77
sidebar_position: 80
---
# `stack zot fix-dates`

View File

@@ -1,6 +1,6 @@
---
title: stack zot fix-fields
sidebar_position: 79
sidebar_position: 82
---
# `stack zot fix-fields`

View File

@@ -1,6 +1,6 @@
---
title: stack zot fix-keys
sidebar_position: 78
sidebar_position: 81
---
# `stack zot fix-keys`

View File

@@ -1,6 +1,6 @@
---
title: stack zot verify-parity
sidebar_position: 80
sidebar_position: 83
---
# `stack zot verify-parity`

View File

@@ -1,6 +1,6 @@
---
title: stack zot
sidebar_position: 75
sidebar_position: 78
---
# `stack zot`

View File

@@ -1303,6 +1303,162 @@ def _(con, mo, pl):
return
@app.cell(hide_code=True)
def _(alt, code, con, mo, not_built, pl):
# ── 7f. Clinical-staff wages and employment ──
from bls.ces import read_series as _read_ces
from bls.oews import read_occupations as _read_oews
from pfs.codetables import read_exposures as _read_exposures_w
from pfs.families import family_of as _family_of_w
_fam = _family_of_w(code)
_key = _fam.key if _fam else code
_occs = {
"31-9092": "Medical assistants",
"29-2061": "Licensed practical / vocational nurses",
"29-1141": "Registered nurses",
"29-1171": "Nurse practitioners",
}
_wages, _ces = [], []
if con is not None:
try:
_wages = _read_oews(con, list(_occs))
except Exception: # noqa: BLE001 — not built yet
_wages = []
try:
_ces = _read_ces(
con, ["CES6562110001", "CES6562100001", "CES6562110003"], annual=True
)
except Exception: # noqa: BLE001
_ces = []
_intro = (
"## 7f. Clinical-staff wages and employment\n\n"
"Care-management codes pay for clinical-staff time — 99490 is twenty "
"minutes of clinical staff directed by the practitioner, APCM's `actor` "
"element is the same staff — so the second outcome series is what that "
"time costs and how many people are employed to supply it. Wages are the "
"BLS OEWS national estimates (`bls.oews`, one May release a year, each "
"cited as its own bib Source); employment is the BLS CES monthly series "
"for offices of physicians and ambulatory care (`bls.ces`, annual means "
"of the seasonally adjusted months, series ids kept on every row). Dashed "
"rules are the family's `becomes-payable` / `ends` dates from section 7c. "
"The SOC revision between the 2018 and 2019 releases can move an "
"occupation's boundary; the four occupations shown kept their codes."
)
if not _wages and not _ces:
_view = mo.md(
_intro
+ "\n\n"
+ not_built("stack bls oews --write` and `stack bls ces --write")
)
else:
try:
_exp = (
_read_exposures_w(con, family=_key)
if _fam
else _read_exposures_w(con, code=code)
)
except Exception: # noqa: BLE001
_exp = []
_rules = pl.DataFrame(
{
"year": [r.year for r in _exp if r.kind in ("becomes-payable", "ends")],
"exposure": [
f"{r.code} {r.kind}"
for r in _exp
if r.kind in ("becomes-payable", "ends")
],
}
)
_panels = [mo.md(_intro)]
if _wages:
_wdf = pl.DataFrame(
{
"occupation": [
f"{r.occ_code} {_occs.get(r.occ_code, r.occ_title)}"
for r in _wages
],
"year": [r.year for r in _wages],
"mean annual wage": [r.a_mean for r in _wages],
"median annual wage": [r.a_median for r in _wages],
"employment": [r.tot_emp for r in _wages],
}
)
_wchart = (
alt.Chart(_wdf.to_pandas())
.mark_line(point=True)
.encode(
x=alt.X("year:O", title="OEWS release year"),
y=alt.Y("mean annual wage:Q", title="Mean annual wage ($)"),
color=alt.Color("occupation:N", title="Occupation"),
tooltip=[
"occupation",
"year",
"mean annual wage",
"median annual wage",
"employment",
],
)
.properties(
width=640,
height=300,
title="Clinical-staff mean annual wage (OEWS)",
)
)
if len(_rules):
_wchart = _wchart + (
alt.Chart(_rules.to_pandas())
.mark_rule(strokeDash=[4, 4], color="gray")
.encode(x="year:O", tooltip=["exposure"])
)
_panels += [
mo.ui.altair_chart(_wchart),
mo.ui.table(_wdf, label="bls.oews — clinical staff"),
]
else:
_panels.append(mo.md(not_built("stack bls oews --write")))
if _ces:
_cdf = pl.DataFrame(
{
"series": [
f"{r.series_id} {r.industry_name} {r.data_type}" for r in _ces
],
"year": [r.year for r in _ces],
"annual mean": [r.value for r in _ces],
}
)
_emp = _cdf.filter(pl.col("series").str.contains("employment"))
_cchart = (
alt.Chart(_emp.to_pandas())
.mark_line(point=True)
.encode(
x=alt.X("year:O", title="Year"),
y=alt.Y(
"annual mean:Q", title="Employees (thousands, SA annual mean)"
),
color=alt.Color("series:N", title="CES series"),
tooltip=["series", "year", "annual mean"],
)
.properties(width=640, height=260, title="Health-care employment (CES)")
)
if len(_rules):
_cchart = _cchart + (
alt.Chart(_rules.to_pandas())
.mark_rule(strokeDash=[4, 4], color="gray")
.encode(x="year:O", tooltip=["exposure"])
)
_panels += [
mo.ui.altair_chart(_cchart),
mo.ui.table(_cdf, label="bls.ces — annual means"),
]
else:
_panels.append(mo.md(not_built("stack bls ces --write")))
_view = mo.vstack(_panels)
_view
return
@app.cell(hide_code=True)
def _(NOTES, REPLICA_PATH, con, mo, pl, q, store):
# ── 8. Provenance ──
@@ -1321,7 +1477,9 @@ def _(NOTES, REPLICA_PATH, con, mo, pl, q, store):
"UNION ALL SELECT 'code_event', count(*) FROM pfs.code_event "
"UNION ALL SELECT 'code_family', count(*) FROM pfs.code_family "
"UNION ALL SELECT 'code_exposure', count(*) FROM pfs.code_exposure "
"UNION ALL SELECT 'utilization', count(*) FROM pfs.utilization"
"UNION ALL SELECT 'utilization', count(*) FROM pfs.utilization "
"UNION ALL SELECT 'bls.oews', count(*) FROM bls.oews "
"UNION ALL SELECT 'bls.ces', count(*) FROM bls.ces"
)
_log = q(
"SELECT run_id, ingested_at, module, table_name, rule_id, source_file, sha256, rows, "

218
src/bls/ces.py Normal file
View File

@@ -0,0 +1,218 @@
"""CES health-care employment and earnings, monthly, from the BLS public
API (#695).
``SERIES`` is the closed list this module pulls: all-employee counts
(thousands, seasonally adjusted) and average hourly earnings for health
care (NAICS 62 less social assistance), ambulatory care, offices of
physicians, home health, hospitals and nursing/residential care — the
settings the care-management families are billed from. The unkeyed API
allows 10 years × 25 series per request and 25 requests a day; with
``BLS_API_KEY`` 20 years × 50 series. Every observation keeps its
series id, so the table itself is the provenance trail back to
https://data.bls.gov/timeseries/<id>.
fetch(client, ids, start, end, api_key="") → {series_id: [obs…]}
to_rows(data, item_keys=…) → CesRow per observation
cite(store, series_id) → bib Source per series
write(con, rows) → replace each series present
read_series(con, ids, annual=False) → rows (or annual means)
"""
from __future__ import annotations
from dataclasses import astuple, dataclass
from typing import Any, Mapping, Sequence
import httpx
from bib.item import Source
from bib.tag import Tag
API = "https://api.bls.gov/publicAPI/v2/timeseries/data/"
TABLE = "bls.ces"
SERIES_URL = "https://data.bls.gov/timeseries/{id}"
@dataclass(frozen=True)
class CesSeries:
series_id: str
industry_code: str
industry_name: str
data_type: str # "employment" | "avg_hourly_earnings"
seasonal: str # "S" | "U"
def _s(industry: str, name: str) -> list[CesSeries]:
"""CES id = ``CES`` + supersector ``65`` (education & health) +
six-digit industry + two-digit data type."""
return [
CesSeries(f"CES65{industry}01", industry, name, "employment", "S"),
CesSeries(f"CES65{industry}03", industry, name, "avg_hourly_earnings", "S"),
]
SERIES: dict[str, CesSeries] = {
s.series_id: s
for s in [
*_s("620000", "Health care"),
*_s("621000", "Ambulatory health care services"),
*_s("621100", "Offices of physicians"),
*_s("621600", "Home health care services"),
*_s("622000", "Hospitals"),
*_s("623000", "Nursing and residential care facilities"),
]
}
@dataclass(frozen=True)
class CesRow:
series_id: str
industry_code: str
industry_name: str
data_type: str
seasonal: str
year: int
month: int
value: float | None
item_key: str
# ── fetch ─────────────────────────────────────────────────────────────
def fetch(
client: httpx.Client,
series_ids: Sequence[str],
start: int,
end: int,
*,
api_key: str = "",
) -> dict[str, list[dict]]:
"""Every observation for *series_ids* between *start* and *end*
inclusive, chunked to the API's per-request limits."""
span, per = (20, 50) if api_key else (10, 25)
out: dict[str, list[dict]] = {sid: [] for sid in series_ids}
windows = [(y, min(y + span - 1, end)) for y in range(start, end + 1, span)]
for i in range(0, len(series_ids), per):
chunk = list(series_ids[i : i + per])
for lo, hi in windows:
body: dict[str, Any] = {
"seriesid": chunk,
"startyear": str(lo),
"endyear": str(hi),
}
if api_key:
body["registrationkey"] = api_key
r = client.post(API, json=body, timeout=120)
r.raise_for_status()
payload = r.json()
if payload.get("status") != "REQUEST_SUCCEEDED":
raise RuntimeError(
f"BLS API {payload.get('status')}: {'; '.join(payload.get('message') or [])}"
)
for s in (payload.get("Results") or {}).get("series", []):
out.setdefault(s["seriesID"], []).extend(s.get("data") or [])
return out
# ── rows ──────────────────────────────────────────────────────────────
def to_rows(
data: Mapping[str, Sequence[dict]], *, item_keys: Mapping[str, str]
) -> list[CesRow]:
out: list[CesRow] = []
for sid, obs in data.items():
meta = SERIES.get(sid)
for o in obs:
period = str(o.get("period") or "")
if not period.startswith("M"):
continue
raw = str(o.get("value") or "").strip()
try:
value: float | None = float(raw)
except ValueError:
value = None
out.append(
CesRow(
series_id=sid,
industry_code=meta.industry_code if meta else "",
industry_name=meta.industry_name if meta else "",
data_type=meta.data_type if meta else "",
seasonal=meta.seasonal if meta else "",
year=int(o["year"]),
month=int(period[1:]),
value=value,
item_key=item_keys.get(sid, ""),
)
)
return out
# ── bib ───────────────────────────────────────────────────────────────
def cite(store: Any, series_id: str) -> str:
meta = SERIES.get(series_id)
what = (
f"{meta.industry_name}, {meta.data_type.replace('_', ' ')}"
if meta
else series_id
)
item = Source(
title=f"Current Employment Statistics series {series_id} — {what} (BLS CES)",
url=SERIES_URL.format(id=series_id),
doc_type="dataset",
)
tags = [Tag.module("bls"), Tag.table(TABLE), Tag.source("bls-website")]
return store.upsert(item, tags=tags)
# ── DuckDB ────────────────────────────────────────────────────────────
def ensure_table(con: Any) -> None:
from bls.table.ces import Ces
con.execute("CREATE SCHEMA IF NOT EXISTS bls")
con.execute(Ces.to_ddl().replace("CREATE TABLE ", "CREATE TABLE IF NOT EXISTS ", 1))
_COLS = "series_id, industry_code, industry_name, data_type, seasonal, year, month, value, item_key"
def write(con: Any, rows: Sequence[CesRow]) -> int:
"""Replace every series that appears in *rows*."""
ensure_table(con)
for sid in sorted({r.series_id for r in rows}):
con.execute(f"DELETE FROM {TABLE} WHERE series_id = ?", [sid])
if not rows:
return 0
con.executemany(
f"INSERT INTO {TABLE} ({_COLS}) VALUES ({','.join('?' * 9)})",
[astuple(r) for r in rows],
)
return len(rows)
def read_series(
con: Any, series_ids: Sequence[str], *, annual: bool = False
) -> list[CesRow]:
"""Monthly rows ordered by series, year, month — or with ``annual``
one row per (series, year) holding the mean of the months (month 0)."""
if not series_ids:
return []
marks = ",".join("?" * len(series_ids))
if annual:
sql = (
"SELECT series_id, any_value(industry_code), any_value(industry_name), "
"any_value(data_type), any_value(seasonal), year, 0, avg(value), any_value(item_key) "
f"FROM {TABLE} WHERE series_id IN ({marks}) AND month BETWEEN 1 AND 12 "
"GROUP BY series_id, year ORDER BY series_id, year"
)
else:
sql = (
f"SELECT {_COLS} FROM {TABLE} WHERE series_id IN ({marks}) "
"ORDER BY series_id, year, month"
)
return [CesRow(*r) for r in con.execute(sql, list(series_ids)).fetchall()]

259
src/bls/oews.py Normal file
View File

@@ -0,0 +1,259 @@
"""OEWS national occupation wages, one May release per year (#695).
The BLS API only serves the *latest* OEWS year (OEWS is not a time
series in BLS's own terms), so history comes from the annual
``oesm{yy}nat.zip`` files on bls.gov — which refuse plain scripted
downloads (403) but serve a request that carries a full browser header
set (``_HEADERS``). Each zip holds ``national_M{year}_dl.xls[x]`` with
one row per SOC occupation; 2013–2018 use ``OCC_GROUP`` (2010 SOC),
2019+ use ``O_GROUP`` plus area/industry columns (2018 SOC), and both
are normalised to :class:`OewsRow`. Suppressed cells (``*`` wage not
available, ``#`` ≥ $115/hr or $239,200/yr, ``**`` employment not
released) load as NULL.
download(client, year, dest) → the zip on disk (validated)
read_national(zip_path, year) → the national sheet as dicts
to_rows(year, records, item_key=…) → OewsRow per national occupation row
cite(store, year) → bib Source for that release
write(con, year, rows) → replace that year
read_occupations(con, codes) → rows per (occupation, year)
"""
from __future__ import annotations
import io
import re
import zipfile
from dataclasses import dataclass
from pathlib import Path
from typing import Any, Sequence
import httpx
from bib.item import Source
from bib.tag import Tag
YEARS = range(2013, 2025)
TABLE = "bls.oews"
_HEADERS = {
"User-Agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0",
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.5",
"Referer": "https://www.bls.gov/oes/tables.htm",
"Sec-Fetch-Dest": "document",
"Sec-Fetch-Mode": "navigate",
"Sec-Fetch-Site": "same-origin",
"Upgrade-Insecure-Requests": "1",
}
_MEMBER = re.compile(r"national_M(\d{4})_dl\.xlsx?$", re.I)
@dataclass(frozen=True)
class OewsRow:
year: int
occ_code: str
occ_title: str
occ_group: str
tot_emp: int | None
emp_prse: float | None
h_mean: float | None
a_mean: float | None
mean_prse: float | None
h_pct10: float | None
h_pct25: float | None
h_median: float | None
h_pct75: float | None
h_pct90: float | None
a_pct10: float | None
a_pct25: float | None
a_median: float | None
a_pct75: float | None
a_pct90: float | None
series_id: str
item_key: str
def zip_url(year: int) -> str:
return f"https://www.bls.gov/oes/special-requests/oesm{year % 100:02d}nat.zip"
def series_id(occ_code: str, *, datatype: str = "") -> str:
"""``OEUN`` + national area ``0000000`` + cross-industry ``000000`` +
the occupation's six digits (+ optional two-digit datatype)."""
digits = re.sub(r"\D", "", occ_code)
return f"OEUN000000000000{digits}{datatype}"
def cache_path(year: int, *, root: Path | None = None) -> Path:
if root is None:
from conf import path
root = path("db.aco").parent / "bls" / "oews"
return Path(root) / f"oesm{year % 100:02d}nat.zip"
# ── fetch ─────────────────────────────────────────────────────────────
def download(client: httpx.Client, year: int, dest: Path) -> Path:
"""Fetch *year*'s national zip to *dest* with browser headers; a
non-200 raises ``httpx.HTTPStatusError`` and a non-zip body (the
bot-check HTML page) raises ``ValueError``."""
r = client.get(zip_url(year), headers=_HEADERS, timeout=120, follow_redirects=True)
r.raise_for_status()
if not r.content.startswith(b"PK"):
raise ValueError(
f"{zip_url(year)}: not a zip ({r.headers.get('content-type', '?')})"
)
dest.parent.mkdir(parents=True, exist_ok=True)
dest.write_bytes(r.content)
return dest
def read_national(zip_path: Path, year: int) -> list[dict]:
"""The ``national_M{year}_dl`` sheet as a list of dicts (header row
→ keys), whichever of xls/xlsx the year shipped."""
import polars as pl
with zipfile.ZipFile(zip_path) as z:
members = [n for n in z.namelist() if _MEMBER.search(n)]
if not members:
raise FileNotFoundError(
f"{zip_path}: no national_M{year}_dl sheet in {z.namelist()}"
)
member = members[0]
raw = z.read(member)
ext = member.rsplit(".", 1)[-1].lower()
df = (
pl.read_excel(io.BytesIO(raw), infer_schema_length=0)
if ext == "xlsx"
else pl.read_excel(io.BytesIO(raw), infer_schema_length=0, engine="calamine")
)
return df.to_dicts()
# ── rows ──────────────────────────────────────────────────────────────
_SUPPRESSED = {"*", "**", "#", "", "-", "N/A"}
def _num(v: Any) -> float | None:
if v is None:
return None
s = str(v).strip().replace(",", "")
if s in _SUPPRESSED:
return None
try:
return float(s)
except ValueError:
return None
def _int(v: Any) -> int | None:
f = _num(v)
return int(f) if f is not None else None
def to_rows(year: int, records: Sequence[dict], *, item_key: str) -> list[OewsRow]:
"""National rows (2019+ files also carry no other areas in the
``nat`` zip, but the AREA column is checked so a mis-filed state
sheet never loads as national)."""
out: list[OewsRow] = []
for rec in records:
area = str(rec.get("AREA", "99") or "99").strip()
if area not in ("99", "0", "00", "0000000"):
continue
code = str(rec.get("OCC_CODE") or "").strip()
if not code:
continue
out.append(
OewsRow(
year=year,
occ_code=code,
occ_title=str(rec.get("OCC_TITLE") or "").strip(),
occ_group=str(rec.get("O_GROUP") or rec.get("OCC_GROUP") or "").strip(),
tot_emp=_int(rec.get("TOT_EMP")),
emp_prse=_num(rec.get("EMP_PRSE")),
h_mean=_num(rec.get("H_MEAN")),
a_mean=_num(rec.get("A_MEAN")),
mean_prse=_num(rec.get("MEAN_PRSE")),
h_pct10=_num(rec.get("H_PCT10")),
h_pct25=_num(rec.get("H_PCT25")),
h_median=_num(rec.get("H_MEDIAN")),
h_pct75=_num(rec.get("H_PCT75")),
h_pct90=_num(rec.get("H_PCT90")),
a_pct10=_num(rec.get("A_PCT10")),
a_pct25=_num(rec.get("A_PCT25")),
a_median=_num(rec.get("A_MEDIAN")),
a_pct75=_num(rec.get("A_PCT75")),
a_pct90=_num(rec.get("A_PCT90")),
series_id=series_id(code),
item_key=item_key,
)
)
return out
# ── bib ───────────────────────────────────────────────────────────────
def cite(store: Any, year: int) -> str:
item = Source(
title=f"Occupational Employment and Wage Statistics, May {year} national estimates (BLS OEWS)",
url=zip_url(year),
doc_type="dataset",
)
tags = [
Tag.module("bls"),
Tag.table(TABLE),
Tag.source("bls-website"),
Tag.year(year),
]
return store.upsert(item, tags=tags)
# ── DuckDB ────────────────────────────────────────────────────────────
def ensure_table(con: Any) -> None:
from bls.table.oews import Oews
con.execute("CREATE SCHEMA IF NOT EXISTS bls")
con.execute(
Oews.to_ddl().replace("CREATE TABLE ", "CREATE TABLE IF NOT EXISTS ", 1)
)
_COLS = (
"year, occ_code, occ_title, occ_group, tot_emp, emp_prse, h_mean, a_mean, mean_prse, "
"h_pct10, h_pct25, h_median, h_pct75, h_pct90, a_pct10, a_pct25, a_median, a_pct75, "
"a_pct90, series_id, item_key"
)
def write(con: Any, year: int, rows: Sequence[OewsRow]) -> int:
ensure_table(con)
con.execute(f"DELETE FROM {TABLE} WHERE year = ?", [year])
if not rows:
return 0
from dataclasses import astuple
con.executemany(
f"INSERT INTO {TABLE} ({_COLS}) VALUES ({','.join('?' * 21)})",
[astuple(r) for r in rows],
)
return len(rows)
def read_occupations(
con: Any, occ_codes: Sequence[str], *, years: Sequence[int] | None = None
) -> list[OewsRow]:
if not occ_codes:
return []
sql = f"SELECT {_COLS} FROM {TABLE} WHERE occ_code IN ({','.join('?' * len(occ_codes))})"
params: list[Any] = list(occ_codes)
if years:
sql += f" AND year IN ({','.join('?' * len(years))})"
params.extend(int(y) for y in years)
sql += " ORDER BY occ_code, year"
return [OewsRow(*r) for r in con.execute(sql, params).fetchall()]

View File

@@ -1 +1,6 @@
__all__ = []
"""BLS Pydantic table models."""
from .ces import Ces as Ces
from .oews import Oews as Oews
__all__ = ["Ces", "Oews"]

37
src/bls/table/ces.py Normal file
View File

@@ -0,0 +1,37 @@
"""BLS Current Employment Statistics — health-care employment and
earnings series, monthly, from the BLS public API (#695)."""
from __future__ import annotations
from conf.table_base import SQLTable
class Ces(SQLTable):
"""One monthly observation of one CES series."""
__schema__ = "bls"
__tablename__ = "ces"
series_id: str | None = None
"""BLS series id, e.g. ``CES6562110001``."""
industry_code: str | None = None
"""CES industry code inside the series id (``656211`` = offices of physicians)."""
industry_name: str | None = None
"""Industry title from the series registry."""
data_type: str | None = None
"""``employment`` (thousands, all employees) or ``avg_hourly_earnings`` (dollars)."""
seasonal: str | None = None
"""``S`` seasonally adjusted / ``U`` not."""
year: int | None = None
month: int | None = None
"""1–12; ``13`` is the BLS annual average row."""
value: float | None = None
item_key: str | None = None
"""bib key of the Source item citing this series."""

65
src/bls/table/oews.py Normal file
View File

@@ -0,0 +1,65 @@
"""BLS Occupational Employment and Wage Statistics — national, cross-industry
estimates per SOC occupation, one May release per year (#695).
Source: the ``oesm{yy}nat.zip`` "special requests" files on bls.gov,
read by ``bls.oews``. The 2013–2018 files (2010 SOC) and 2019+ files
(2018 SOC) are normalised to one shape; SOC code changes across that
boundary are the caller's problem to reconcile, not hidden here.
"""
from __future__ import annotations
from conf.table_base import SQLTable
class Oews(SQLTable):
"""National OEWS estimates per occupation and release year."""
__schema__ = "bls"
__tablename__ = "oews"
year: int | None = None
"""Release year (May reference period)."""
occ_code: str | None = None
"""SOC occupation code, e.g. ``31-9092``."""
occ_title: str | None = None
"""SOC occupation title."""
occ_group: str | None = None
"""``total`` / ``major`` / ``minor`` / ``broad`` / ``detailed``."""
tot_emp: int | None = None
"""Estimated employment (rounded to the nearest 10)."""
emp_prse: float | None = None
"""Employment percent relative standard error."""
h_mean: float | None = None
"""Mean hourly wage."""
a_mean: float | None = None
"""Mean annual wage."""
mean_prse: float | None = None
"""Wage percent relative standard error."""
h_pct10: float | None = None
h_pct25: float | None = None
h_median: float | None = None
h_pct75: float | None = None
h_pct90: float | None = None
a_pct10: float | None = None
a_pct25: float | None = None
a_median: float | None = None
a_pct75: float | None = None
a_pct90: float | None = None
series_id: str | None = None
"""BLS API series-id prefix for the national cross-industry estimate
(``OEUN000000000000`` + occupation digits); append the two-digit
datatype (``01`` employment, ``04`` annual mean, ``13`` hourly mean)."""
item_key: str | None = None
"""bib key of the Source item citing this release's zip."""

View File

@@ -14,6 +14,7 @@ import typer
from cli.api import app as api_app
from cli.bib import app as bib_app
from cli.bls import app as bls_app
from cli.comments import app as comments_app
from cli.db import app as db_app
from cli.docs import app as docs_app
@@ -43,6 +44,11 @@ app.command()(health)
app.add_typer(load_app, name="load", help="Ingest data (CCLF, BCDA, seeds).")
app.add_typer(generate_app, name="generate", help="Regenerate code artefacts.")
app.add_typer(bib_app, name="bib", help="Bibliography operations.")
app.add_typer(
bls_app,
name="bls",
help="BLS occupation wages (OEWS) and health-care employment (CES).",
)
app.add_typer(
comments_app,
name="comments",

225
src/cli/bls.py Normal file
View File

@@ -0,0 +1,225 @@
"""stack bls — BLS occupation wages (OEWS) and health-care employment (CES).
uv run stack bls oews --write # every May release 2013–
uv run stack bls oews --occ 31-9092 --occ 29-1141
uv run stack bls ces --write # the registry's series, 2006–
uv run stack bls ces --series CES6562110001 --annual
"""
from __future__ import annotations
import os
from typing import Any
import typer
from bls.ces import SERIES as CES_SERIES
from bls.ces import cite as cite_ces
from bls.ces import fetch as fetch_ces
from bls.ces import read_series as read_ces
from bls.ces import to_rows as ces_rows
from bls.ces import write as write_ces
from bls.oews import YEARS as OEWS_YEARS
from bls.oews import cache_path as oews_cache_path
from bls.oews import cite as cite_oews
from bls.oews import download as download_oews
from bls.oews import read_national as read_oews_national
from bls.oews import read_occupations as read_oews
from bls.oews import to_rows as oews_rows
from bls.oews import write as write_oews
from pfs.codetables import is_missing_table_error
app = typer.Typer(no_args_is_help=True)
# ── indirections the tests monkeypatch ───────────────────────────────
def _batch() -> Any:
from conf.connect import duckdb_batch
return duckdb_batch("aco")
def _read() -> Any:
from conf.connect import duckdb
return duckdb("aco", read_only=True)
def _store() -> Any:
from conf.connect import bib
return bib()
def _publish() -> None:
from conf.connect import publish_replica
typer.echo(f"replica → {publish_replica('aco')}")
def _http() -> Any:
import httpx
return httpx.Client()
def _fmt(v: float | None, money: bool = False) -> str:
if v is None:
return "—"
return f"${v:,.2f}" if money else f"{v:,.0f}"
@app.command()
def oews(
year: list[int] = typer.Option(
[], "--year", help="Release year(s); default 2013–."
),
occ: list[str] = typer.Option(
[], "--occ", help="SOC code(s) to print, e.g. 31-9092."
),
write: bool = typer.Option(
False, "--write", help="Download, load bls.oews, cite, republish."
),
offline: bool = typer.Option(
False, "--offline", help="With --write: use the cached zips."
),
) -> None:
"""OEWS national occupation wages (bls.oews, #695): one May release
per year from the bls.gov oesm{yy}nat.zip files, every SOC occupation,
employment and hourly/annual wage distribution, cited in bib per
release and logged in cms.ingest_log. --occ prints the series."""
codes = [c.strip() for c in occ]
if write:
years = sorted(year) if year else list(OEWS_YEARS)
bad = [y for y in years if y not in OEWS_YEARS]
if bad:
raise typer.BadParameter(
f"no OEWS release for {bad}; known {OEWS_YEARS.start}–{OEWS_YEARS.stop - 1}"
)
store = _store()
staged: list[tuple[int, list[Any], str, str]] = []
client = None if offline else _http()
try:
for y in years:
zp = oews_cache_path(y)
if not offline:
download_oews(client, y, zp)
key = cite_oews(store, y)
rows = oews_rows(y, read_oews_national(zp, y), item_key=key)
staged.append((y, rows, key, str(zp)))
typer.echo(f"{y}: {len(rows)} occupations ← {key}")
finally:
if client is not None:
client.close()
from cms.ingest_log import log_ingest
with _batch() as con:
for y, rows, key, zp in staged:
n = write_oews(con, y, rows)
log_ingest(
con,
module="bls.oews",
table_name="bls.oews",
rows=n,
source_file=zp,
pincite_key=key,
)
_publish()
if not codes:
return
if not codes:
raise typer.BadParameter("pass --occ (repeatable) or --write")
con = _read()
try:
try:
rows = read_oews(con, codes)
except Exception as exc:
if is_missing_table_error(exc):
typer.echo("no bls.oews yet — run with --write")
return
raise
for r in rows:
typer.echo(
f"{r.occ_code} {r.year} emp {_fmt(r.tot_emp):>10} "
f"mean {_fmt(r.h_mean, True):>8}/hr {_fmt(r.a_mean, True):>12}/yr "
f"median {_fmt(r.a_median, True):>12}/yr {r.occ_title}"
)
typer.echo(f"oews: {len(rows)} (occupation, year) rows")
finally:
con.close()
@app.command()
def ces(
series: list[str] = typer.Option(
[], "--series", help="CES series id(s) to print; default all in the registry."
),
start: int = typer.Option(2006, "--start", help="First year to pull with --write."),
end: int = typer.Option(2025, "--end", help="Last year to pull with --write."),
annual: bool = typer.Option(
False, "--annual", help="Print annual means instead of months."
),
write: bool = typer.Option(
False,
"--write",
help="Pull the registry's series, load bls.ces, cite, republish.",
),
) -> None:
"""CES health-care employment and earnings (bls.ces, #695): monthly,
seasonally adjusted, for health care / ambulatory care / offices of
physicians / home health / hospitals / nursing care, via the BLS
public API (BLS_API_KEY widens the per-request limits), one bib
Source per series."""
ids = [s.strip().upper() for s in series] or list(CES_SERIES)
if write:
store = _store()
client = _http()
try:
data = fetch_ces(
client,
list(CES_SERIES),
start,
end,
api_key=os.environ.get("BLS_API_KEY", ""),
)
finally:
client.close()
keys = {sid: cite_ces(store, sid) for sid in CES_SERIES}
rows = ces_rows(data, item_keys=keys)
from cms.ingest_log import log_ingest
with _batch() as con:
n = write_ces(con, rows)
log_ingest(
con,
module="bls.ces",
table_name="bls.ces",
rows=n,
source_file="",
pincite_key=",".join(sorted(set(keys.values()))),
)
_publish()
per = {}
for r in rows:
per[r.series_id] = per.get(r.series_id, 0) + 1
for sid in sorted(per):
typer.echo(f"{sid}: {per[sid]} months ← {keys.get(sid, '')}")
typer.echo(f"ces: {n} observations, {len(per)} series")
return
con = _read()
try:
try:
rows = read_ces(con, ids, annual=annual)
except Exception as exc:
if is_missing_table_error(exc):
typer.echo("no bls.ces yet — run with --write")
return
raise
for r in rows:
when = f"{r.year}" if annual else f"{r.year}-{r.month:02d}"
typer.echo(
f"{r.series_id} {when} {_fmt(r.value, r.data_type == 'avg_hourly_earnings')} {r.industry_name} {r.data_type}"
)
typer.echo(f"ces: {len(rows)} rows")
finally:
con.close()

213
tests/bls/test_ces.py Normal file
View File

@@ -0,0 +1,213 @@
"""bls.ces — Current Employment Statistics health-care series via the BLS API (#695)."""
from __future__ import annotations
import json
import duckdb
import httpx
import pytest
from bib.store import Store
from bls.ces import (
SERIES,
CesRow,
cite,
fetch,
read_series,
to_rows,
write,
)
def _payload(series_ids, years, *, message=None):
return {
"status": "REQUEST_SUCCEEDED",
"responseTime": 1,
"message": message or [],
"Results": {
"series": [
{
"seriesID": sid,
"data": [
{
"year": str(y),
"period": f"M{m:02d}",
"periodName": "x",
"value": f"{y}.{m}",
"footnotes": [{}],
}
for y in years
for m in (12, 1)
],
}
for sid in series_ids
]
},
}
class TestRegistry:
def test_series_cover_health_care_employment_and_earnings(self):
assert "CES6562000001" in SERIES and "CES6562110001" in SERIES
assert SERIES["CES6562110001"].industry_name.startswith("Offices of physicians")
assert SERIES["CES6562000003"].data_type == "avg_hourly_earnings"
assert all(s.seasonal == "S" for s in SERIES.values())
class TestFetch:
def test_chunks_years_and_series_and_posts_json(self):
calls = []
def handler(request):
body = json.loads(request.content)
calls.append(body)
return httpx.Response(
200,
json=_payload(
body["seriesid"],
range(int(body["startyear"]), int(body["endyear"]) + 1),
),
)
client = httpx.Client(transport=httpx.MockTransport(handler))
ids = [f"CES{i:010d}" for i in range(30)]
out = fetch(client, ids, 2006, 2025)
# unkeyed: 10 years and 25 series per request → 2 year windows × 2 series chunks
assert len(calls) == 4
assert {(c["startyear"], c["endyear"]) for c in calls} == {
("2006", "2015"),
("2016", "2025"),
}
assert max(len(c["seriesid"]) for c in calls) == 25
assert "registrationkey" not in calls[0]
assert len(out) == 30 and all(len(v) == 40 for v in out.values())
def test_key_widens_the_windows(self):
calls = []
def handler(request):
body = json.loads(request.content)
calls.append(body)
return httpx.Response(
200,
json=_payload(
body["seriesid"],
range(int(body["startyear"]), int(body["endyear"]) + 1),
),
)
client = httpx.Client(transport=httpx.MockTransport(handler))
fetch(client, ["CES6562000001"], 2006, 2025, api_key="k")
assert len(calls) == 1 and calls[0]["registrationkey"] == "k"
def test_api_failure_status_raises(self):
client = httpx.Client(
transport=httpx.MockTransport(
lambda r: httpx.Response(
200,
json={
"status": "REQUEST_NOT_PROCESSED",
"message": ["daily threshold"],
"Results": {},
},
)
)
)
with pytest.raises(RuntimeError, match="daily threshold"):
fetch(client, ["CES6562000001"], 2016, 2025)
def test_http_error_raises(self):
client = httpx.Client(
transport=httpx.MockTransport(lambda r: httpx.Response(500))
)
with pytest.raises(httpx.HTTPStatusError):
fetch(client, ["CES6562000001"], 2016, 2025)
class TestToRows:
def test_rows_carry_series_metadata(self):
data = {
"CES6562110001": [
{
"year": "2024",
"period": "M03",
"periodName": "March",
"value": "3010.5",
"footnotes": [{}],
}
]
}
rows = to_rows(data, item_keys={"CES6562110001": "K"})
assert len(rows) == 1
r = rows[0]
assert isinstance(r, CesRow)
assert (r.series_id, r.year, r.month, r.value) == (
"CES6562110001",
2024,
3,
3010.5,
)
assert (
r.industry_code == "621100"
and r.data_type == "employment"
and r.item_key == "K"
)
def test_annual_average_and_blank_values(self):
data = {
"CES6562000001": [
{"year": "2024", "period": "M13", "value": "-"},
{"year": "2024", "period": "M01", "value": "1"},
]
}
rows = to_rows(data, item_keys={})
assert [(r.month, r.value) for r in rows] == [(13, None), (1, 1.0)]
def test_unknown_series_still_loads(self):
rows = to_rows(
{"CES9999999999": [{"year": "2024", "period": "M01", "value": "2"}]},
item_keys={},
)
assert rows[0].industry_code == "" and rows[0].data_type == ""
@pytest.fixture
def con():
c = duckdb.connect(":memory:")
yield c
c.close()
class TestWriteRead:
def test_write_replaces_per_series(self, con):
rows = to_rows(
{
"CES6562110001": [
{"year": "2024", "period": "M01", "value": "1"},
{"year": "2024", "period": "M02", "value": "2"},
],
"CES6562000001": [{"year": "2024", "period": "M01", "value": "9"}],
},
item_keys={},
)
assert write(con, rows) == 3
assert write(con, rows) == 3
assert con.execute("SELECT count(*) FROM bls.ces").fetchone()[0] == 3
out = read_series(con, ["CES6562110001"])
assert [(r.year, r.month, r.value) for r in out] == [
(2024, 1, 1.0),
(2024, 2, 2.0),
]
assert read_series(con, ["CES6562110001"], annual=True)[0].value == 1.5
class TestCite:
def test_one_source_per_series(self, tmp_path):
s = Store(":memory:", storage_dir=tmp_path / "st")
k1, k2 = cite(s, "CES6562110001"), cite(s, "CES6562110001")
assert k1 == k2
item = s.get(k1)
assert "CES6562110001" in item.title and item.url.endswith("CES6562110001")
assert {"module:bls", "table:bls.ces", "source:bls-website"} <= set(item.tags)
s.close()

382
tests/bls/test_oews.py Normal file
View File

@@ -0,0 +1,382 @@
"""bls.oews — OEWS national occupation wages from the annual zips (#695)."""
from __future__ import annotations
import io
import zipfile
import duckdb
import httpx
import openpyxl
import pytest
from bib.store import Store
from bls.oews import (
YEARS,
OewsRow,
cite,
download,
read_national,
read_occupations,
series_id,
to_rows,
write,
zip_url,
)
OLD_COLS = [
"OCC_CODE",
"OCC_TITLE",
"OCC_GROUP",
"TOT_EMP",
"EMP_PRSE",
"H_MEAN",
"A_MEAN",
"MEAN_PRSE",
"H_PCT10",
"H_PCT25",
"H_MEDIAN",
"H_PCT75",
"H_PCT90",
"A_PCT10",
"A_PCT25",
"A_MEDIAN",
"A_PCT75",
"A_PCT90",
"ANNUAL",
"HOURLY",
]
NEW_COLS = [
"AREA",
"AREA_TITLE",
"AREA_TYPE",
"PRIM_STATE",
"NAICS",
"NAICS_TITLE",
"I_GROUP",
"OWN_CODE",
"OCC_CODE",
"OCC_TITLE",
"O_GROUP",
"TOT_EMP",
"EMP_PRSE",
"JOBS_1000",
"LOC_QUOTIENT",
"PCT_TOTAL",
"PCT_RPT",
"H_MEAN",
"A_MEAN",
"MEAN_PRSE",
"H_PCT10",
"H_PCT25",
"H_MEDIAN",
"H_PCT75",
"H_PCT90",
"A_PCT10",
"A_PCT25",
"A_MEDIAN",
"A_PCT75",
"A_PCT90",
"ANNUAL",
"HOURLY",
]
def _xlsx(cols, rows) -> bytes:
wb = openpyxl.Workbook()
ws = wb.active
ws.append(cols)
for r in rows:
ws.append(r)
buf = io.BytesIO()
wb.save(buf)
return buf.getvalue()
def _zip(year: int, cols, rows) -> bytes:
buf = io.BytesIO()
with zipfile.ZipFile(buf, "w") as z:
z.writestr(
f"oesm{year % 100:02d}nat/field_descriptions.xlsx", _xlsx(["a"], [["b"]])
)
z.writestr(
f"oesm{year % 100:02d}nat/national_M{year}_dl.xlsx", _xlsx(cols, rows)
)
return buf.getvalue()
OLD_ROWS = [
[
"00-0000",
"All Occupations",
"total",
132588810,
0.1,
"22.33",
46440,
0.1,
"8.74",
"10.9",
"16.87",
"27.34",
"41.7",
18180,
22670,
35080,
56860,
86730,
None,
None,
],
[
"31-9092",
"Medical Assistants",
"detailed",
571690,
0.8,
"14.8",
30780,
0.3,
"10.23",
"12.09",
"14.24",
"17.12",
"20.5",
21270,
25150,
29610,
35610,
42650,
None,
None,
],
[
"29-1069",
"Physicians and Surgeons, All Other",
"detailed",
308570,
1.2,
"#",
"*",
0.5,
"#",
"#",
"#",
"#",
"#",
"*",
"*",
"*",
"*",
"*",
None,
None,
],
]
NEW_ROWS = [
[
"99",
"U.S.",
"1",
"US",
"000000",
"Cross-industry",
"cross-industry",
"1235",
"31-9092",
"Medical Assistants",
"detailed",
763040,
0.5,
5.025,
1.0,
None,
None,
20.2,
42000,
0.2,
15.1,
17.3,
19.75,
22.1,
25.5,
31400,
36000,
41070,
46000,
53000,
None,
None,
],
[
"99",
"U.S.",
"1",
"US",
"000000",
"Cross-industry",
"cross-industry",
"1235",
"29-1141",
"Registered Nurses",
"detailed",
3175390,
0.4,
20.9,
1.0,
None,
None,
45.42,
94480,
0.3,
30.1,
35.9,
42.1,
52.4,
66.5,
63720,
75000,
86070,
108900,
132680,
None,
None,
],
]
class TestYearsAndIds:
def test_years_and_zip_urls(self):
assert list(YEARS) == list(range(2013, 2025))
assert zip_url(2023) == "https://www.bls.gov/oes/special-requests/oesm23nat.zip"
assert zip_url(2013).endswith("oesm13nat.zip")
def test_series_id_is_the_national_cross_industry_prefix(self):
# OEUN + area 0000000 + industry 000000 + occupation digits; datatype appended by callers
assert series_id("31-9092") == "OEUN000000000000319092"
assert series_id("29-1141", datatype="04") == "OEUN00000000000029114104"
class TestReadNational:
def test_old_layout(self, tmp_path):
p = tmp_path / "oesm13nat.zip"
p.write_bytes(_zip(2013, OLD_COLS, OLD_ROWS))
recs = read_national(p, 2013)
assert len(recs) == 3
assert recs[1]["OCC_CODE"] == "31-9092" and recs[1]["OCC_GROUP"] == "detailed"
def test_new_layout(self, tmp_path):
p = tmp_path / "oesm23nat.zip"
p.write_bytes(_zip(2023, NEW_COLS, NEW_ROWS))
recs = read_national(p, 2023)
assert len(recs) == 2 and recs[0]["O_GROUP"] == "detailed"
def test_missing_member(self, tmp_path):
p = tmp_path / "bad.zip"
with zipfile.ZipFile(p, "w") as z:
z.writestr("oesm23nat/other.xlsx", b"x")
with pytest.raises(FileNotFoundError):
read_national(p, 2023)
class TestToRows:
def test_old_layout_rows_and_suppression(self):
rows = to_rows(2013, [dict(zip(OLD_COLS, r)) for r in OLD_ROWS], item_key="K")
assert [r.occ_code for r in rows] == ["00-0000", "31-9092", "29-1069"]
ma = rows[1]
assert isinstance(ma, OewsRow)
assert (ma.year, ma.occ_group, ma.tot_emp) == (2013, "detailed", 571690)
assert ma.h_mean == pytest.approx(14.8) and ma.a_mean == 30780.0
assert ma.a_median == 29610.0 and ma.h_pct90 == pytest.approx(20.5)
assert ma.series_id == "OEUN000000000000319092" and ma.item_key == "K"
phys = rows[2]
assert (
phys.h_mean is None and phys.a_mean is None and phys.a_pct10 is None
) # "#"/"*" → NULL
assert phys.tot_emp == 308570
def test_new_layout_rows(self):
rows = to_rows(2023, [dict(zip(NEW_COLS, r)) for r in NEW_ROWS], item_key="K")
assert [(r.occ_code, r.occ_group) for r in rows] == [
("31-9092", "detailed"),
("29-1141", "detailed"),
]
assert rows[0].a_mean == 42000.0 and rows[0].h_median == pytest.approx(19.75)
assert rows[1].a_pct90 == 132680.0
def test_non_national_rows_are_skipped(self):
state = dict(zip(NEW_COLS, NEW_ROWS[0]))
state.update(AREA="39", AREA_TITLE="Ohio", AREA_TYPE="2")
assert to_rows(2023, [state], item_key="K") == []
class TestDownload:
def test_sends_browser_headers_and_writes_file(self, tmp_path):
seen = {}
def handler(request):
seen["ua"] = request.headers.get("user-agent", "")
seen["accept"] = request.headers.get("accept", "")
seen["url"] = str(request.url)
return httpx.Response(200, content=_zip(2023, NEW_COLS, NEW_ROWS))
client = httpx.Client(transport=httpx.MockTransport(handler))
dest = tmp_path / "oesm23nat.zip"
out = download(client, 2023, dest)
assert out == dest and zipfile.is_zipfile(dest)
assert "Mozilla" in seen["ua"] and "text/html" in seen["accept"]
assert seen["url"] == zip_url(2023)
def test_403_raises(self, tmp_path):
client = httpx.Client(
transport=httpx.MockTransport(lambda r: httpx.Response(403, text="denied"))
)
with pytest.raises(httpx.HTTPStatusError):
download(client, 2023, tmp_path / "x.zip")
def test_non_zip_body_raises(self, tmp_path):
client = httpx.Client(
transport=httpx.MockTransport(
lambda r: httpx.Response(200, text="<html>bot check</html>")
)
)
with pytest.raises(ValueError, match="not a zip"):
download(client, 2023, tmp_path / "x.zip")
@pytest.fixture
def con():
c = duckdb.connect(":memory:")
yield c
c.close()
class TestWriteRead:
def test_write_replaces_year_and_read_filters_occupations(self, con):
r13 = to_rows(2013, [dict(zip(OLD_COLS, r)) for r in OLD_ROWS], item_key="A")
r23 = to_rows(2023, [dict(zip(NEW_COLS, r)) for r in NEW_ROWS], item_key="B")
assert write(con, 2013, r13) == 3 and write(con, 2023, r23) == 2
assert write(con, 2013, r13) == 3
assert con.execute("SELECT count(*) FROM bls.oews").fetchone()[0] == 5
out = read_occupations(con, ["31-9092"])
assert [(r.year, r.a_mean) for r in out] == [(2013, 30780.0), (2023, 42000.0)]
assert (
read_occupations(con, ["31-9092", "29-1141"], years=[2023])[0].occ_code
== "29-1141"
)
class TestCite:
def test_one_source_per_release(self, tmp_path):
s = Store(":memory:", storage_dir=tmp_path / "st")
k1, k2, k3 = cite(s, 2023), cite(s, 2023), cite(s, 2013)
assert k1 == k2 != k3
item = s.get(k1)
assert "May 2023" in item.title and item.url == zip_url(2023)
assert {
"module:bls",
"table:bls.oews",
"source:bls-website",
"year:2023",
} <= set(item.tags)
s.close()

View File

@@ -4,4 +4,4 @@
def test_import():
import bls.table
assert bls.table.__all__ == []
assert bls.table.__all__ == ["Ces", "Oews"]

234
tests/cli/test_bls_cli.py Normal file
View File

@@ -0,0 +1,234 @@
"""stack bls — oews / ces (#695)."""
from __future__ import annotations
import io
import zipfile
from unittest.mock import MagicMock
import duckdb
import openpyxl
import pytest
from typer.testing import CliRunner
import cli.bls as bls_cli
from bib.store import Store
from cli import app
runner = CliRunner()
COLS = [
"OCC_CODE",
"OCC_TITLE",
"OCC_GROUP",
"TOT_EMP",
"EMP_PRSE",
"H_MEAN",
"A_MEAN",
"MEAN_PRSE",
"H_PCT10",
"H_PCT25",
"H_MEDIAN",
"H_PCT75",
"H_PCT90",
"A_PCT10",
"A_PCT25",
"A_MEDIAN",
"A_PCT75",
"A_PCT90",
"ANNUAL",
"HOURLY",
]
ROWS = [
[
"31-9092",
"Medical Assistants",
"detailed",
571690,
0.8,
"14.8",
30780,
0.3,
"10.23",
"12.09",
"14.24",
"17.12",
"20.5",
21270,
25150,
29610,
35610,
42650,
None,
None,
]
]
def _zip_bytes(year: int) -> bytes:
wb = openpyxl.Workbook()
ws = wb.active
ws.append(COLS)
for r in ROWS:
ws.append(r)
xb = io.BytesIO()
wb.save(xb)
zb = io.BytesIO()
with zipfile.ZipFile(zb, "w") as z:
z.writestr(f"oesm{year % 100:02d}nat/national_M{year}_dl.xlsx", xb.getvalue())
return zb.getvalue()
@pytest.fixture
def con(monkeypatch, tmp_path):
c = duckdb.connect(":memory:")
class _Batch:
def __enter__(self):
return c
def __exit__(self, *a):
return False
monkeypatch.setattr(bls_cli, "_batch", lambda: _Batch())
monkeypatch.setattr(bls_cli, "_read", lambda: c.cursor())
store = Store(":memory:", storage_dir=tmp_path / "st")
monkeypatch.setattr(bls_cli, "_store", lambda: store)
published: list[bool] = []
monkeypatch.setattr(bls_cli, "_publish", lambda: published.append(True))
monkeypatch.setattr(bls_cli, "_http", lambda: MagicMock())
monkeypatch.setattr(
bls_cli,
"oews_cache_path",
lambda y, root=None: tmp_path / f"oesm{y % 100:02d}nat.zip",
)
def fake_download(client, year, dest):
dest.write_bytes(_zip_bytes(year))
return dest
monkeypatch.setattr(bls_cli, "download_oews", fake_download)
try:
yield c, store, published, tmp_path
finally:
store.close()
c.close()
class TestOews:
def test_write_loads_cites_logs_publishes(self, con):
c, store, published, tmp_path = con
res = runner.invoke(
app, ["bls", "oews", "--write", "--year", "2013", "--year", "2018"]
)
assert res.exit_code == 0, res.output
assert c.execute(
"SELECT count(*), count(distinct year) FROM bls.oews"
).fetchone() == (2, 2)
keys = [
r[0]
for r in c.execute(
"SELECT DISTINCT item_key FROM bls.oews ORDER BY 1"
).fetchall()
]
assert len(keys) == 2 and all(
"year:" in " ".join(store.get(k).tags) for k in keys
)
logs = c.execute(
"SELECT module, rows, source_file FROM cms.ingest_log ORDER BY source_file"
).fetchall()
assert [l[0] for l in logs] == ["bls.oews", "bls.oews"] and logs[0][2].endswith(
"oesm13nat.zip"
)
assert published == [True]
assert "2013: 1 occupations" in res.output
def test_offline_rerun_uses_cache_and_is_idempotent(self, con, monkeypatch):
c, store, published, tmp_path = con
runner.invoke(app, ["bls", "oews", "--write", "--year", "2013"])
monkeypatch.setattr(
bls_cli,
"download_oews",
lambda *a: (_ for _ in ()).throw(AssertionError("no download")),
)
res = runner.invoke(
app, ["bls", "oews", "--write", "--year", "2013", "--offline"]
)
assert res.exit_code == 0, res.output
assert c.execute("SELECT count(*) FROM bls.oews").fetchone()[0] == 1
def test_occ_reads_without_write_lock(self, con, monkeypatch):
c, store, published, tmp_path = con
runner.invoke(app, ["bls", "oews", "--write", "--year", "2013"])
published.clear()
monkeypatch.setattr(
bls_cli, "_batch", lambda: (_ for _ in ()).throw(AssertionError("no batch"))
)
res = runner.invoke(app, ["bls", "oews", "--occ", "31-9092"])
assert res.exit_code == 0, res.output
assert "31-9092 2013" in res.output and "$30,780.00/yr" in res.output
assert published == []
def test_unknown_year_rejected(self, con):
res = runner.invoke(app, ["bls", "oews", "--write", "--year", "1999"])
assert res.exit_code != 0
def test_no_table_yet(self, monkeypatch):
bare = duckdb.connect(":memory:")
monkeypatch.setattr(bls_cli, "_read", lambda: bare)
try:
res = runner.invoke(app, ["bls", "oews", "--occ", "31-9092"])
assert res.exit_code == 0 and "no bls.oews yet" in res.output
finally:
bare.close()
class TestCes:
def test_write_and_read(self, con, monkeypatch):
c, store, published, tmp_path = con
seen = {}
def fake_fetch(client, ids, start, end, *, api_key=""):
seen["ids"] = list(ids)
seen["key"] = api_key
return {
sid: [
{"year": "2024", "period": f"M{m:02d}", "value": str(m)}
for m in (1, 2)
]
for sid in ids
}
monkeypatch.setattr(bls_cli, "fetch_ces", fake_fetch)
monkeypatch.setenv("BLS_API_KEY", "abc")
res = runner.invoke(app, ["bls", "ces", "--write"])
assert res.exit_code == 0, res.output
assert seen["key"] == "abc" and "CES6562110001" in seen["ids"]
n_series = len(bls_cli.CES_SERIES)
assert c.execute("SELECT count(*) FROM bls.ces").fetchone()[0] == 2 * n_series
assert c.execute("SELECT module, rows FROM cms.ingest_log").fetchone() == (
"bls.ces",
2 * n_series,
)
assert published == [True]
assert f"ces: {2 * n_series} observations, {n_series} series" in res.output
published.clear()
monkeypatch.setattr(
bls_cli, "_batch", lambda: (_ for _ in ()).throw(AssertionError("no batch"))
)
res = runner.invoke(
app, ["bls", "ces", "--series", "CES6562110001", "--annual"]
)
assert res.exit_code == 0, res.output
assert "CES6562110001 2024 2 Offices of physicians employment" in res.output
assert published == []
def test_no_table_yet(self, monkeypatch):
bare = duckdb.connect(":memory:")
monkeypatch.setattr(bls_cli, "_read", lambda: bare)
try:
res = runner.invoke(app, ["bls", "ces"])
assert res.exit_code == 0 and "no bls.ces yet" in res.output
finally:
bare.close()

View File

@@ -52,6 +52,7 @@ def test_cells_are_anonymous_and_banners_present():
"7c",
"7d",
"7e",
"7f",
"8",
]