Harmonizing the historical and current codebooks

How RICF’s 2005–2017 historical tables and HTML filings and 2019–2025 Ministry of Civil Affairs filings were reconciled into one panel, what evidence justified each mapping, and how the result was validated.

Snapshot 2026-07-29. All figures in this document are reproducible from the commands in Reproducing this analysis.

The current build first expanded collection to report-form types 0 and 1, raising foundation filings from 23,883 to 31,521. Stage 3f then integrated the evidence-mapped financial projection of the historical HTML corpora. The canonical result is 11,159,082 retained source observations across 16,712 organizations and 198 variables. Every identity test below was re-run against the expanded snapshot. The six ambiguous published-table columns remain unmapped, and HTML contexts without explicit label/header evidence remain source_native_unmapped.

Contents

  1. The problem
  2. Design principles
  3. Step 1 — Inventory the vocabularies
  4. Step 2 — Match labels
  5. Step 3 — Establish unlabeled column meanings by test
  6. Step 4 — Assemble the panel
  7. Coverage results
  8. Presenting the result
  9. Validation
  10. Limitations
  11. Reproducing this analysis

The problem

RICF holds two bodies of evidence that had never shared a variable.

The published historical tables use three different header conventions across four years, and the financial tables add period and restriction suffixes on top:

Year Basic profile Financial statements
2013 RICF letter codes (ba_cn, ba_rdt) Chinese labels for position; letter codes for activities and cash flow
2014 Chinese labels (机构名称, 数据年度) Chinese labels throughout
2015 Chinese labels Both: 2015_financial_activities.tsv uses Chinese labels, fa_2015.tsv uses fa_dine_last-style codes
2016 RICF letter codes Letter codes with _begin/_end and _last/_this suffixes
2017 (companion) Letter codes Letter codes with no suffix

The current MCA tables are long-format. Each row is one line of the official form, and the columns are opaque platform identifiers:

annual_financial_position   920,320 rows
  akbe1347  货币资金 / 存货 / 长期股权投资 …   ← the official line label
  akbe1348  1 / 8 / 21 …                      ← the line's ordinal in the form
  akbe1349  9358464.34                        ← unlabeled value column
  akbe1350  19534467.34                       ← unlabeled value column

The project had treated the whole akbe* family as source_native_unmapped. That was correct for genuinely opaque fields, but it obscured a fact worth acting on: the financial tables are self-describing. akbe1347 holds the form’s own Chinese label, which is the same vocabulary the historical codebook indexes. What remained unknown was only the meaning of the unlabeled value columns.

The consequence was visible on every foundation profile. The coverage matrix showed historical H rows for 2013–2016 and current M rows for 2020–2024 that never shared a variable, and the only trend the site could draw for current years was labelled akbe1349 / akbe1350.

The same break showed up in the profile’s tables even after the panel existed, because the tables did not read the crosswalk. A historical financial-position row was named fp_csh; the current row carrying the same measurement was a Chinese label under four anonymous akbe* headings. Publishing the crosswalk’s vocabulary to the browser removed that asymmetry; see Presenting the result.

Design principles

Four constraints shaped the method, all from AGENTS.md:

  1. No guessing. A source field receives a human-readable meaning only when evidence supports it. Every mapping records the grade of evidence that justified it, so a reviewer can audit the weakest links first.
  2. No fuzzy matching. String similarity is never used. Mechanical normalization is allowed and documented; a genuine wording difference needs an explicit, hand-reviewed entry with its own justification.
  3. Nothing is discarded. Conflicting and duplicate evidence stays available. The panel is a derived layer; the published TSVs and the source-native MCA tables are untouched.
  4. The mapping is separable from the data. The crosswalk is a tracked JSON document reviewable on its own, and the panel builder reads nothing else.

Evidence grades

Grade Definition Panel rows
codebook_exact Source label and codebook field identical after Unicode NFKC normalization and whitespace removal 4,903,450
codebook_normalized Matched after documented mechanical folding only: list numbering (四、), item numbering ((一)), accounting prefixes (减:, 其中:), separator punctuation 585,159
reviewed_synonym Wording genuinely differs; each alias listed explicitly with the evidence for it 343,011
verified_identity An unlabeled column’s meaning established by an empirical test, with the pass rate recorded 120,848
current_only_variable The line exists only in the current form; the code is assigned from the form ordinal and no historical equivalent is claimed 69,444
source_native_unmapped No evidence supports a mapping; the field is retained verbatim

Step 1 — Inventory the vocabularies

The historical vocabulary comes from the four field sheets of RICF Codebook.xlsx, which give letter code, Chinese label, and English label for 197 variables. The current vocabulary is the distinct set of values in each long table’s label column, read from the snapshot.

Component Codebook variables Distinct current labels
Financial position 45 35
Financial activities 42 20
Cash flow 35 24
Basic profile 119 78

The current form holds fewer lines than the codebook because subtotals (流动资产合计, 资产合计, 收入合计) are computed by the platform rather than stored. Those variables remain in the crosswalk with no current mapping; researchers can derive them from their components.

One structural difference had to be resolved before any matching. The historical financial-activities codebook splits every income and expense line into two variables by restriction — fa_dine (捐赠收入_限定) and fa_dinu (捐赠收入_非限定) — while the current form states the line once and splits the restriction across value columns. The panel therefore keys financial-activities rows on the base variable (fa_din) and carries the restriction in measure, exactly as the balance sheet uses begin and end.

Step 2 — Match labels

Labels were matched through a three-rung ladder, stopping at the first rung that resolves.

Rung 1 — exact. NFKC-normalize, strip whitespace, compare. Resolves 31 of 35 financial-position labels immediately.

Rung 2 — mechanical normalization. Strip only decoration that carries no meaning: leading list numbering, parenthesized item numbering, accounting prefixes, and separator punctuation. This is deliberately conservative — it never changes a word.

四、汇率变动对现金的影响额  →  汇率变动对现金的影响额
减:累计折旧                →  累计折旧
(一)业务活动成本           →  业务活动成本
其中:捐赠收入              →  捐赠收入

Rung 3 — reviewed synonyms. Twenty-three wording differences remained. Each was resolved by hand and recorded with the evidence that justifies it — never by similarity scoring. The evidence is that both labels occupy the same slot in the same official form, verified against the snapshot’s ordinal column.

Component Aliases Example
Financial position 3 应收账款应收款项, both at form ordinal 3
Financial activities 1 限定性净资产转为非限定性资产…非限定性净资产, both at ordinal 40
Cash flow 9 收取会费收到的现金收取会员费收到的现金, both the second operating inflow
Basic profile 29 住所基金会地址, current form heading for the registered address; 志愿者数量志愿者数, 2014/2015 published wording

The cash-flow case is the clearest demonstration that this is not fuzzy matching. Every current cash-flow row maps to the codebook in strictly increasing order, with only the platform’s unstored subtotals skipped:

Ordinal Current label Codebook label Code Difference
2 收取会费收到的现金 收取会员费收到的现金 cf_soc 会费 / 会员费
5 政府补助收到的现金 政府补贴收到的现金 cf_govc 补助 / 补贴
8 收到的其他与业务活动有关的现金 收到的与其他业务活动有关的现金 cf_otca word order
16 购买商品、接受服务支付的现金 购买商品接受服务支付的现金 cf_psc punctuation
30 收到的其他与投资活动有关的现金 收到的其他与投资有关的现金 cf_oc 活动 inserted
52 偿付利息所支付的现金 偿还利息所支付的现金 cf_inpc 偿付 / 偿还
60 四、汇率变动对现金的影响额 汇率变动对现金的影响 cf_exch numbering, trailing 额
61 五、现金及现金等价物净增加额 现金及现金等价物净增加额 cf_ninc numbering

That ordering isomorphism is the evidence. A similarity score would have produced the same answers with none of the justification.

Result: every observed current label in all three financial components resolves. Zero remain unmapped.

Step 3 — Establish unlabeled column meanings by test

Matching labels identifies which line a row describes. It says nothing about what the numeric columns beside it mean. Those were resolved empirically, and each test records its pass rate so a reviewer audits a number rather than an assertion.

3.1 Balance sheet: which column is the opening balance?

annual_financial_position has two unlabeled value columns.

Test. A balance sheet’s opening balance for year T must equal the same organization’s closing balance in its year T−1 filing. Join every filing to its predecessor on organization and form ordinal, and compare.

Reading Agreement n
akbe1349 = opening, akbe1350 = closing 92.2% 104,486
Columns reversed 10.1%  
Columns unchanged year over year 14.7%  

The accepted reading beats the best alternative by a factor of 9.4.

3.2 Financial activities: four unlabeled columns

annual_financial_activities has four. Three independent tests were run, each capable of falsifying the others.

Test A — cross-filing. If one pair carries the prior reporting period, it must reproduce the same organization’s previous-year filing.

Reading Agreement n
akbe1364/akbe1365 = prior period, akbe1366/akbe1367 = current 85.4% 129,237
Current columns carry the prior period 1.1%  
Columns unchanged year over year 2.1%  
Restriction columns swapped 0.7%  

A separation of roughly 45× against every alternative. The 13.6% shortfall is not noise in the test — it is restatement, which is itself recorded (see 3.4).

Test B — transfer-line sign. On the line 限定性净资产转为非限定性净资产, value moves into unrestricted net assets and out of restricted, so the two columns must be exact negations with known signs.

Check Rate n
akbe1366 = −akbe1367 exactly 91.9% 4,111
akbe1366 ≥ 0 (consistent with unrestricted) 89.8%  
akbe1367 ≤ 0 (consistent with restricted) 89.2%  

This fixes the orientation that Test A leaves open.

Test C — accounting identity. If each pair is a complete statement of activities for its period, then within that pair net asset change must equal total income less total expenses.

Period Identity holds n
Current-period columns 72.4% 27,190
Prior-period columns 69.6% 25,049

Both pairs satisfy the identity independently and at similar rates, which is what a genuine period decomposition predicts and what an arbitrary column grouping would not produce.

Conclusion. akbe1364 = prior-period unrestricted, akbe1365 = prior-period restricted, akbe1366 = current unrestricted, akbe1367 = current restricted. Three tests, three different failure modes, one consistent answer.

3.3 An undocumented historical column family

The 2015–2017 financial-activities tables carry three restriction variants per base variable — fa_dine, fa_dinu, and fa_dins — but the original codebook documents only e (限定) and u (非限定). The s variant is undocumented.

Test. If s is the 合计 total, then s == e + u must hold.

Result Value
Agreement across all 27 base/period combinations in the 2016 table 99.97%
Non-zero rows tested 6,193
Combinations at exactly 100% 25 of 27

fa_*s_* is mapped as the period total. The mapping is accepted only where both documented variants of the same base exist, so the rule cannot over-generalize.

3.4 Restatements are preserved, not averaged away

The 13.6% of cross-filing comparisons that disagree are real: a later filing restated an earlier year. The panel prefers a year’s own filing, so those disagreements would otherwise vanish. They are written to panel_restatements.parquet instead.

Component Measure Restatements
Financial activities unrestricted 14,298
Financial activities restricted 4,549
Financial position closing balance 8,111
Total   26,958

Median absolute restatement: ¥78,370. This table is a research object in its own right — it identifies exactly which organization-years were revised and by how much.

Step 4 — Assemble the panel

organization_panel.parquet is built from the crosswalk and nothing else. Its grain is one measurement per (unified_credit_code, data_year, component, variable, measure, source row). The deprecated foundation_panel.parquet compatibility alias has identical bytes; it is not independently built.

Historical rows that carry only an RICF identifier are linked to a credit code through the reviewed legacy_id_links.parquet, extended by a prepass over any historical table that carries both identifiers. An identifier resolving to more than one credit code is ambiguous evidence and is excluded rather than resolved. In this snapshot, zero identifiers were ambiguous; 4,921 legacy rows carry no resolvable identity and are counted in the summary.

Observation kinds, and recovering 2018

Chinese annual reports restate the prior period alongside the current one, and balance sheets carry an opening balance. Both are evidence about a neighbouring year. The panel records how each value was observed:

Kind Meaning Rows
reported Filed for this data year 8,658,689
prior_period From the prior-period column of a later filing 863,029
opening_balance Opening balance of the next filing, equal to this year’s close 251,717

A prior_period or opening_balance row is emitted only for a year with no filing of its own. Where the year has a filing, that filing is the better evidence and the disagreement, if any, goes to the restatement table. This rule suppressed 668,602 redundant opening balances and 287,291 redundant prior-period rows, and it is what prevents a naive groupby(year).sum() from double counting.

The practical consequence is that 2018 is no longer empty. The 2019 filings supply 162,310 observations for 2018 — clearly labelled, never presented as filed. 2018 still has no filing of its own, and the site’s coverage matrix continues to show that gap.

Coverage results

Historical source columns

Layer Tables Columns mapped
Published legacy 2013–2016 17 821 / 890
Companion 2017 4 159 / 164
Total 21 980 / 1,054 (93.0%)

The denominator counts every source data column except the identity keys, which carry no measurement, and the import’s own provenance columns. Mapped means the column appears as a source_field in the panel.

The unmapped remainder is identity, provenance, and six genuinely ambiguous columns listed in Limitations.

The HTML layer uses the same evidence rule at context rather than spreadsheet column grain. The crosswalk records 594 form label/header mappings and 317 Shanghai table-position mappings. Their privacy-safe financial projection contributes 3,751,523 rows; all other parsed form controls and table cells remain source-native internal evidence.

Variables with a current-source mapping

Component Mapped Total Unmapped are
Financial position 33 45 12 platform-computed subtotals
Financial activities 32 42 10 subtotals and totals
Cash flow 24 35 11 subtotals and net-flow lines

The panel

Dimension Value
Retained source observations 11,159,082
Organizations 16,712
Variables 198
Data years 2004–2025, continuous
From published legacy 2,302,184
From historical HTML 3,751,523
From current MCA 3,719,728
Foundation / non-foundation observations 9,662,534 / 1,496,548
File size (Parquet + Zstandard) 82 MiB

Observations by data year:

2005 2006 2007 2008 2009 2010 2011
19,362 34,228 65,070 134,575 175,984 215,941 225,361
2012 2013 2014 2015 2016 2017 2018
248,649 1,207,900 1,245,333 1,277,013 728,600 475,691 162,310
2019 2020 2021 2022 2023 2024 2025
346,423 351,057 341,029 465,158 659,645 706,008 688,098

Years before 2013 are explicitly marked sparse_early_year; they are weighted toward Beijing and chinanpo and must not be read as a national trend.

Presenting the result

A panel that shares a variable is only half the work. The profile page also has to show both eras as the same thing, and for a while it did not: it rendered historical letter codes on one side of the 2018 gap and anonymous akbe* columns on the other, which invited a reader to conclude the two were incomparable.

The crosswalk already contained everything needed to fix that, so nothing new was inferred. ricf build-pages-data publishes the relevant subset as site/data/panel/panel_fields.json.gz (about 12 KiB compressed, 990 historical column mappings and 79 current label mappings in this snapshot), and the site reads it rather than deciding anything itself. Three consequences:

One row identity in both eras. A current filing row is resolved to its letter code through current.labels, keyed by the normalized official line label with the documented canonical foldings as a fallback. Every observed label in this snapshot resolves, which is the same fact as unmapped_current_labels being empty for all three statement components. A row the crosswalk does not cover keeps its source-native field codes in a separate block; programs and public_fundraising have no crosswalk entry, so they stay wholly source-native.

The measure becomes a column, not a code suffix. historical_columns maps each published historical column to {variable, measure, year_offset}, so fa_dine_last resolves to fa_din / restricted / prior period. A 2016 financial-activities panel that used to render as 60 rows of suffixed codes now renders as 15 rows with restricted and unrestricted columns, matching the shape the current form already uses. The exact column names stay listed under each table, so nothing becomes unverifiable.

Prior-period columns are separated, not merged. A year_offset of -1 reports the neighbouring year, so those values are shown in their own table labeled with the year they describe. This is the display counterpart of the panel rule in Step 4: a carried value is never presented as a filed one.

Row order and section grouping come from the codebook sheets, which list fields in official statement order and separate blocks with unlettered heading rows. That is presentation only — it orders variables the crosswalk already carries. The financial-position sheet originally left its six heading rows untranslated; an English gloss was added for each so the grouping reads in either language.

Validation

Steps 2 and 3 establish that the mapping is structurally consistent. That is not the same as showing the two eras measure the same quantity. Three independent checks address that.

V1 — Cross-era rank agreement

If a historical variable and its current counterpart measure the same thing, then for a given organization the last historical value and the first current value should be strongly related, despite the 2–5 year gap between them. If the mapping were wrong, or the units differed, they would not be.

The control is a permutation: the same two value distributions, paired to the wrong organizations, with a fixed seed so the result is reproducible. The script streams the panel and applies the canonical filed-value filter before retaining only each organization’s last published-legacy and first current-MCA observation.

python scripts/validate_harmonization.py --snapshot 2026-07-29
Variable Measure n Spearman ρ Permuted ρ Median |log₁₀ ratio| Within 2× Permuted within 2×
fp_csh cash closing 2,612 0.732 0.029 0.212 57.8% 23.7%
fa_din donation income unrestricted 1,961 0.487 0.004 0.523 33.1% 16.7%
fa_acco activity cost unrestricted 2,261 0.683 −0.017 0.306 49.4% 21.2%
cf_doc donations received amount 2,057 0.631 −0.014 0.405 39.7% 17.9%
cf_stpc staff payments amount 1,249 0.744 0.065 0.223 59.3% 25.7%

Every variable shows strong rank agreement across the era boundary against a permuted control at essentially zero. The median ratio near 10⁰ also rules out a unit mismatch: had one era reported in 万元 and the other in 元, the ratios would cluster at 10⁴.

Agreement is not perfect and should not be. Two to five years separate the paired observations, and real change intervenes — which is precisely why the flow variables (fa_din, cf_doc, volatile year to year) agree less closely than the stock and structural variables (fp_csh, cf_stpc).

V2 — Internal series continuity

For any organization with consecutive current filings, the panel’s begin value for year T must equal its end value for T−1. This is the same identity as test 3.1, now checked on the assembled output rather than the source, confirming the builder did not scramble the mapping. It holds by construction for the 92.24% of source rows that satisfy it, and the residual is exactly the restatement set.

A worked example — cash for 532200006687540807, a foundation with evidence in both eras:

Data year Measure Value Kind Source
2014 end 14,479,831.21 reported 2014/2014_financial_position.tsv
2019 end 14,410,444.13 opening balance annual_financial_position
2020 begin 14,410,444.13 reported annual_financial_position
2020 end 13,491,548.37 reported annual_financial_position
2021 begin 13,491,548.37 reported annual_financial_position
       
2024 end 11,651,157.38 reported annual_financial_position
2025 begin 11,651,157.38 reported annual_financial_position
2025 end 10,868,245.21 reported annual_financial_position

Each year’s opening equals the previous year’s close, the 2019 point is correctly attributed to the next filing, and a single series now spans a decade that previously broke into two unconnected blocks.

V3 — Regression tests

The mapping logic is covered by regression tests in tests/test_crosswalk.py and tests/test_panel.py, including:

Limitations

Duplicate evidence is ranked rather than deleted. An observation key is (unified_credit_code, data_year, variable, measure, observation_kind). Measured on the completed 2026-07-29 panel:

Metric Value
Duplicate observation keys / retained rows 828,347 / 1,875,071
Cross-source / within-table keys 594,380 / 233,967
Agree / rounding / substantive conflicts 828,171 / 76 / 6,303
Identity-conflict keys / quarantined rows 718 / 1,525

Every duplicate key receives contiguous ranks and exactly one rank 1. The 594,380 cross-source keys use the tracked lineage policy: merged historical lineages outrank either single-source lineage, and chinanpo outranks CFC where the two single-source lineages meet. Within-table keys prefer direct identity, then completeness and source-modification evidence, with deterministic stable source order as the final fallback. Twenty-one identity-conflict contexts are quarantined rather than silently resolved.

The independent evidence does not make the policy infallible. A 2016 prior-period reading of 2015 matches CFC within one cent in 97.6735% of 37,266 comparisons and chinanpo in 97.4749%; CFC is strictly closer 103 times and chinanpo 19 times. The uniform maintainer decision remains chinanpo-first, and the slightly contrary evidence is retained rather than turned into hidden per-variable overrides.

That comparison is taken over all comparable observations, where the two lineages are near-indistinguishable. Restricting it to the keys where the candidates disagree with each other tells a much sharper story, and scripts/audit_duplicate_precedence.py reproduces it from the snapshot. Of the 6,250 conflicting organization-year-variable groups, 1,109 can be compared against the same figure observed with a different observation_kind; excluding the 1,522 arbiter observations that share a lineage with one of the candidates — where agreement is plumbing rather than evidence — leaves 326 to 373 decidable cases depending on rounding tolerance. On those the rank-1 value is the one the independent observation backs in 10.2% to 11.3% of cases, by class 5.9% for one_side_zero and 15.3% to 20.3% for non_zero_ratio, stable from a 0.01% to a 2% tolerance.

The policy is unchanged on this evidence and deliberately so: the arbiter is a third opinion rather than ground truth, since a later filing can legitimately restate a year, and it reaches only about 5% of the conflicting keys. Nor does the cheap correction hold up — only 56.1% of the 3,309 one_side_zero conflicts publish the zero today, so the failure is not simply a preference for zero. The honest reading is narrower and still useful: on a conflicting key the rank is a deterministic tie-break, not a finding about which figure is true, and it is measurably not neutral. Analysis that cannot tolerate that should filter values_disagree == false rather than rely on rank 1.

Two ways of pushing that 5% coverage higher were tested and both are closed. A third source lineage cannot serve as a witness, because panel_duplicate_observations.parquet holds every row for a duplicated key: a lineage that observed the key is a candidate, not an outside observer. And an accounting identity is circular. fp_ntt == fp_unls + fp_lims holds for 93.99% of unconflicted triples and reaches 439 conflicting keys, where it appears to contradict the rank-1 choice in 99.3% of cases — but all three terms for an organization-year come from one filing, and that filing is a disputed source, so the identity only re-states a candidate’s internal consistency. The independent count is zero. The script records that route and its zero explicitly, because a circular arbiter produces a more emphatic number than a sound one and is easy to mistake for a result.

The sector chart and profile trend now select rank 1 and reject identity-conflicted rows. Researchers should do the same. Counts in this document are retained-source-row counts; canonical values are a filtered view, not a destructive rewrite.

The current basic-information table is outside the panel. It is published in full on the profile, and every one of its 76 published columns now resolves to a codebook variable, so the component runs 2013–2025 rather than stopping where the published legacy tables do. Its counts and amounts carry a number as well as the filed text; its contact and identity fields stay text. 41 codebook basic variables have no counterpart in the current form at all, so those series still end in 2017.

Six historical columns remain unmapped because the evidence is genuinely ambiguous:

Column File Why
cf_inpc.1 2013 cash flow A duplicated column in the source table
GIS编码源地址 2015 basic profile No codebook equivalent
党员数 2015 basic profile Could be ba_pm (all party members) or ba_bpm (board/supervisors); no evidence distinguishes them
ba_evi, ba_seceml, cfc_oid 2016 basic Undocumented codes

Six current-only lines (税金及附加, 商品销售成本, 提供服务成本, 政府补助支出, 其他支出, 其中:捐赠项目成本) exist in the current expense breakdown but not in the historical form. They receive deterministic fa_ord<N> codes from their form ordinal and keep the platform’s own Chinese label. No historical equivalent is claimed for them.

Platform-computed subtotals have no current values. Variables like fp_asto (总资产) and fp_ntt (净资产合计) are present historically but the current form does not store them. They can be derived from their components; the panel does not derive them, because a derived total would be indistinguishable from a reported one.

Panel coverage is 16,712 organizations and is jurisdictionally uneven. The 1,302,542 rows before 2013 carry sparse_early_year; Beijing, Shanghai, and chinanpo are denser than other jurisdictions. Public profile shards remain foundation-scoped under O2, so non-foundations are available in the canonical bulk panel and registry search but do not receive thin longitudinal profiles.

Pass rates are not error rates. A test at 85.42% does not mean 14.58% of the mapping is wrong; it means 14.58% of comparisons differ for a substantive reason, principally restatement, which is separately recorded.

Reproducing this analysis

SNAPSHOT=2026-07-29

# Regenerate the crosswalk, re-running every identity test against the snapshot
python scripts/build_variable_crosswalk.py --snapshot "$SNAPSHOT"

# Fail instead of writing if the tracked crosswalk is stale
python scripts/build_variable_crosswalk.py --snapshot "$SNAPSHOT" --check

# Rebuild the panel and the restatement table
ricf --snapshot "$SNAPSHOT" build-panel

# Regenerate the codebook's Panel Crosswalk sheet
python scripts/update_codebook.py --snapshot "$SNAPSHOT"
ricf --snapshot "$SNAPSHOT" inventory

# Reproduce the cross-era validation table
python scripts/validate_harmonization.py --snapshot "$SNAPSHOT"

# Run the mapping regression tests
python -m pytest tests/test_crosswalk.py tests/test_panel.py -v

Where things live

Artifact Path
The mapping metadata/variable_crosswalk.json
Matching logic and reviewed synonyms src/ricf/crosswalk.py
Identity tests src/ricf/evidence.py
Cross-era validation scripts/validate_harmonization.py
Panel builder src/ricf/panel.py
Human-readable mapping RICF Codebook.xlsx, sheet Panel Crosswalk 面板对照
Panel table data/processed/<snapshot>/organization_panel.parquet
Deprecated compatibility alias data/processed/<snapshot>/foundation_panel.parquet
Restatements data/processed/<snapshot>/panel_restatements.parquet

Using the panel

import pandas as pd

panel = pd.read_parquet("organization_panel.parquet")

# Canonical filed values — the safe default for most analyses
filed = panel[
    (panel.observation_kind == "reported")
    & (panel.precedence_rank == 1)
    & ~panel.identity_conflict
]

# Include neighbouring-filing evidence to bridge the 2018 gap
bridged = panel[
    panel.observation_kind.isin(["reported", "prior_period"])
    & (panel.precedence_rank == 1)
    & ~panel.identity_conflict
]

# One organization's cash series across both eras
series = panel[
    (panel.variable == "fp_csh")
    & (panel.measure == "end")
    & (panel.unified_credit_code == "532200006687540807")
    & (panel.precedence_rank == 1)
    & ~panel.identity_conflict
].sort_values("data_year")

# Restrict to the strongest evidence
strict = panel[panel.mapping_evidence.isin(["codebook_exact", "codebook_normalized"])]

Cite the snapshot identifier alongside the data paper. Values in a snapshot are immutable; live platform totals are not.