Harmonizing the historical and current codebooks
How RICF’s 2005–2017 historical tables and HTML filings and 2019–2025 Ministry of Civil Affairs filings were reconciled into one panel, what evidence justified each mapping, and how the result was validated.
Snapshot 2026-07-29. All figures in this document are reproducible from the
commands in Reproducing this analysis.
The current build first expanded collection to report-form types 0 and 1,
raising foundation filings from 23,883 to 31,521. Stage 3f then integrated the
evidence-mapped financial projection of the historical HTML corpora. The
canonical result is 11,159,082 retained source observations across 16,712
organizations and 198 variables. Every identity test below was re-run against
the expanded snapshot. The six ambiguous published-table columns remain
unmapped, and HTML contexts without explicit label/header evidence remain
source_native_unmapped.
Contents
- The problem
- Design principles
- Step 1 — Inventory the vocabularies
- Step 2 — Match labels
- Step 3 — Establish unlabeled column meanings by test
- Step 4 — Assemble the panel
- Coverage results
- Presenting the result
- Validation
- Limitations
- Reproducing this analysis
The problem
RICF holds two bodies of evidence that had never shared a variable.
The published historical tables use three different header conventions across four years, and the financial tables add period and restriction suffixes on top:
| Year | Basic profile | Financial statements |
|---|---|---|
| 2013 | RICF letter codes (ba_cn, ba_rdt) |
Chinese labels for position; letter codes for activities and cash flow |
| 2014 | Chinese labels (机构名称, 数据年度) |
Chinese labels throughout |
| 2015 | Chinese labels | Both: 2015_financial_activities.tsv uses Chinese labels, fa_2015.tsv uses fa_dine_last-style codes |
| 2016 | RICF letter codes | Letter codes with _begin/_end and _last/_this suffixes |
| 2017 (companion) | Letter codes | Letter codes with no suffix |
The current MCA tables are long-format. Each row is one line of the official form, and the columns are opaque platform identifiers:
annual_financial_position 920,320 rows
akbe1347 货币资金 / 存货 / 长期股权投资 … ← the official line label
akbe1348 1 / 8 / 21 … ← the line's ordinal in the form
akbe1349 9358464.34 ← unlabeled value column
akbe1350 19534467.34 ← unlabeled value column
The project had treated the whole akbe* family as
source_native_unmapped. That was correct for genuinely opaque fields, but it
obscured a fact worth acting on: the financial tables are self-describing.
akbe1347 holds the form’s own Chinese label, which is the same vocabulary the
historical codebook indexes. What remained unknown was only the meaning of the
unlabeled value columns.
The consequence was visible on every foundation profile. The coverage matrix
showed historical H rows for 2013–2016 and current M rows for 2020–2024
that never shared a variable, and the only trend the site could draw for
current years was labelled akbe1349 / akbe1350.
The same break showed up in the profile’s tables even after the panel existed,
because the tables did not read the crosswalk. A historical financial-position
row was named fp_csh; the current row carrying the same measurement was a
Chinese label under four anonymous akbe* headings. Publishing the crosswalk’s
vocabulary to the browser removed that asymmetry; see
Presenting the result.
Design principles
Four constraints shaped the method, all from
AGENTS.md:
- No guessing. A source field receives a human-readable meaning only when evidence supports it. Every mapping records the grade of evidence that justified it, so a reviewer can audit the weakest links first.
- No fuzzy matching. String similarity is never used. Mechanical normalization is allowed and documented; a genuine wording difference needs an explicit, hand-reviewed entry with its own justification.
- Nothing is discarded. Conflicting and duplicate evidence stays available. The panel is a derived layer; the published TSVs and the source-native MCA tables are untouched.
- The mapping is separable from the data. The crosswalk is a tracked JSON document reviewable on its own, and the panel builder reads nothing else.
Evidence grades
| Grade | Definition | Panel rows |
|---|---|---|
codebook_exact |
Source label and codebook field identical after Unicode NFKC normalization and whitespace removal | 4,903,450 |
codebook_normalized |
Matched after documented mechanical folding only: list numbering (四、), item numbering ((一)), accounting prefixes (减:, 其中:), separator punctuation |
585,159 |
reviewed_synonym |
Wording genuinely differs; each alias listed explicitly with the evidence for it | 343,011 |
verified_identity |
An unlabeled column’s meaning established by an empirical test, with the pass rate recorded | 120,848 |
current_only_variable |
The line exists only in the current form; the code is assigned from the form ordinal and no historical equivalent is claimed | 69,444 |
source_native_unmapped |
No evidence supports a mapping; the field is retained verbatim | — |
Step 1 — Inventory the vocabularies
The historical vocabulary comes from the four field sheets of
RICF Codebook.xlsx, which give letter code, Chinese label, and English label
for 197 variables. The current vocabulary is the distinct set of values in each
long table’s label column, read from the snapshot.
| Component | Codebook variables | Distinct current labels |
|---|---|---|
| Financial position | 45 | 35 |
| Financial activities | 42 | 20 |
| Cash flow | 35 | 24 |
| Basic profile | 119 | 78 |
The current form holds fewer lines than the codebook because subtotals
(流动资产合计, 资产合计, 收入合计) are computed by the platform rather
than stored. Those variables remain in the crosswalk with no current mapping;
researchers can derive them from their components.
One structural difference had to be resolved before any matching. The
historical financial-activities codebook splits every income and expense line
into two variables by restriction — fa_dine (捐赠收入_限定) and fa_dinu
(捐赠收入_非限定) — while the current form states the line once and splits the
restriction across value columns. The panel therefore keys financial-activities
rows on the base variable (fa_din) and carries the restriction in
measure, exactly as the balance sheet uses begin and end.
Step 2 — Match labels
Labels were matched through a three-rung ladder, stopping at the first rung that resolves.
Rung 1 — exact. NFKC-normalize, strip whitespace, compare. Resolves 31 of 35 financial-position labels immediately.
Rung 2 — mechanical normalization. Strip only decoration that carries no meaning: leading list numbering, parenthesized item numbering, accounting prefixes, and separator punctuation. This is deliberately conservative — it never changes a word.
四、汇率变动对现金的影响额 → 汇率变动对现金的影响额
减:累计折旧 → 累计折旧
(一)业务活动成本 → 业务活动成本
其中:捐赠收入 → 捐赠收入
Rung 3 — reviewed synonyms. Twenty-three wording differences remained. Each was resolved by hand and recorded with the evidence that justifies it — never by similarity scoring. The evidence is that both labels occupy the same slot in the same official form, verified against the snapshot’s ordinal column.
| Component | Aliases | Example |
|---|---|---|
| Financial position | 3 | 应收账款 ↔ 应收款项, both at form ordinal 3 |
| Financial activities | 1 | 限定性净资产转为非限定性资产 ↔ …非限定性净资产, both at ordinal 40 |
| Cash flow | 9 | 收取会费收到的现金 ↔ 收取会员费收到的现金, both the second operating inflow |
| Basic profile | 29 | 住所 ↔ 基金会地址, current form heading for the registered address; 志愿者数量 ↔ 志愿者数, 2014/2015 published wording |
The cash-flow case is the clearest demonstration that this is not fuzzy matching. Every current cash-flow row maps to the codebook in strictly increasing order, with only the platform’s unstored subtotals skipped:
| Ordinal | Current label | Codebook label | Code | Difference |
|---|---|---|---|---|
| 2 | 收取会费收到的现金 | 收取会员费收到的现金 | cf_soc |
会费 / 会员费 |
| 5 | 政府补助收到的现金 | 政府补贴收到的现金 | cf_govc |
补助 / 补贴 |
| 8 | 收到的其他与业务活动有关的现金 | 收到的与其他业务活动有关的现金 | cf_otca |
word order |
| 16 | 购买商品、接受服务支付的现金 | 购买商品接受服务支付的现金 | cf_psc |
punctuation |
| 30 | 收到的其他与投资活动有关的现金 | 收到的其他与投资有关的现金 | cf_oc |
活动 inserted |
| 52 | 偿付利息所支付的现金 | 偿还利息所支付的现金 | cf_inpc |
偿付 / 偿还 |
| 60 | 四、汇率变动对现金的影响额 | 汇率变动对现金的影响 | cf_exch |
numbering, trailing 额 |
| 61 | 五、现金及现金等价物净增加额 | 现金及现金等价物净增加额 | cf_ninc |
numbering |
That ordering isomorphism is the evidence. A similarity score would have produced the same answers with none of the justification.
Result: every observed current label in all three financial components resolves. Zero remain unmapped.
Step 3 — Establish unlabeled column meanings by test
Matching labels identifies which line a row describes. It says nothing about what the numeric columns beside it mean. Those were resolved empirically, and each test records its pass rate so a reviewer audits a number rather than an assertion.
3.1 Balance sheet: which column is the opening balance?
annual_financial_position has two unlabeled value columns.
Test. A balance sheet’s opening balance for year T must equal the same organization’s closing balance in its year T−1 filing. Join every filing to its predecessor on organization and form ordinal, and compare.
| Reading | Agreement | n |
|---|---|---|
akbe1349 = opening, akbe1350 = closing |
92.2% | 104,486 |
| Columns reversed | 10.1% | |
| Columns unchanged year over year | 14.7% |
The accepted reading beats the best alternative by a factor of 9.4.
3.2 Financial activities: four unlabeled columns
annual_financial_activities has four. Three independent tests were run, each
capable of falsifying the others.
Test A — cross-filing. If one pair carries the prior reporting period, it must reproduce the same organization’s previous-year filing.
| Reading | Agreement | n |
|---|---|---|
akbe1364/akbe1365 = prior period, akbe1366/akbe1367 = current |
85.4% | 129,237 |
| Current columns carry the prior period | 1.1% | |
| Columns unchanged year over year | 2.1% | |
| Restriction columns swapped | 0.7% |
A separation of roughly 45× against every alternative. The 13.6% shortfall is not noise in the test — it is restatement, which is itself recorded (see 3.4).
Test B — transfer-line sign. On the line
限定性净资产转为非限定性净资产, value moves into unrestricted net assets
and out of restricted, so the two columns must be exact negations with known
signs.
| Check | Rate | n |
|---|---|---|
akbe1366 = −akbe1367 exactly |
91.9% | 4,111 |
akbe1366 ≥ 0 (consistent with unrestricted) |
89.8% | |
akbe1367 ≤ 0 (consistent with restricted) |
89.2% |
This fixes the orientation that Test A leaves open.
Test C — accounting identity. If each pair is a complete statement of activities for its period, then within that pair net asset change must equal total income less total expenses.
| Period | Identity holds | n |
|---|---|---|
| Current-period columns | 72.4% | 27,190 |
| Prior-period columns | 69.6% | 25,049 |
Both pairs satisfy the identity independently and at similar rates, which is what a genuine period decomposition predicts and what an arbitrary column grouping would not produce.
Conclusion. akbe1364 = prior-period unrestricted, akbe1365 =
prior-period restricted, akbe1366 = current unrestricted, akbe1367 =
current restricted. Three tests, three different failure modes, one consistent
answer.
3.3 An undocumented historical column family
The 2015–2017 financial-activities tables carry three restriction variants per
base variable — fa_dine, fa_dinu, and fa_dins — but the original
codebook documents only e (限定) and u (非限定). The s variant is
undocumented.
Test. If
sis the 合计 total, thens == e + umust hold.
| Result | Value |
|---|---|
| Agreement across all 27 base/period combinations in the 2016 table | 99.97% |
| Non-zero rows tested | 6,193 |
| Combinations at exactly 100% | 25 of 27 |
fa_*s_* is mapped as the period total. The mapping is accepted only where
both documented variants of the same base exist, so the rule cannot
over-generalize.
3.4 Restatements are preserved, not averaged away
The 13.6% of cross-filing comparisons that disagree are real: a later filing
restated an earlier year. The panel prefers a year’s own filing, so those
disagreements would otherwise vanish. They are written to
panel_restatements.parquet instead.
| Component | Measure | Restatements |
|---|---|---|
| Financial activities | unrestricted | 14,298 |
| Financial activities | restricted | 4,549 |
| Financial position | closing balance | 8,111 |
| Total | 26,958 |
Median absolute restatement: ¥78,370. This table is a research object in its own right — it identifies exactly which organization-years were revised and by how much.
Step 4 — Assemble the panel
organization_panel.parquet is built from the crosswalk and nothing else. Its
grain is one measurement per (unified_credit_code, data_year, component,
variable, measure, source row). The deprecated foundation_panel.parquet
compatibility alias has identical bytes; it is not independently built.
Historical rows that carry only an RICF identifier are linked to a credit code
through the reviewed legacy_id_links.parquet, extended by a prepass over any
historical table that carries both identifiers. An identifier resolving to more
than one credit code is ambiguous evidence and is excluded rather than
resolved. In this snapshot, zero identifiers were ambiguous; 4,921 legacy rows
carry no resolvable identity and are counted in the summary.
Observation kinds, and recovering 2018
Chinese annual reports restate the prior period alongside the current one, and balance sheets carry an opening balance. Both are evidence about a neighbouring year. The panel records how each value was observed:
| Kind | Meaning | Rows |
|---|---|---|
reported |
Filed for this data year | 8,658,689 |
prior_period |
From the prior-period column of a later filing | 863,029 |
opening_balance |
Opening balance of the next filing, equal to this year’s close | 251,717 |
A prior_period or opening_balance row is emitted only for a year with no
filing of its own. Where the year has a filing, that filing is the better
evidence and the disagreement, if any, goes to the restatement table. This
rule suppressed 668,602 redundant opening balances and 287,291 redundant
prior-period rows, and it is what prevents a naive groupby(year).sum() from
double counting.
The practical consequence is that 2018 is no longer empty. The 2019 filings supply 162,310 observations for 2018 — clearly labelled, never presented as filed. 2018 still has no filing of its own, and the site’s coverage matrix continues to show that gap.
Coverage results
Historical source columns
| Layer | Tables | Columns mapped |
|---|---|---|
| Published legacy 2013–2016 | 17 | 821 / 890 |
| Companion 2017 | 4 | 159 / 164 |
| Total | 21 | 980 / 1,054 (93.0%) |
The denominator counts every source data column except the identity keys, which
carry no measurement, and the import’s own provenance columns. Mapped means the
column appears as a source_field in the panel.
The unmapped remainder is identity, provenance, and six genuinely ambiguous columns listed in Limitations.
The HTML layer uses the same evidence rule at context rather than spreadsheet column grain. The crosswalk records 594 form label/header mappings and 317 Shanghai table-position mappings. Their privacy-safe financial projection contributes 3,751,523 rows; all other parsed form controls and table cells remain source-native internal evidence.
Variables with a current-source mapping
| Component | Mapped | Total | Unmapped are |
|---|---|---|---|
| Financial position | 33 | 45 | 12 platform-computed subtotals |
| Financial activities | 32 | 42 | 10 subtotals and totals |
| Cash flow | 24 | 35 | 11 subtotals and net-flow lines |
The panel
| Dimension | Value |
|---|---|
| Retained source observations | 11,159,082 |
| Organizations | 16,712 |
| Variables | 198 |
| Data years | 2004–2025, continuous |
| From published legacy | 2,302,184 |
| From historical HTML | 3,751,523 |
| From current MCA | 3,719,728 |
| Foundation / non-foundation observations | 9,662,534 / 1,496,548 |
| File size (Parquet + Zstandard) | 82 MiB |
Observations by data year:
| 2005 | 2006 | 2007 | 2008 | 2009 | 2010 | 2011 |
|---|---|---|---|---|---|---|
| 19,362 | 34,228 | 65,070 | 134,575 | 175,984 | 215,941 | 225,361 |
| 2012 | 2013 | 2014 | 2015 | 2016 | 2017 | 2018 |
|---|---|---|---|---|---|---|
| 248,649 | 1,207,900 | 1,245,333 | 1,277,013 | 728,600 | 475,691 | 162,310 |
| 2019 | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 |
|---|---|---|---|---|---|---|
| 346,423 | 351,057 | 341,029 | 465,158 | 659,645 | 706,008 | 688,098 |
Years before 2013 are explicitly marked sparse_early_year; they are weighted
toward Beijing and chinanpo and must not be read as a national trend.
Presenting the result
A panel that shares a variable is only half the work. The profile page also has
to show both eras as the same thing, and for a while it did not: it rendered
historical letter codes on one side of the 2018 gap and anonymous akbe*
columns on the other, which invited a reader to conclude the two were
incomparable.
The crosswalk already contained everything needed to fix that, so nothing new
was inferred. ricf build-pages-data publishes the relevant subset as
site/data/panel/panel_fields.json.gz (about 12 KiB compressed, 990 historical
column mappings and 79 current label mappings in this snapshot), and the site
reads it rather than deciding anything itself. Three consequences:
One row identity in both eras. A current filing row is resolved to its
letter code through current.labels, keyed by the normalized official line
label with the documented canonical foldings as a fallback. Every observed
label in this snapshot resolves, which is the same fact as
unmapped_current_labels being empty for all three statement components. A row
the crosswalk does not cover keeps its source-native field codes in a separate
block; programs and public_fundraising have no crosswalk entry, so they stay
wholly source-native.
The measure becomes a column, not a code suffix. historical_columns maps
each published historical column to {variable, measure, year_offset}, so
fa_dine_last resolves to fa_din / restricted / prior period. A 2016
financial-activities panel that used to render as 60 rows of suffixed codes now
renders as 15 rows with restricted and unrestricted columns, matching the shape
the current form already uses. The exact column names stay listed under each
table, so nothing becomes unverifiable.
Prior-period columns are separated, not merged. A year_offset of -1
reports the neighbouring year, so those values are shown in their own table
labeled with the year they describe. This is the display counterpart of the
panel rule in Step 4: a carried value is never
presented as a filed one.
Row order and section grouping come from the codebook sheets, which list fields in official statement order and separate blocks with unlettered heading rows. That is presentation only — it orders variables the crosswalk already carries. The financial-position sheet originally left its six heading rows untranslated; an English gloss was added for each so the grouping reads in either language.
Validation
Steps 2 and 3 establish that the mapping is structurally consistent. That is not the same as showing the two eras measure the same quantity. Three independent checks address that.
V1 — Cross-era rank agreement
If a historical variable and its current counterpart measure the same thing, then for a given organization the last historical value and the first current value should be strongly related, despite the 2–5 year gap between them. If the mapping were wrong, or the units differed, they would not be.
The control is a permutation: the same two value distributions, paired to the wrong organizations, with a fixed seed so the result is reproducible. The script streams the panel and applies the canonical filed-value filter before retaining only each organization’s last published-legacy and first current-MCA observation.
python scripts/validate_harmonization.py --snapshot 2026-07-29
| Variable | Measure | n | Spearman ρ | Permuted ρ | Median |log₁₀ ratio| | Within 2× | Permuted within 2× |
|---|---|---|---|---|---|---|---|
fp_csh cash |
closing | 2,612 | 0.732 | 0.029 | 0.212 | 57.8% | 23.7% |
fa_din donation income |
unrestricted | 1,961 | 0.487 | 0.004 | 0.523 | 33.1% | 16.7% |
fa_acco activity cost |
unrestricted | 2,261 | 0.683 | −0.017 | 0.306 | 49.4% | 21.2% |
cf_doc donations received |
amount | 2,057 | 0.631 | −0.014 | 0.405 | 39.7% | 17.9% |
cf_stpc staff payments |
amount | 1,249 | 0.744 | 0.065 | 0.223 | 59.3% | 25.7% |
Every variable shows strong rank agreement across the era boundary against a permuted control at essentially zero. The median ratio near 10⁰ also rules out a unit mismatch: had one era reported in 万元 and the other in 元, the ratios would cluster at 10⁴.
Agreement is not perfect and should not be. Two to five years separate the
paired observations, and real change intervenes — which is precisely why the
flow variables (fa_din, cf_doc, volatile year to year) agree less closely
than the stock and structural variables (fp_csh, cf_stpc).
V2 — Internal series continuity
For any organization with consecutive current filings, the panel’s
begin value for year T must equal its end value for T−1. This is the
same identity as test 3.1, now checked on the assembled output rather than the
source, confirming the builder did not scramble the mapping. It holds by
construction for the 92.24% of source rows that satisfy it, and the residual is
exactly the restatement set.
A worked example — cash for 532200006687540807, a foundation with evidence
in both eras:
| Data year | Measure | Value | Kind | Source |
|---|---|---|---|---|
| 2014 | end | 14,479,831.21 | reported | 2014/2014_financial_position.tsv |
| 2019 | end | 14,410,444.13 | opening balance | annual_financial_position |
| 2020 | begin | 14,410,444.13 | reported | annual_financial_position |
| 2020 | end | 13,491,548.37 | reported | annual_financial_position |
| 2021 | begin | 13,491,548.37 | reported | annual_financial_position |
| … | ||||
| 2024 | end | 11,651,157.38 | reported | annual_financial_position |
| 2025 | begin | 11,651,157.38 | reported | annual_financial_position |
| 2025 | end | 10,868,245.21 | reported | annual_financial_position |
Each year’s opening equals the previous year’s close, the 2019 point is correctly attributed to the next filing, and a single series now spans a decade that previously broke into two unconnected blocks.
V3 — Regression tests
The mapping logic is covered by regression tests in
tests/test_crosswalk.py and
tests/test_panel.py, including:
- a near-miss that is not a reviewed synonym must stay unmapped, so the matcher cannot drift toward fuzziness;
- the undocumented
stotal variant is accepted only when both documented variants exist; - every recorded identity test must beat its alternatives by at least 2×;
- an unmapped source label must be excluded from the panel, not guessed;
- an opening balance must fill only a year with no filing of its own.
Limitations
Duplicate evidence is ranked rather than deleted. An observation key is
(unified_credit_code, data_year, variable, measure, observation_kind).
Measured on the completed 2026-07-29 panel:
| Metric | Value |
|---|---|
| Duplicate observation keys / retained rows | 828,347 / 1,875,071 |
| Cross-source / within-table keys | 594,380 / 233,967 |
| Agree / rounding / substantive conflicts | 828,171 / 76 / 6,303 |
| Identity-conflict keys / quarantined rows | 718 / 1,525 |
Every duplicate key receives contiguous ranks and exactly one rank 1. The 594,380 cross-source keys use the tracked lineage policy: merged historical lineages outrank either single-source lineage, and chinanpo outranks CFC where the two single-source lineages meet. Within-table keys prefer direct identity, then completeness and source-modification evidence, with deterministic stable source order as the final fallback. Twenty-one identity-conflict contexts are quarantined rather than silently resolved.
The independent evidence does not make the policy infallible. A 2016 prior-period reading of 2015 matches CFC within one cent in 97.6735% of 37,266 comparisons and chinanpo in 97.4749%; CFC is strictly closer 103 times and chinanpo 19 times. The uniform maintainer decision remains chinanpo-first, and the slightly contrary evidence is retained rather than turned into hidden per-variable overrides.
That comparison is taken over all comparable observations, where the two
lineages are near-indistinguishable. Restricting it to the keys where the
candidates disagree with each other tells a much sharper story, and
scripts/audit_duplicate_precedence.py reproduces it from the snapshot. Of the
6,250 conflicting organization-year-variable groups, 1,109 can be compared
against the same figure observed with a different observation_kind; excluding
the 1,522 arbiter observations that share a lineage with one of the candidates —
where agreement is plumbing rather than evidence — leaves 326 to 373 decidable
cases depending on rounding tolerance. On those the rank-1 value is the one the
independent observation backs in 10.2% to 11.3% of cases, by class 5.9% for
one_side_zero and 15.3% to 20.3% for non_zero_ratio, stable from a 0.01% to
a 2% tolerance.
The policy is unchanged on this evidence and deliberately so: the arbiter is a
third opinion rather than ground truth, since a later filing can legitimately
restate a year, and it reaches only about 5% of the conflicting keys. Nor does
the cheap correction hold up — only 56.1% of the 3,309 one_side_zero conflicts
publish the zero today, so the failure is not simply a preference for zero. The
honest reading is narrower and still useful: on a conflicting key the rank is a
deterministic tie-break, not a finding about which figure is true, and it is
measurably not neutral. Analysis that cannot tolerate that should filter
values_disagree == false rather than rely on rank 1.
Two ways of pushing that 5% coverage higher were tested and both are closed. A
third source lineage cannot serve as a witness, because
panel_duplicate_observations.parquet holds every row for a duplicated key: a
lineage that observed the key is a candidate, not an outside observer. And an
accounting identity is circular. fp_ntt == fp_unls + fp_lims holds for 93.99%
of unconflicted triples and reaches 439 conflicting keys, where it appears to
contradict the rank-1 choice in 99.3% of cases — but all three terms for an
organization-year come from one filing, and that filing is a disputed source, so
the identity only re-states a candidate’s internal consistency. The independent
count is zero. The script records that route and its zero explicitly, because a
circular arbiter produces a more emphatic number than a sound one and is easy to
mistake for a result.
The sector chart and profile trend now select rank 1 and reject identity-conflicted rows. Researchers should do the same. Counts in this document are retained-source-row counts; canonical values are a filtered view, not a destructive rewrite.
The current basic-information table is outside the panel. It is published in full on the profile, and every one of its 76 published columns now resolves to a codebook variable, so the component runs 2013–2025 rather than stopping where the published legacy tables do. Its counts and amounts carry a number as well as the filed text; its contact and identity fields stay text. 41 codebook basic variables have no counterpart in the current form at all, so those series still end in 2017.
Six historical columns remain unmapped because the evidence is genuinely ambiguous:
| Column | File | Why |
|---|---|---|
cf_inpc.1 |
2013 cash flow | A duplicated column in the source table |
GIS编码源地址 |
2015 basic profile | No codebook equivalent |
党员数 |
2015 basic profile | Could be ba_pm (all party members) or ba_bpm (board/supervisors); no evidence distinguishes them |
ba_evi, ba_seceml, cfc_oid |
2016 basic | Undocumented codes |
Six current-only lines (税金及附加, 商品销售成本, 提供服务成本,
政府补助支出, 其他支出, 其中:捐赠项目成本) exist in the current
expense breakdown but not in the historical form. They receive deterministic
fa_ord<N> codes from their form ordinal and keep the platform’s own Chinese
label. No historical equivalent is claimed for them.
Platform-computed subtotals have no current values. Variables like
fp_asto (总资产) and fp_ntt (净资产合计) are present historically but the
current form does not store them. They can be derived from their components;
the panel does not derive them, because a derived total would be
indistinguishable from a reported one.
Panel coverage is 16,712 organizations and is jurisdictionally uneven.
The 1,302,542 rows before 2013 carry sparse_early_year; Beijing, Shanghai,
and chinanpo are denser than other jurisdictions. Public profile shards remain
foundation-scoped under O2, so non-foundations are available in the canonical
bulk panel and registry search but do not receive thin longitudinal profiles.
Pass rates are not error rates. A test at 85.42% does not mean 14.58% of the mapping is wrong; it means 14.58% of comparisons differ for a substantive reason, principally restatement, which is separately recorded.
Reproducing this analysis
SNAPSHOT=2026-07-29
# Regenerate the crosswalk, re-running every identity test against the snapshot
python scripts/build_variable_crosswalk.py --snapshot "$SNAPSHOT"
# Fail instead of writing if the tracked crosswalk is stale
python scripts/build_variable_crosswalk.py --snapshot "$SNAPSHOT" --check
# Rebuild the panel and the restatement table
ricf --snapshot "$SNAPSHOT" build-panel
# Regenerate the codebook's Panel Crosswalk sheet
python scripts/update_codebook.py --snapshot "$SNAPSHOT"
ricf --snapshot "$SNAPSHOT" inventory
# Reproduce the cross-era validation table
python scripts/validate_harmonization.py --snapshot "$SNAPSHOT"
# Run the mapping regression tests
python -m pytest tests/test_crosswalk.py tests/test_panel.py -v
Where things live
| Artifact | Path |
|---|---|
| The mapping | metadata/variable_crosswalk.json |
| Matching logic and reviewed synonyms | src/ricf/crosswalk.py |
| Identity tests | src/ricf/evidence.py |
| Cross-era validation | scripts/validate_harmonization.py |
| Panel builder | src/ricf/panel.py |
| Human-readable mapping | RICF Codebook.xlsx, sheet Panel Crosswalk 面板对照 |
| Panel table | data/processed/<snapshot>/organization_panel.parquet |
| Deprecated compatibility alias | data/processed/<snapshot>/foundation_panel.parquet |
| Restatements | data/processed/<snapshot>/panel_restatements.parquet |
Using the panel
import pandas as pd
panel = pd.read_parquet("organization_panel.parquet")
# Canonical filed values — the safe default for most analyses
filed = panel[
(panel.observation_kind == "reported")
& (panel.precedence_rank == 1)
& ~panel.identity_conflict
]
# Include neighbouring-filing evidence to bridge the 2018 gap
bridged = panel[
panel.observation_kind.isin(["reported", "prior_period"])
& (panel.precedence_rank == 1)
& ~panel.identity_conflict
]
# One organization's cash series across both eras
series = panel[
(panel.variable == "fp_csh")
& (panel.measure == "end")
& (panel.unified_credit_code == "532200006687540807")
& (panel.precedence_rank == 1)
& ~panel.identity_conflict
].sort_values("data_year")
# Restrict to the strongest evidence
strict = panel[panel.mapping_evidence.isin(["codebook_exact", "codebook_normalized"])]
Cite the snapshot identifier alongside the data paper. Values in a snapshot are immutable; live platform totals are not.
Related documentation
- Data architecture — where the panel sits among the layers
- Known gaps and decisions — the open questions
- Provenance and merge policy — how source precedence is decided