CastBench · AI-governance benchmark

From public sources to a causal graph, evaluated out-of-time

How AI-policy bills, actors and impacts are collected from real APIs, assembled into a typed causal graph, split out-of-time, and scored — balanced accuracy, recall and F1 only, never AUC. Ends with two trained models: an actor-action forecaster (stage 6) and an actor-relevance classifier (stage 7).

2,170 region bills · 6 regions 2,170 causal graphs 3,186 impact samples split test ≥ 2023 every field source-cited 2 trained models · forecaster + relevance
1

Data sources

real public APIs & databases — no fabrication
OECD.AI Observatory
api.oecdai.org
1,979 non-US AI policies, 76 countries + content/motivation
Congress.gov
api.congress.gov
US federal AI bills, sponsors, outcomes
SEC EDGAR
efts.sec.gov (full-text)
10-K/10-Q/20-F disclosures of law impacts
UN Comtrade
comtradeapi.un.org
trade-flow outcomes of export controls
Federal Register
federalregister.gov
downstream agency actions (generated events)
USAspending · GSA · DoD · DOE
usaspending.gov + releases
federal AI contract awards (funding+)
GDPR Tracker · EC · DPAs
enforcementtracker + ec.europa.eu
fines / market-remedy orders
News · regulator · IR releases
web-mined via research agents
firm responses, bans, deals — URL per fact
2

Corpus & labels

bills · actors · sourced outcomes

Bill records

Content, motivation, antecedent events, outcome + passage-mechanism flags.

must_pass_vehicleexecutive_orderstrong_standalone

Actor cards

25 featurized firms & governments: disposition (openness / race-pull / internalization) + harm profile.

Nvidia · TSMC · MetaOpenAI · AnthropicUS_GOV · CHINA_GOV · EU_GOV

Gold impact labels

Verified actor outcomes coded (aspect, ±direction), every field cited.

131 pairs63 laws · 25 actors75− / 45+

Orientation + stance

Every bill labeled restrictive/enabling/neutral from real OECD text; support/oppose stances sourced.

1,979 orientation+ 20 stance
Field completeness — % of records where each field is populated

Master bills corpus

ai_bills_policies · n = 660
core (content · impact · actors · outcome · flags)100%
outcome_reason (routed by outcome_type)100%
↳ why_enacted 15% · why_failed 84% · why_vetoed 1% — mutually exclusive
antecedent_events90%
sponsor85%
cosponsor_party72%
source_quotes18%
generated_events12%
recorded_votes8%
fr_documents2%
enacted → why_enacted · vetoed → why_vetoed · failed → why_failed. Every bill carries exactly one → 100%.

International corpus

intl_ai_policies · train+test · n = 1,979
core (title · content · motivation · type · status · source · actors)100%
outcome / adoption (from OECD status)100%
↳ in_force 95% · proposed 5% · not_adopted 0% — already official laws
sectors97%
responsible_org90%
binding82%
tags81%
end_year (only if ended)49%
Empty ≠ missing. These are strategies without a set end date — not gaps to fill.

Gold impact labels

actor_bill_impacts · verified · n = 131
bill_id · actor · aspect · direction100%
event · date · grade100%
source (URL / API query)100%
magnitude ($ figure where stated)10%
Every gold row is fully specified & source-cited; only a numeric magnitude is optional.

The sparse fields are structural, not missing data: why_enacted exists only for enacted bills, generated_events / fr_documents / recorded_votes only for US bills with downstream federal activity, and magnitude only where a source states a figure. Nothing here is a blank waiting to be fabricated.

Train set — what was "unfillable", and what stage 6 resolved

✓ Complete from source

OECD record + derivation + taxonomy
  • title · content · motivation · description100%
  • type · status · outcome (adoption) · source100%
  • document_type · legal_force · level (new)100%
  • train/test split · effective year (new)100%
  • sectors · responsible_org · binding · tags76–97%

◐ Now grounded-filled — was 0

graph-fields · verbatim-quote-verified from OECD text
  • motivation · generated_actions~94%
  • impacted_actors (structured)~94% was "inferred-only" → now source-grounded per bill
  • generated_impacts (impact · domain · ±dir)~94% each carries a verbatim evidence_quote
  • independent per-actor GOLD285 · 589 grew 131 → 285 sourced (A+B) / 589 actor-attributed — still the scarce tier, search-gated, never fabricated

Update: the fields once marked "cannot fill" are now grounded-filled for ~94% of bills — extracted from each OECD record's own text with every quote verbatim-verified (stage 6), not fabricated. What stays limited is independent per-actor gold, which grew 131 → 285 sourced but remains gated on real documented events.

3

Typed causal graph → GNN

one graph per bill · every node/edge traced to a field + source
antecedentprior events
motivates
motivationwhy proposed
reason
billprovisions
proposal
outcome ①enacted / vetoed / failed
actorsponsor · firms · govs
supports / opposes
outcome ①mechanism + why
do(enact)
impactworld_impact
impacts
outcome ②per-actor (aspect, ±dir)
structural node edge / relation outcome ① passage outcome ② impact data source
A · Source → field → node mind map (which field from which source builds which node)
OECD.AI Observatory api.oecdai.org
englishName / titlebill
description · overviewmotivation
type · extentBindingoutcome ① (orientation)
gaiinCountryactor (country gov)
Congress.gov api.congress.gov · BILLSTATUS
title · crs_summarybill · motivation
sponsor · cosponsor_party · n_cosponsorsactor
legislative milestonesantecedent_event
status → outcome_type + must_pass / exec_order / strong_standaloneoutcome ①
Federal Register federalregister.gov
fr_documentsimpact
generated_eventsoutcome ② world
SEC EDGAR efts.sec.gov · 10-K/10-Q/20-F
filing disclosure (charge · revenue-at-risk)outcome ② actor (aspect, ±dir)
UN Comtrade comtradeapi.un.org
trade flows (HS 8542 / 8486)outcome ② actor (trade ±)
USAspending · GSA · DoD · DOE usaspending.gov + releases
contract award $ · ceilingoutcome ② actor (funding +)
GDPR Tracker · EC · DPAs enforcementtracker · ec.europa.eu
fine amount · remedy orderoutcome ② actor (compliance / market_access −)
News · regulator · IR (web agents) → actor_bill_responses / _impacts
stance (support / oppose) + sourcesupports/opposes edge
impact event + source urloutcome ② actor
why_enacted · world_impactimpact · motivation
actor_features (real events) Comtrade · WorldBank · LDA · BLS
disposition · harm · leverageactor (61-dim features)
B · Edge provenance — how each relation is created
EdgeFrom → ToCreated from
motivatesantecedent → motivationeach antecedent_events item
reason to proposemotivation → billalways
proposalbill → outcome ①always
decideactor → outcome ①sponsor / key_actors present
supports / opposesactor → outcome ①actor_bill_responses stance (+/−) · sourced
do(enact) / counterfactualoutcome ① → impactenacted flag (else "would-have")
impacts (dir, grade)impact → outcome ②aspect map / generated_events — dir & grade attached
bears impactstance-actor → outcome ② actoractor also carries a gold impact
C · GNN feature tensor — the relevant-actor GNN (advance-past-committee, leak-free)
Nodes: bill + antecedent + bill-relevant actors (content matched to the 51-actor roster). Per-node vector = type one-hot (3)process (7)actor disposition/harm (9)
  • bill node = n_cosponsors · bipartisan · sponsor party · chamber (majority-sponsor + sponsor-prior gated OFF — hurt 0.487; must_pass OFF — leak; text-embed OFF — hurt 0.599)
  • antecedent node = antecedent_events text length
  • relevant-actor node = disposition (race-pull · internalization · openness) + harm (DISEMP·LOSSC·POWER·SOCID·MISUSE) + is_gov, from actor_features.json
Edges: star — every antecedent & relevant-actor node → the bill node. Training: 12-seed ensemble · pool all pre-2023 congresses · label = advanced-past-committee.
Fields from: Congress.gov (title / sponsor / cosponsor_party / n_cosponsors) + actor_features roster  •  PENDING: sponsor effectiveness · leadership · seniority from the 60k-bill general crawl (`congress_general_crawl` + `augment_congress`) — being tested before wiring into the bill node.
model: 2-layer GCN (LayerNorm + dropout, mean readout → advance logit). Result ~0.75 leak-free.
worked example CHIPS and Science Act us-117-hr4346 · one real bill through every node
antecedent_event
US–China chip rivalry · 2020 export controls · pandemic shortage
← antecedent_events
motivation
revitalize US semiconductor manufacturing & secure supply chains
← background / content
bill
$52.7B semiconductor incentives + 25% investment tax credit
← title + content
outcome ① passage
ENACTED · signed Aug 2022 · bipartisan
← outcome_type + strong_standalone
impact
direct CHIPS awards to chipmakers → US fab build-out
outcome ② — per-actor, each cited
TSMCfunding +$6.6B· Commerce Samsungfunding +$4.74B· Commerce SK Hynixfunding +$458M· Senate Micronfunding +$6.1B· SEC Intelfunding +$7.9B· SEC 10-K TSMCcompliance −upside-share· Tax Foundation
4

Train / test splits

out-of-time, leakage-safe
TaskSplit ruleTrainTest
Passage (enacted/failed)2015–2022 → 2023–2025 · US-federal AI165399
International corpus · 95% already official lawsstartYear < 2024 → ≥ 2024 · 76 countries1,435544
Actor stanceby bill137
Impact (gold)zero-shot — scored whole131
Proposal-time features only. The international set is 95% already-adopted official laws (OECD status → in_force), so it trains actor-impact / coverage — not passage; US bills carry the enacted/vetoed/failed outcome. Bills with no documented firm impact keep the rule-derived orientation label, never an invented outcome.
5

Evaluation results

Precision · Recall · F1 — never AUC
F-score summary — our model vs LLM vs baseline (leak-free · fair matched candidate sets · pre-2023 LLM · no AUC)
Task — metricBaseline F1Our model F1LLM F1Winner
Actor relevance — set-F1 (fair 47-roster)0.210.310.22our model
Impact, actor-gold — set-F1 (9-aspect vocab)~0.240.340.43LLM
Advance-past-committee — F1 (advanced class)0.37our GNN
Passage, enacted — bal-acc (out-of-time)0.50~0.50~0.48chance ✗
F1 = unbudgeted set Precision·Recall·F1 (aspects / actors) or advanced-class F1 (GNN). Our model owns advance-past-committee; the LLM wins impact (higher recall, over-predicts). Actor-relevance is match-mode-dependent: our rule matcher wins on EXACT F1 (0.31 vs 0.22), the LLM wins on HIERARCHICAL F1 (0.32 vs 0.28, parent-agency / population→firm credit). Passage is unforecastable for everyone once de-leaked. Full P/R per metric below.
Passage — every method, LEAK-FREE out-of-time US-federal balanced accuracy · test 2023–2025 · only 10 positives · outcome leakage removed
↑ higher better
always-NO baseline0.500
our LogReg (tabular)0.500
our causal-graph GNN0.544
LLM · title-only (pre-2023)0.554
LLM · reads content (pre-2023)0.483
Honest negative result. Once outcome leakage is removed, no method beats chance (~0.48–0.55) — LLM, GNN and LogReg alike. Earlier LLM 0.79 / 0.85 read our world_impact field, which literally states the outcome ("Did not take effect / Had it become law…") for 95% of bills — pure leak, discarded. Passage-as-enacted isn't forecastable at 10 positives — but reframing the task fixes it ↓
Passage REFRAMED — "advance past committee" · our GNN works balanced accuracy · seed-ensemble · ~51 test positives · leak-free
↑ higher better
always-NO / chance0.500
process-feature LogReg (ceiling)0.590
relevant-actor GNN0.668
+ seed-ensemble + more data0.752
Our GNN: chance → 0.752, leak-free. Reframe "enacted" (unforecastable) → "advanced past committee"; add legislative-process features + bill-relevant actor nodes (matched to the 51-actor roster with real disposition/harm features); then ensemble over 12 seeds + pool all pre-2023 congresses (0.668→0.752). Plain process features cap at 0.590, so the graph + actor nodes carry the gain. Advanced-class P/R/F1 = 0.32 / 0.45 / 0.37. Text embeddings & a train-tuned threshold hurt (overfit 18 positives) and were dropped. Honest ceiling ~0.75 (balanced acc, high variance) — reaching 0.85 would need leaky features, so we stop here.
Actor stance support / oppose · sourced ground truth
accuracy
majority base0.55
all0.80
held-out test0.857
Actor-impact prediction aspect-level · vs raw generated-events 0.09
recall / F1
raw events (F1)0.09
M1 aspect+dir0.62
M1 aspect-only0.72
M2 change-detect F10.77
Actor-action forecast hybrid: persistence + LLM-novel
recall
persistence base0.49
hybrid · raw0.57
hybrid · distinct0.57
6

Multi-region expansion

6 regions · classified · graph-filled · strict-2023 split — new this session
New data sources — WebFetch-verified & snapshotted offline
OECD.AI raw cache
output/raw · 2,464 records
full API dump — the verbatim text every field is grounded on
Official harvest · r.jina.ai proxy
China Law Translate · DigiChina · e-Gov
JS / 403 government pages read through a reader proxy
Wikipedia · china-briefing · gov portals
WebFetch (research agents)
national / provincial / municipal reg discovery
Source-page snapshots
output/raw/harvest_pages · 76
raw page per source_url → verify every quote offline
New components

Region corpus · 6 buckets

EU · US · East Asia · China · Singapore · India — OECD base + verified harvest + native bills, deduped by title.

2,170 bills28 jurisdictions

Governance taxonomy

Every row classified — document_type (22-type) · jurisdiction_level · legal_force · province / city · issuing_authority.

policy_taxonomy.py

Grounded graph-fields

4 graph fields per bill (motivation · actions · impacted_actors · impacts); every evidence_quote verbatim-verified.

~94% filled

Region causal graphs

Typed graph per bill: motivation→bill→outcome + action / actor / impact nodes with signed direction.

2,170 graphs15.3k nodes · 15.6k edges

Unified impact dataset

Tiers A gold · A2 grounded-actor · B cited · C grounded — one temporal split, explicit counts.

3,186589 actor-attributed

Raw provenance & gate

snapshot script + README manifest; admission gate rejects unverifiable (Baidu-search) rows.

reproducible from cache
Source → field → node (multi-region additions)
OECD.AI record text description · overview
motivation.text (+quote)motivation
generated_actionsaction
impacted_actorsactor
generated_impacts (domain, ±dir)impact ②
taxonomy + status policy_taxonomy · adoption
document_type · legal_force · levelbill attrs
outcome (in_force / proposed / not_adopted)outcome ①
effective year → splittrain / test tag
Strict temporal split — every region in both (test = effective year ≥ 2023)
RegionTrain <2023Test ≥2023
EU713603
US131329
East Asia · JP·KR·TW·HK5452
China5357
Singapore6539
India3638
Total1,0521,118
0 leakage verified: nothing ≥2023 in train, nothing <2023 in test; undated rows kept in train so no data is dropped.
Evaluation on the region split — stratified, never AUC
Passage — US, the real proposal-stage task balanced accuracy · 7% pass · leak-free · out-of-time (test ≥ 2023)
↑ higher better
majority baseline0.500
trained LogReg0.500
our causal-graph GNN0.544
LLM · reads bill content (pre-2023)0.483
Honest negative result. On leak-free inputs no method beats chance (~0.48–0.55). The earlier 0.79/0.85 came from our world_impact text stating the outcome (95% of bills) — a leak, now removed. Whether an AI bill passes isn't recoverable from its text; it depends on the legislative process.
Impact — actor-gold test · our model vs LLM Metric 1 aspect recall · 210 test pairs · leak-free + memorization-audited
↑ higher better
majority baseline0.237
our trained model (actor-prior)0.243
LLM · pre-2023 cutoff0.561
The LLM is the signal; our trained model isn't. Our actor-prior model barely clears the baseline (0.243 vs 0.237), the LLM (pre-2023-cutoff gpt-3.5-0613, leak-free) more than doubles it. Full P/R/F1 (actor-gold): LLM 0.28/0.90/0.425 · our actor-prior 0.34/0.34/0.344 · baseline ~0.24 — the LLM over-predicts (high recall, low precision); our model is more precise. Fair: both choose from the same fixed 9-aspect vocab (a runtime guard asserts it). Genuine reasoning, not recall (pre-2023 cutoff). (Enrichment domains B: F1 0.28, near baseline.)
Actor relevance — which actors a bill affects title-only · leak-free · 75 bills · FAIR: both choose from the full 47-actor roster · M2 balanced accuracy
↑ higher better
frequency baseline0.593
our rule matcher0.672
LLM · pre-2023 cutoff0.747
Fair candidate pool (47-actor roster, with distractors); winner depends on match mode. EXACT set-F1: baseline 0.21 · our rule matcher 0.31 (wins) · LLM 0.22 — the LLM over-predicts into distractors so its exact F1 falls below our precise matcher. HIERARCHICAL F1 (parent-agency / population→firm credit, shared actor_hierarchy): baseline 0.19 · rule 0.28 · LLM 0.32 (wins) — broad predictions get expanded credit. M2 bal-acc 0.593 / 0.672 / 0.747. Title-only, no leak, pre-2023 cutoff, guarded identical candidate set.
6

Actor-action forecaster

predict what each actor DOES · temporal split (train <2024 / test ≥2024)

macro-F1 0.459

vs baseline 0.08 · gpt-3.5 0.24. 10 action types.

top-1 72%

per-actor F1 0.57.

10 / 10 classes

every action type carries signal (F1 > 0).

310 / 363

train / test actions, leak-free.

Per-action-type F1held-out
LOBBY0.88
RETALIATE0.80
LITIGATE0.50
BUILD0.43
RESTRICT0.40
PARTNER0.36
FUND0.35
REGULATE0.34
ENFORCE0.31
COMPLY0.22
Per-actor forecast — predicted top action across six policies
ActorSignatureChip export controls on ChinaCHIPS Act fab incentivesEU AI Act obligationsState deepfake / content lawAI compute infrastructureFederal AI safety rules
United StatesLOBBYRESTRICT 29%FUND 62%LOBBY 41%LOBBY 37%REGULATE 44%LOBBY 43%
ChinaLOBBYRETALIATE 41%FUND 51%LOBBY 50%LOBBY 45%REGULATE 38%LOBBY 55%
European UnionLOBBYREGULATE 30%FUND 45%LOBBY 44%LOBBY 39%REGULATE 39%LOBBY 48%
United KingdomFUNDRETALIATE 21%FUND 53%LOBBY 28%FUND 25%PARTNER 36%LOBBY 32%
NvidiaLOBBYRESTRICT 35%BUILD 45%LOBBY 73%LOBBY 33%PARTNER 71%LOBBY 77%
TsmcBUILDRESTRICT 70%BUILD 88%BUILD 49%BUILD 34%PARTNER 61%LOBBY 44%
IntelLOBBYRESTRICT 29%BUILD 72%LOBBY 40%LOBBY 47%PARTNER 85%LOBBY 58%
MicronLOBBYRESTRICT 60%BUILD 78%LOBBY 55%LOBBY 24%PARTNER 73%LOBBY 65%
SamsungLOBBYRESTRICT 52%BUILD 73%LOBBY 61%LOBBY 25%PARTNER 79%LOBBY 67%
AsmlBUILDRESTRICT 44%BUILD 59%ENFORCE 26%BUILD 21%PARTNER 79%PARTNER 21%
MicrosoftLOBBYPARTNER 29%BUILD 63%LOBBY 65%LOBBY 52%PARTNER 91%LOBBY 61%
GoogleLOBBYRESTRICT 27%BUILD 59%LOBBY 52%LITIGATE 36%PARTNER 89%LOBBY 51%
MetaLOBBYRESTRICT 31%BUILD 49%LOBBY 50%LITIGATE 43%PARTNER 81%LOBBY 44%
AmazonLOBBYPARTNER 28%BUILD 65%LOBBY 63%LOBBY 51%PARTNER 90%LOBBY 52%
AppleBUILDRESTRICT 30%BUILD 59%ENFORCE 35%BUILD 24%PARTNER 84%PARTNER 27%
OpenaiLOBBYPARTNER 23%BUILD 60%LOBBY 58%LOBBY 34%PARTNER 87%LOBBY 55%
AnthropicLOBBYRESTRICT 30%BUILD 67%LOBBY 54%LOBBY 37%PARTNER 89%LOBBY 53%
HuaweiRESTRICTRESTRICT 54%BUILD 52%ENFORCE 31%RESTRICT 24%PARTNER 66%ENFORCE 22%
DeepseekBUILDBUILD 21%BUILD 43%ENFORCE 27%LITIGATE 41%PARTNER 83%LITIGATE 21%
7

Actor-relevance classifier

which actors relate to a bill · complete gold · large held-out split

F1 0.756

precision 0.67 · recall 0.86, on 8,977 held-out pairs.

bal-acc 0.814

complete thorough-relatedness gold.

20,821 train

443 bills · 7,504 related pairs.

8,977 test

191 held-out bills · 3,213 related.

F1 progressionvs complete gold
Cosine similarity + rank0.69
+ reduced embedding dims0.72
+ shared-basis embeddings & product0.74
+ stronger GBT, F1-tuned threshold0.75
+ 2× data, large held-out split0.76

LLM-embedding semantics of bill vs. each actor (TF-IDF can't see "chip export controls" ~ NVIDIA; the embedding can), + shared-basis interaction features. Measured against a complete gold labeled by a strong annotator (sonnet-4.5, distinct from the classifier — not circular), superseding the earlier sparse-gold 0.31. F1 0.75 is the honest number; the residual is nuanced policy judgment only a reasoning model captures.