Technical report · run fedjev-2026-09-20

fedjev-bench

Pairwise textual hawkishness scores for FOMC chair openings, evaluated against same-day funds-rate changes

Abstract

Same-day changes in the federal funds target are a convenient but incomplete label for the hawkishness of Federal Open Market Committee (FOMC) communication. Policy actions and textual stance often co-move on scheduled action days. They need not coincide when the Committee leaves the target unchanged.

This note reports a pre-registered evaluation of TypeSafe/Jev on chair press-conference openings under the fixed pairwise criterion more hawkish about inflation. Primary quantities include standard errors (s.e.) and bootstrap percentile confidence intervals. Gate 4 is a construct-validity result: same-day funds-rate changes are an incomplete label for textual hawkishness when the target is unchanged. Gate 7 is Spearman’s rank correlation with FedLock — an independent published text-scoring project — using that project’s published press-conference scores (raw TrueSkill mean m; era-adjusted ma as a sensitivity). Gate 7 reads those scores; it does not re-run FedLock’s tournament. A separate TrueSkill replica on the 95 openings is a different claim.

Measurement

Two constructs are distinguished throughout. The behavioral measure is the same-day change in the funds target (d_same — hike, hold, or cut; identically zero on holds). The textual measure is the stance in the Chair’s opening remarks, recovered by dual-order Choice and Bradley–Terry aggregation (a pairwise strength model on a fixed gold-pair graph), with a secondary Score pass. FedLock’s published scores are a second text measure, not a behavioral label.

Gates 1, 3, 4, and 6 are pre-registered pass/fail tests. Gates 2, 5, and 7 are report-only or secondary. Gate 7 is not a TrueSkill replication. The exercise does not identify a causal effect of communication on rates. A plain-language scoreboard guide explains what each gate answers and how to read inversion rates, rank correlations, Score gaps, and cost.

Selected estimates

GateEstimate
1 Easy-pair inversionPASS — 0.000 (n=40, STE=0.000)
3 Action-day Spearman (BT vs d_same)PASS — +0.851 (n=24, STE=0.074)
4 Holds − cuts (BT mean gap)PASS — +0.519 (STE=0.368)
6 Name/order stabilityPASS — Δ inversion = 0.000
7 FedLock consistency (score_jev vs published raw m; not a TrueSkill replication)report — Spearman +0.944 (n=90, s.e.=0.016)

Selected figures

Text scores versus same-day target change
Figure 6. BT (n=46) and score_jev (n=93) against d_same. Holds stack at zero. Spearman ρ and bootstrap STE are the Gate 3 all-scheduled estimates.
Mean scores for holds, cuts, and hikes
Figure 8. Mean text scores by same-day action, ± STE. Holds sit above cuts on both the BT and Score axes.

Documents