(row_id_a, row_id_b, score) tuples.
Scorer reference
The
name_freq_weighted_jw / given_name_aliased_jw scorers ship as part of the bundled reference-data packs and are picked automatically by auto-config when a column matches the relevant name pattern AND its profiled col_type agrees. See Reference Data for the full pack overview, refinement rules, and the col_type gate.
Date fields
Do not score an ISO date withjaro_winkler. The fixed YYYY-MM-DD shape, the shared digit alphabet, and the common 19../20.. prefix push unrelated birthdays to 0.80+, so a threshold that admits a one-digit typo also admits a different person — precision collapses on any date column that blocking co-locates by birth year.
The date scorer parses both sides as ISO dates and compares them by Damerau-Levenshtein edit distance over the eight canonical digits (a swapped-digit typo is one edit), mapped so a single-digit typo stays high while an unrelated date drops to 0:
levenshtein. Auto-config already skips date columns as fuzzy fields, so date is for hand-written / Splink-converted configs; a preflight check warns when a name-oriented scorer (jaro_winkler / token_sort / …) is placed on a date field.
The date scorer is edit-distance and therefore magnitude-blind — 1990-01-02 vs 1991-01-02 is one digit-edit → 0.90, over-scoring a full-year gap. The date_diff scorer instead parses both sides to a day-ordinal and bands by day-distance (same → 1.0, ≤1 day → 0.92, ≤1 month → 0.80, ≤1 year → 0.60, ≤5 years → 0.30, else 0), with an MM/DD-transposition floor and the same levenshtein fallback on unparseable input. On the Fellegi-Sunter path, GOLDENMATCH_FS_DOMAIN_COMPARATORS=1 makes auto-config admit date columns as date_diff instead of levenshtein (default off; byte-identical when off).
Numeric and coordinate fields
The same magnitude-blindness afflicts numbers and coordinates:levenshtein("100","900") ≈ 0.67 reads two very different amounts as near-agreement, and string similarity on a lat,long pair is meaningless. Two more domain comparators (behind the same GOLDENMATCH_FS_DOMAIN_COMPARATORS flag) fix this on the Fellegi-Sunter path:
numeric_diffparses both sides to a float and maps the distance to a monotone[0,1]ramp —numeric_diff:abs:<eps>for an absolute band (|a−b|) ornumeric_diff:pct:<frac>for a relative band (|a−b| / max(|a|,|b|)); barenumeric_diff=pct:0.1. Auto-config admits numeric columns asnumeric_diff:pct:0.1.geo_haversineparses one combined"lat,long"field per side and bands the great-circle (haversine) km distance (≤0.1 km → 1.0, ≤1 km → 0.85, ≤10 km → 0.5, ≤100 km → 0.2, else 0). Auto-config admits a column whose sampled values parse as coordinates. (Two separate lat/long columns are a deferred cross-field comparator.)
None for non-null values), so the scalar and vectorized scoring paths agree by construction — and, like date_diff, they are scale-neutral: just new scorers flowing through the unchanged level → m/u → weight machinery, leaving blocking, the pair set, memory, and clustering untouched. Default off is byte-identical.
Fuzzy scoring
Fuzzy matching usesrapidfuzz.process.cdist for vectorized NxN scoring within each block. This is the core scoring engine for weighted matchkeys.
Weighted matchkeys
Each field gets a scorer, weight, and optional transforms. The overall score is a weighted average:overall_score = sum(field_score * weight) / sum(weight)
Pairs with overall_score >= threshold are matched.
Exact scoring
Exact matching uses Polars self-join for high performance. No threshold needed.Probabilistic scoring (Fellegi-Sunter)
EM-trained m/u probabilities with comparison vectors. Match weights are log-likelihood ratios.level_thresholds generalizes the fixed agree/partial/disagree banding to any
number of levels: pass descending similarity cutoffs (len == levels - 1), and
a pair’s comparison level is the count of thresholds its similarity satisfies
(satisfying all of them = the top “agree” level, none = “disagree”). Omit it to
keep the classic 2/3-level behavior (partial_threshold covers the 3-level
case). The native FS kernel scores level_thresholds matchkeys natively from
goldenmatch-native >= 0.1.14 (older wheels fall back to the pure-Python
scoring path automatically); the fused-match path scores them natively from
goldenmatch-native >= 0.1.15 (capability const
FUSED_FS_SUPPORTS_LEVEL_THRESHOLDS; older wheels fall back).
- u-probabilities estimated from random pairs and fixed during EM (Splink approach)
- Blocking fields must be excluded from training (always agree within blocks)
- Comparison vectors apply field transforms before scoring
- Achieves P=0.978 / R=0.958 / F1=0.968 on DBLP-ACM (full-block vectorized scoring). The old “98.8% / 57.6%” figure was a benchmark artifact — a per-block size cap, not the scorer.
Negative evidence on Fellegi-Sunter matchkeys
negative_evidence is valid on type: probabilistic matchkeys (not just weighted/exact). Each NE field is a constrained EM-learned dimension: it fires when both values are present and scorer(a, b) < threshold (strict <), contributing log2(m_fired/u_fired) bits when fired and exactly 0 when not fired — that fired-else-zero clamp is what makes it negative evidence rather than a normal scored field (agreement never adds weight; only a hard disagreement subtracts it).
- EM-learned (default): omit
penalty_bits. The weight is estimated from data duringtrain_emthe same way every other field’s m/u is —m_fired = P(fired | match),u_fired = P(fired | non-match), both from the same EM loop and random-pair sample as the regular fields. penalty_bits(fixed override): a literal log2 LLR in bits; the fired contribution is exactly-abs(penalty_bits), no EM training for that dimension. Useful when you know the right veto strength (e.g. migrating a Splink config) without waiting on EM to converge on enough data.
penalty (the weighted/exact knob) and penalty_bits are mutually exclusive by matchkey type: weighted/exact matchkeys still require penalty and reject penalty_bits; probabilistic matchkeys reject penalty and accept penalty_bits (or neither, for EM-learned).
An unregistered/unknown NE scorer on a probabilistic matchkey fails loudly at train/score time (score_field raises on unknown scorers; there is no _NE_BROKEN swallow on FS) — unlike the weighted path, which silently swallows a broken NE scorer and warns. This is intentional: FS negative evidence is new enough that a silent no-op would be worse than a hard error.
Guards: the native kernel scores NE-bearing FS matchkeys from goldenmatch-native >= 0.1.15 (capability const FS_SUPPORTS_NE; requires every NE scorer to be FS-native), and the fused-match path does likewise; older wheels keep the pure-Python fallback automatically. The fast-path per-pair scorer does not implement NE and defers to the standard path. Persisted/imported EM models trained before this feature (including imported Splink models) fail loudly on load if the matchkey has NE fields without penalty_bits and the model lacks the corresponding trained weights, rather than silently scoring NE at weight 0 — retrain, or set penalty_bits to bypass EM entirely. NE is not supported on the continuous/Winkler EM path (train_em_continuous), which rejects it with a clear error.
Splink-parity surface
The Fellegi-Sunter matchkey is a full probabilistic-linkage engine, not just a scorer:- Model lifecycle (train-once, reuse).
EMResult.save_json/load_json+MatchkeyConfig.model_path(ordedupe_df(fs_model_path=...)) train EM once and reuse the model — no retraining per run. - Supervised m from labels.
estimate_m_from_labels(df, mk, labels)(Splink’sestimate_m_from_label_column) estimates m directly from known matches; adapters pull labels straight from the review-queue / memory corrections store. - Match-weight waterfall.
explain_pair_fsdecomposes a pair into per-comparisonlog2(m/u)bits + prior + posterior, surfaced ingoldenmatch explain --pairand the lineage sidecar. - Calibration.
GOLDENMATCH_FS_CALIBRATED=posteriorturns the score into a true match probability1/(1+2^-(log2(λ/(1-λ)) + ΣW));linear(default) is monotonic in the summed weight. - Accuracy analysis from labels.
goldenmatch evaluate --threshold-sweepemits the precision/recall/F1 operating-point curve, a recommended cut,probability_two_random_records_match, and the per-comparison m/u match-weight report. - Config migration from Splink.
goldenmatch import-splink settings.json -o goldenmatch.yaml [--model-out model.json](orgm.from_splink(...)in Python, or theconvert_splink_configMCP tool) converts a Splink settings or trained-model JSON into a GoldenMatch config; trained m/u probabilities import directly so no re-training is needed, and anything lossy is reported in aConversionReport. - Splink migration upgrade pass.
goldenmatch import-splink settings.json --upgrade data.parquet -o goldenmatch.yaml --model-out model.json(orgm.upgrade_splink_conversion(conversion, data)) runs a data-aware pass over the converted config with four levers: it computes term-frequency tables from the data (Splink model exports don’t include them), re-derives Levenshtein distance thresholds from measured string lengths, applies fan-out defenses (a risk-gated negative-evidence suggestion — an unused identity-grade column whose disagreement contradicts pairs the imported model would confidently merge is added asnegative_evidencewith posterior-weighted__ne__weights — plusgolden_rules.max_cluster_sizetuned from the reference clusters), and calibrates link/review thresholds from the blocked-pair score distribution (NE-aware). The faithful baseline conversion is always written alongside as*.baseline.*, and a baseline-vs-upgraded delta table is printed (add--splink-clusters/--labelsfor agreement / truth F1;--sample-cap,--no-measure, and--id-columncontrol the measurement pass). Measured on wild Splink configs, the upgraded conversion beat native Splink on all three benchmark pairs (pairwise F1 vs truth: 0.633 vs 0.601, 0.766 vs 0.699, 0.740 vs 0.686).
Scale-out
Probabilistic matchkeys ride thebucket backend’s hash-bucketed parallel scorer (the same path that carries the Ray / DataFusion distribution wiring), so they scale single-node the same way weighted matchkeys do. Measured 6M-row dedupe on a 16c/64GB node: 162.6 s (native) / 288.5 s (numpy), 11.3 GB peak RSS, F1 1.000 on synthetic data with entity-clique ground truth. An opt-in Rust kernel (GOLDENMATCH_FS_NATIVE=1) is ~10.8× on the scoring step for tiny-block workloads. See scale envelope.
Embedding scorers on the probabilistic path. embedding and record_embedding are first-class Fellegi-Sunter field scorers: they both train (EM) and score on the vectorized matrix path (their cosine-similarity matrix flows through the same comparison-level logic as string scorers, so training and scoring agree). Because they are matrix-only, a matchkey carrying one always runs vectorized — the GOLDENMATCH_FS_VECTORIZED=0 debug fallback only applies to string scorers.
Probabilistic auto-config v2 (beats Splink head-to-head)
The probabilistic auto-config path (auto_configure_probabilistic_df / build_probabilistic_matchkeys) builds Fellegi-Sunter configs that beat Splink on every dataset Splink scores in the shared bench_er_headtohead panel. This is default-on as of FS auto-config v2; set GOLDENMATCH_FS_AUTOCONFIG_V2=0 to restore the legacy auto-config byte-identically. It touches only the probabilistic path — the weighted/DQbench path and zero-config dedupe_df are unchanged. Numbers are deterministic as of #829 (which fixed a non-deterministic EM training-pair sample that previously swung historical_50k F1 between 0.64 and 0.80).
GoldenMatch also wins at the cluster level on historical_50k (B-cubed F1 0.844 vs 0.789). The full three-engine accuracy + performance bake-off (incl. the zero-config controller path, wall, peak RSS, throughput) is at
docs/benchmarks/2026-06-09-splink-bakeoff.md.
Four levers drive the gain:
- Admit
dob/ date fields aslevenshteininstead of leaving them exact-only, so near-miss dates still contribute weight. - Drop redundant name composites (
full_name/first_and_surname) when atomicgiven+familyfields already exist — no double-counting the same signal. - Additively diversify blocking onto orthogonal stable keys (date-year + postcode/zip) so a single noisy key doesn’t cap recall.
- Admit
description/ multi-name fields astoken_sort, which lifts the bibliographic case (DBLP-ACM) from 0.003 → 0.377 — a large relative gain, but still recall-bound.
These are pairwise F1 under one shared evaluator (
bench_er_headtohead). The often-cited ~0.97 Splink figure on historical_50k is a cluster/entity-level metric, not exhaustive within-cluster pairwise F1 — Splink itself scores ~0.75 pairwise on this dataset under the same harness, and the pairwise blocking ceiling for any engine on these columns is ~0.93. The honest claim is “matches/beats Splink head-to-head on the same evaluator,” not “0.97 pairwise.” Where Splink still leads: it is 3-19x faster on these datasets, runs distributed FS at 1B+ rows on Spark, and ships a mature interactive m/u comparison-viewer charting UI.For bibliographic data (DBLP-ACM), use the weighted path, not probabilistic. Splink skips dblp_acm, and the probabilistic auto-config is weak there (pairwise F1 0.377, recall-bound). The zero-config weighted controller scores 0.964 on DBLP-ACM and is the right tool for that shape. The probabilistic path targets PII / person-record linkage (historical_50k, febrl3, synthetic_person).
LLM scoring
Send borderline pairs to GPT-4o-mini or Claude for scoring. Two modes:Pairwise mode
Score individual pairs. Best for small candidate sets.Cluster mode
Send entire borderline blocks to the LLM for in-context clustering. More efficient for large blocks.Cross-encoder reranking
Re-score borderline pairs with a pre-trained cross-encoder for higher precision.threshold +/- rerank_band get reranked. Requires pip install goldenmatch[embeddings].
Parallel scoring
Fuzzy blocks are scored concurrently viaThreadPoolExecutor. RapidFuzz’s cdist releases the GIL, so threads provide real parallelism.
- Blocks are independent — frozen
exclude_pairssnapshot avoids race conditions - For 2 or fewer blocks, threading overhead is skipped (sequential execution)
- All call sites (pipeline, engine, chunked) use the shared
score_blocks_parallelhelper - Ray backend (
score_blocks_ray) distributes blocks across Ray tasks for cluster-level scaling
Intra-field early termination
After scoring each expensive field, the scorer checks if the remaining fields can push any pair above the threshold. If not, it breaks early. This reduces 100K fuzzy matching from ~100s to ~39s (2.5x speedup).Embedding scoring
Requirespip install goldenmatch[embeddings].