Backend by row count
GoldenMatch also has a
chunked backend used for very large single-machine dedupe (streaming CSV reader plus a Polars cross-chunk join plus a block-keyed index). See backends and scale.Measured numbers
These are quoted from the repository and were measured on the versions noted. Re-measure for your own hardware.Block-size failure modes
Most ER blow-ups come from a blocking key that produces a few huge blocks. GoldenMatch enforces guard rails:max_block_size: 5000in hand-writtenBlockingConfig.max_safe_block = 1000in auto-config (blocks over 1000 can OOM ensemble scorers).skip_oversized: falseby default. When true, oversized blocks are ANN-sub-blocked or skipped.
Cardinality guards (v1.2.7)
The common-email trap
Usingexact=["email"] as the only matchkey creates oversized clusters around shared values like info@, noreply@, or null. Symptoms: one giant cluster, stalled scoring, memory pressure. Fix it by standardizing or stripping common patterns, or by adding a second blocking pass.
Checklist before scaling
1
Profile your blocking key
Check its
cardinality_ratio. Auto-config prints this in its postflight health report.2
Respect the block-size cap
Stay under
max_block_size=1000 for auto-config, or 5000 for hand-written configs.3
Pick the backend from row count
< 500K → Polars; 500K – 50M → DuckDB; ≥ 50M or many large blocks → Ray.
4
Re-measure if exact numbers matter
The published figures are baselines on specific versions and hardware.
Tuning & opt-ins
The flags that make a scale run fast — the native gate (
GOLDENMATCH_NATIVE=1), the address-normalize fast path, the distributed pipeline contract, and every GOLDENMATCH_* default. Set these before a large run.