Skip to main content
This is the one place that lists every runtime knob GoldenMatch reads, native acceleration, backend selection, the distributed pipeline, the prep/scoring perf opt-ins, and the full GOLDENMATCH_* environment-variable table. For each, you get what it does, its default, and when to set it (or leave it alone). If you are tuning a large run and just want the answer, jump to Recommended setups. If something ran far slower than you expected, read Gotchas that cost a day first. Every item there is a real footgun.
Most GoldenMatch knobs are environment variables, not config fields, because they switch a code path (native vs Python, in-memory vs distributed) rather than describe your data. Data-shape settings, matchkeys, blocking keys, golden rules, live in Configuration. This page is the operational layer on top of that.
You almost never need to set anything for correctness: auto-config and the planner choose sane defaults. These blocks are about speed at scale and reproducibility, not behavior.

Local / interactive (up to ~500K rows)

Set nothing. Zero-config is the supported path. If you want the compiled kernels for a measurable speed-up on the scoring/clustering hot loops, install the optional native wheel:
With the wheel present, the default GOLDENMATCH_NATIVE=auto already runs every component that has a kernel natively (native is the reference; pure-Python is the lossy fallback). You do not need to set GOLDENMATCH_NATIVE=1.

Single box at scale (1M–200M rows, 16-core / 64 GB class)

The planner picks backend=bucket for this lane automatically when the native kernel is enabled. This is the exact env used by the quality-invariant scale ladder (1K → 100M), so it is the validated single-box recipe.
Setting backend="ray" does NOT distribute the run. On its own it falls back to an in-memory path that materializes the whole frame on the driver. Real distribution needs GOLDENMATCH_DISTRIBUTED_PIPELINE=2 plus a Ray Dataset input, see Distributed pipeline and Gotchas. At single-box scale, prefer bucket.

Distributed (Ray cluster, 50M+ rows)

Phase 5 has a hard contract: a Ray Dataset input with a global __row_id__ column, and it writes golden records to output_path rather than returning an in-memory clusters dict. The recall-complete path (block-shuffle scoring + randomized-contraction WCC) is default-on and was validated end-to-end at 100M rows on a real 5-node GCP cluster (9.2 min, 20M clusters recovered exactly, 0.36 GB driver RSS). See Distributed pipeline for the full contract. At single-box scale, prefer bucket. Only reach for a Ray cluster when the data genuinely won’t fit one box.

Native acceleration

Rust is the reference implementation (as of the 2026-07 reference-mode change). The Rust/PyO3 kernel is the source of truth for GoldenMatch’s hot-path primitives; pure-Python is the lossy fallback used only when a native wheel isn’t installed (the wheel ships as pip install goldenmatch[native]). A central gate (core/_native_loader.py) selects per component. GOLDENMATCH_NATIVE (core/_native_loader.py): Under auto, native runs for every component that has a kernel symbol. clustering, block_scoring, pairs, featurize, hashing, autoconfig, sketch, pprl_bloom, and the multimodal perceptual hashes. Each cleared bit-/ULP-level parity against the Python reference, so the fallback is output-identical (golden-vector verified). (The former per-component allowlist, _GATED_ON, is retained in the loader as sign-off documentation but no longer governs auto.) The sketch component covers both MinHash/LSH (#1081) and SimHash-over-embeddings (#1090) band hashing (~29x faster; the single source of truth across Python/Rust/TS, so semantic blocking runs on the Rust sketch-core kernel by default). Fellegi-Sunter block scoring is native-authoritative by default as of goldenmatch 2.6.0 (GOLDENMATCH_FS_NATIVE=0 opts out to the reproducible numpy path); it is scoped to the probabilistic path only. The autoconfig component is the shared, pyo3-free decision core (goldenmatch-autoconfig-core) behind auto-config, the planner, column classifier (with the adaptive 1 − 1/√n identifier-cardinality floor), the sample→full pair-count extrapolation (the corrected ratio² + Chao1 estimate), and the adaptive decision thresholds (the sparse-match floor and the per-type exact-matchkey floor). It runs native under auto and is byte-identical to the pure-Python oracle (golden-vector parity across Python / native / wasm-TS). The same crate compiles to WebAssembly so the TypeScript surface consumes the identical decision logic, one source of truth, zero cross-surface drift. When to set =1: any large or benchmarked run where a silent Python fallback would be a disaster. The cost of a missing wheel is enormous (the scoring kernel is the dominant stage), and auto hides it. When to leave it auto: normal local use. If the wheel is installed you already get native; if it isn’t, you get working Python.

GOLDENMATCH_NATIVE_ADDRESS_NORMALIZE

Default off. When set to 1, address-normalize transforms run as a Polars-native expression instead of the per-row Python refdata plugin. On address-heavy data at scale the per-row plugin is a real bottleneck in the prep stage; the native path removes it.
  • Set it (=1) when you have address columns and are running at scale. Output is parity-tested on representative inputs.
  • Leave it off for small runs, or if you hit an address-canonicalization edge case (it is opt-in precisely because one integration-level cluster-count drift hasn’t been fully root-caused).

GOLDENMATCH_NATIVE_RAYON_MIN_PAIRS

Default 20000000. Candidate-pair count per kernel call below which the native block-scorer runs in the calling thread instead of fanning out to rayon. This exists to avoid a futex-park stall observed on some 8-core Linux CPUs (issue #688). The Python caller already parallelizes across buckets, so the sequential path loses nothing on small calls. Leave it at the default unless you are reproducing #688; 0 forces rayon always, a huge value forces sequential always.

GOLDENMATCH_NATIVE_SKETCH_RAYON_MIN_ROWS

Default 10000. Row count below which the native MinHash/LSH batch sketcher (#1081) runs in the calling thread instead of fanning out to rayon, the same #688 calling-thread-vs-fanout guard, applied to the sketch kernel. The SimHash batch sketcher (#1090) shares this threshold. Relevant whenever the sketch kernel runs native, which is the default under auto since #1090 (set GOLDENMATCH_NATIVE=0 to force the Python path), or under GOLDENMATCH_NATIVE=1. 0 forces rayon always, a huge value forces sequential always.

GOLDENMATCH_ER_MODEL_DIR

Default ~/.cache/goldenmatch/er-matcher. Directory where the local ER-matcher LLM boost caches its downloaded GGUF weights. Set it to relocate the cache, e.g. to a shared volume, a larger disk, or a read-only mount that ships the weights pre-seeded. The pinned model is sha256-verified on resolve regardless of location, so a relocated (or externally populated) cache is checked the same way as the default one.

Backend selection

Same pipeline, different execution engine. Auto-config picks one from your row count; you can override. Force one in config or on the CLI:
Auto-selection knobs:
The old single-threshold GOLDENMATCH_AUTOCONFIG_BACKEND / GOLDENMATCH_AUTOCONFIG_BACKEND_THRESHOLD env overrides were removed in v2.0. Backend selection is now the v3 planner’s job, it decides from (ComplexityProfile, RuntimeProfile). To force a backend, set it explicitly on GoldenMatchConfig(backend=...) or the CLI --backend.

ANN blocker backend (embedding/semantic blocking)

The embedding ANN blocker (ANNBlocker) resolves one of three inner-product backends per index build: native HNSW → FAISS → numpy. Native HNSW is the optional goldenmatch-hnsw wheel, a pure-Rust IndexHNSWFlat with zero C dependencies (no FAISS); FAISS (goldenmatch[embeddings]) is exact IndexFlatIP; numpy is the always-available all-pairs fallback. Scores are the raw inner product on the HNSW/FAISS paths and cosine on numpy, identical when vectors are L2-normalized (the default embedder output). In auto mode HNSW is chosen only where it wins: a large corpus and a small neighbor count per probe. Below that gate, brute force is both faster and exact, so small indexes (and the retrieve-all pattern behind the persistent VectorIndex) stay on the exact path unchanged.
goldenmatch-hnsw is an optional wheel, installed separately (pip install goldenmatch-hnsw) exactly like FAISS. Without it, the ANN blocker uses FAISS if present, else the numpy fallback. No behavior change.

Distributed pipeline

This is the knob most people get wrong, so read the contract before you run a 100M job. GOLDENMATCH_DISTRIBUTED_PIPELINE selects how run_dedupe_pipeline_distributed (the Ray Dataset entry point) executes: Phase 5 (=2) contract: different from in-memory dedupe:
  • Input must carry a global __row_id__ column. Without it, each partition synthesizes local IDs that collide and the weakly-connected-components step merges unrelated clusters.
  • Output is written to output_path as parquet golden records. It does not return an in-memory clusters dict the way dedupe_df does.
  • It still materializes two things on the driver: the ~5K-row auto-config sample, and cluster aggregates for golden-field synthesis. These are the last “cheat-lines” being retired.
GOLDENMATCH_ENABLE_DISTRIBUTED_RAY (default 0) is separate: it only lets the v3 planner pick ray during backend selection. It does not turn on the streaming pipeline. Explicit backend="ray" always works without it. The planner stopped auto-picking ray after the Distributed Plan v1 stack failed the 5M kill criterion against the bucket backend, so ray is opt-in. Recall-complete clustering at scale (#844, shipped 2026-06-11). Phase 5’s default path now block-shuffles scoring so true duplicates co-locate, then runs a relational randomized-contraction weakly-connected-components pass that clusters correctly across input-partition boundaries. This replaced two earlier WCC implementations that died at 100M (a driver head-wedge and a Ray streaming-executor deadlock). It is default-on and validated end-to-end at 100M (9.2 min, 20M clusters exact). On a multi-node cluster the WCC checkpoints each round to GOLDENMATCH_DISTRIBUTED_WCC_SCRATCH, which must be a shared path (e.g. gs://...). A node-local scratch breaks the cross-node parquet reads and now raises rather than silently under-merging.
Only scoring needs distribution at this scale; clustering usually doesn’t. The scored pair set is small (~(id_a, id_b, score) per match), so at 100M rows it’s ~1.8 GB, it fits one node. Clustering therefore routes to the in-memory scipy connected-components pass (seconds) automatically whenever the pairs fit driver RAM, and only falls back to the distributed WCC (a multi-hour tail) when they genuinely don’t. You don’t need to set GOLDENMATCH_DISTRIBUTED_CLUSTERING_THRESHOLD for this, it’s the default.
Correlated survivorship (field_groups and conditional field_rules) does not run on the distributed or Sail paths. The golden builder that resolves lock-step groups and conditional predicates runs on the driver. If your config includes field_groups or list-form field_rules, the Ray Phase-5 streaming pipeline (GOLDENMATCH_DISTRIBUTED_PIPELINE=2) and the Sail backend will refuse with a clear error rather than silently mis-merging. Run those configs on the in-memory path.
Distributed sub-knobs (only relevant on a real Ray cluster):
On a Ray cluster, exporting these four in your shell does not reach the run. ray submit hands the driver a fresh shell, so a value set around the submit never arrives. Pass them as flags on the driver script instead — --score-num-cpus, --score-concurrency, --score-project / --no-score-project, --op-reservation, --shuffle-parts — which travel in the argv the submit already ships. The engine reads all four at call time, so a driver that sets them after importing goldenmatch.distributed is fine. scripts/quality_invariant_scale.py carries these flags and records the effective values in its output JSON under score_knobs.
If the distributed score stage leaves workers idle (MapBatches(_score): … [backpressured:tasks(ResourceBudget)] with the object store well under capacity), _score is reservation-capped, not CPU-bound. Lower GOLDENMATCH_DISTRIBUTED_OP_RESERVATION (→ 0.2) and/or raise GOLDENMATCH_DISTRIBUTED_SHUFFLE_PARTS, then re-measure. The optimal pair is cluster-shaped. (Shrinking Ray Data’s target_max_block_size is not a safe lever: _score keeps each co-located partition whole via batch_size=None, so a split block would split a group and under-score.)
You bring the Ray cluster, GoldenMatch ships the pipeline, not the bootstrap. The GCP recipe is in docs/distributed-ray-cluster-setup.md. At single-box scale the bucket backend is validated through 200M and is simpler, use it unless the data won’t fit one box.

Fast path at scale (distribute only what doesn’t fit)

The single most expensive distributed-tuning mistake is distributing a stage that fits in memory. “The job is big” does not mean “every stage must be distributed.” Ask it per stage: Concrete: don’t force the distributed WCC on a pair set that fits. By default GOLDENMATCH_DISTRIBUTED_CLUSTERING_THRESHOLD is unset and the routing is memory-aware, clustering stays on the fast in-memory scipy.csgraph path whenever the scored pair set fits available driver RAM, and only distributes when it genuinely won’t. At 100M the scored pair set measures 609.4M pairs and the in-memory route’s driver peak measures ~81 GB, so the default keeps clustering in-memory on a 128-256 GB head and distributes on a 64 GB one. Set an explicit integer only if you want to force a hard pair-count boundary either way:
Setting it to 0 does the opposite: it forces the slow distributed WCC on everything. Recover a result without re-running. The Phase-5 connected-components clustering is independent of the WCC algorithm, so a completed run’s cluster assignments (written to gs:// via the assignments_output_path hook) are the result. If a run finishes but the output is lost, or you want to re-score, download the assignments and score them against your ground truth on a 64 GB box, no cluster, no re-run. The 100M group_by for that scoring OOMs a 16 GB box, so run it on a large runner, not your laptop.

Prep & scoring perf opt-ins

These trade nothing for speed/memory at scale and are safe to leave at their defaults. Two are already on. Diagnostic-only: these can SLOW the run:

Fused kernel auto-routing

The pipeline can route covered slow-path runs to the fused Arrow-native kernels, byte-identically, with no config change. There are two seams with very different maturity: golden routing is a broad default-on win; match routing is a narrow capacity-survival path that fires only under memory pressure.

Auto-config & planning

If auto-config raises ControllerNotConfidentError, that’s by design at ≥100K rows on a RED profile. It’s refusing to run a noise-result dedupe. See Passing an explicit config for recovery (the in-repo docs/explicit-config.md has the failing-sub-profile table). There are also a number of experimental blocking and probabilistic knobs (GOLDENMATCH_BLOCKING_*, GOLDENMATCH_FS_*, GOLDENMATCH_NOISE_AWARE_*, GOLDENMATCH_QUALITY_AWARE_BLOCKING, GOLDENMATCH_FD_NEGATIVE_EVIDENCE). These are research/tuning levers, default-safe, and listed in the full table below. Most users never touch them.

Config-suggestion (healer) knobs

These tune the healer: the review_config config-suggestion loop, which is now wired into the default dedupe_df pipeline as an advisory surface (a free controller signal gates a cheap attach; see GOLDENMATCH_SUGGEST_ON_DEDUPE). The healer needs goldenmatch[native] and degrades gracefully without it. You almost never need to change these, the defaults are the values the suggester gym selected.

Servers, security & storage

When you expose any server (MCP, REST, A2A, web UI) on a non-loopback host, it is fail-closed: it refuses to start without the matching bearer token. Storage / persistence: Product analytics (opt-in): Anonymous usage analytics: off by default. Emits only when GOLDENMATCH_ANALYTICS is on and a POSTHOG_API_KEY is set, so a user’s machine stays silent unless they opt in (the project’s own hosted services set the flag server-side). PII-free by construction: only coarse scalars (scale bands, backend, version) ever leave, never record values, column names, or paths. Full collection list + guarantees in packages/python/goldenmatch/ANALYTICS.md.

Gotchas that cost a day

Read these before a big run. Each is a real trap that produces a slow or wrong result with no error.
backend="ray" (or dedupe_df(backend="ray")) alone runs the Phase 2 cheat-line: it collects the whole frame to the driver and dedupes in memory. The cluster sits idle. To actually stream across workers you need GOLDENMATCH_DISTRIBUTED_PIPELINE=2, a Ray Dataset input with a global __row_id__, and you read golden output from output_path. Not the return value. On a multi-node cluster also set a shared GOLDENMATCH_DISTRIBUTED_WCC_SCRATCH (e.g. gs://...). For single-box scale, use bucket, not ray.
Under the default GOLDENMATCH_NATIVE=auto, if the native wheel isn’t importable (not installed, ABI mismatch, stale build) the scoring kernel silently runs in Python. Orders of magnitude slower. On any run where speed matters, set GOLDENMATCH_NATIVE=1 so a missing wheel raises instead of crawling. Confirm with python -c "import goldenmatch_native" or check goldenmatch.core._native_loader.native_available().
The per-row Python address-normalize plugin is a real bottleneck at scale. Set GOLDENMATCH_NATIVE_ADDRESS_NORMALIZE=1 to use the Polars-native expression instead. It’s off by default, so you have to opt in.
The scored pair set fits one node on a large head (609.4M pairs at 100M, ~81 GB driver peak measured — fits 128-256 GB, not 64 GB). By default GOLDENMATCH_DISTRIBUTED_CLUSTERING_THRESHOLD is unset and routing is memory-aware, clustering stays in-memory (scipy.csgraph, ~60s) whenever the pair set fits driver RAM, so a 100M run should already keep clustering on the fast path. The multi-hour tail happens if something forces the distributed WCC (whose gs://-checkpoint rounds are slow): check you haven’t set the threshold to 0 (forces the slow path on everything) or to a value below your pair count. See Fast path at scale. Distribute the scoring (the rows don’t fit), not the clustering (the pairs do).
Cross-run auto-config memory (GOLDENMATCH_AUTOCONFIG_MEMORY=1, the default) lets a prior run influence the next. For reproducible benchmarks set GOLDENMATCH_AUTOCONFIG_MEMORY=0.
Polars runs a WMI CPU probe at import that can hang on some Windows boxes. Set POLARS_SKIP_CPU_CHECK=1. (Add PYTHONIOENCODING=utf-8 if non-ASCII output mangles the console.)

Throughput tier (sketch-then-verify)

New in #1083, builds on #1080. An opt-in high-throughput path for large text corpora where you care about near-duplicate recall but can trade precision for speed. Instead of the full fuzzy/Fellegi-Sunter pipeline, it blocks with MinHash/LSH (lsh) or SimHash (simhash) and confirms candidate pairs by cheap sketch distance before any field-level scoring. Think: training data dedup, web crawl dedup, paper-corpus cleaning.

Enabling it

Pass throughput= to dedupe_df:
Or via environment variables:

What it does

  1. Auto-selects blocking strategy. Picks simhash (cosine/semantic) when an embedder is reachable, otherwise falls back to lsh (MinHash/Jaccard on shingles). Uses the longest text column in the frame.
  2. Blocks with sketch buckets. Records that share at least one LSH band become candidates. The recall_target controls the banding split via the LSH probability curve.
  3. Verifies by sketch distance only. Pairs that clear similarity_threshold are kept; the field-level fuzzy/FS scorer does not run.
  4. Raises ThroughputNotApplicableError when the frame has no text column. Fix: pass an explicit BlockingConfig to direct it to a column.

Posture and honest reporting

The throughput tier does not measure F1 or precision. DedupeResult.throughput_posture (and the controller telemetry throughput block) carry the metrics the tier can honestly report: expected_recall is the LSH-theoretic probability that a pair at exactly similarity_threshold Jaccard/cosine collides in at least one band. It is not a measured F1 or recall against ground truth. For pairs above threshold the curve is tighter; for pairs below it, collision probability drops rapidly. This is an honest bound, not a precision measurement.

Measured throughput (benchmark)

The bench-corpus-dedup harness (scripts/bench_corpus_dedup/, .github/workflows/bench-corpus-dedup.yml) measures the tier end-to-end against a real public corpus, with injected ground-truth near-dups so recall is measurable. On a slice of FineWeb (web crawl), the zero-config dedupe_df(df, throughput=…) path measured: docs/sec rises with scale because the wall is dominated by a fixed ~50s of zero-config auto-config at these sizes; the sketch dedup itself is ~9s at 70k (≈7,800 docs/sec raw). Measured recall ≈ 0.43 is the honest LSH recall over a mix of exact / partial-overlap / paraphrase near-dups, the paraphrase and partial dups often do not collide in the LSH bands. (Contrast expected_recall above: the per-pair collision probability at the similarity threshold, not a recall over a real corpus, which is exactly why the bench measures recall directly.) A deterministic per-PR CI perf gate (throughput-gate job, via scripts/bench_corpus_dedup/throughput_perf_gate.py) guards regression on the machine-independent cost, candidate pairs, reduction ratio, and measured recall on a vendored offline corpus, against a committed baseline, so a change can’t quietly blow up the candidate count or drop recall.
Scale note: the published numbers are the validated 14k–70k range. 100k+ corpora currently bottleneck on the survivorship/golden stage; lifting that ceiling (and a datatrove head-to-head) are tracked follow-ups.

Scope

The throughput tier is single-node only. Distributed throughput (#1084) and the product surface (#1085) are follow-on issues. Default-off (throughput=None) is byte-identical to today in every metric.

Configuration precedence

When the same setting is reachable multiple ways, the more specific wins:
  1. Explicit API / CLI argument (dedupe_df(..., backend=...), --backend): highest.
  2. Config field (GoldenMatchConfig(backend=...), YAML).
  3. Environment variable (GOLDENMATCH_*).
  4. Built-in default (auto-config / planner choice): lowest.
So an env var sets the default behavior for a process, a config field overrides it for a run, and an explicit call argument overrides it for one call.

Full environment-variable reference

Every GOLDENMATCH_* variable the package reads (plus the one GOLDEN_* exception, GOLDEN_DIAGNOSTICS), grouped by who should care. Defaults verified against source. Most users only ever touch the recommended setups.

Everyday & scale

Distributed sub-knobs

Perf opt-ins & diagnostics

Auto-config & blocking (advanced)

Probabilistic Fellegi-Sunter (advanced)

A type: probabilistic matchkey trains EM once, then its candidate pairs are scored by one of four paths, chosen automatically. You rarely set these by hand, but this is the map when you need to:
  • bucket (default): memory-bounded field-hash buckets; the #1798-safe route for static/multi_pass blocking on the default backend.
  • external-blocks. Memory-bounded scorer for strategy-generated candidates (lsh/ann/learned/canopy/sorted_neighborhood) that buckets can’t re-derive from field hashes.
  • batched. The legacy per-block path. Reached by GOLDENMATCH_FS_DEFAULT_BUCKET=0 or an explicit non-default backend. Eager build_blocks, so memory grows with dataset size.
  • fused. A single Arrow-native FFI call (block + score + dedup + cluster) under est-RSS pressure on an artifact-free config; 2–19× leaner peak RSS. See Fused kernel auto-routing.
Within whichever path is chosen, the per-block math runs on the native Rust FS kernel by default (GOLDENMATCH_FS_NATIVE), falling back to the vectorized numpy scorer (GOLDENMATCH_FS_VECTORIZED) and then the scalar loop. Every knob below is an escape hatch or an advanced tuning lever, the defaults are the recommended path. Path routing Scoring engine (native vs numpy vs scalar) EM training Model & math Other advanced (scorer / golden / cluster: not FS-specific)

Servers, security, storage, LLM/embedding

Not a GOLDENMATCH_* var but relevant: set OPENAI_API_KEY or ANTHROPIC_API_KEY for LLM scoring and the LLM auto-config fallback, and POLARS_SKIP_CPU_CHECK=1 on Windows to skip the Polars import-time CPU probe.

Backends and scale

The measured numbers behind each backend choice.

Scale envelope

The block-size failure modes that matter more than the backend.

Distributed routing

The planner decides scoring, clustering, and golden routing per stage, keyed off driver RAM (not cluster total). The lint_routing / explain_routing MCP tools surface these decisions and flag overrides that force a slow path; a slow-path override at scale refuses unless allow_slow_path=true.

Single box

With no Ray cluster connected, every stage runs in-memory / in-process. There is nothing to distribute to, so the projections below do not apply.

Driver-RAM projection

When a cluster is present, each materializing stage distributes only when its working set exceeds the driver-RAM budget (available_ram_gb × 0.6):
  • Clustering / WCC distributes when the projected edge set (estimated_pair_count × 16 bytes) exceeds the budget. At 100M rows the 609.4M-pair edge set fits a large head, so WCC runs in-memory (scipy): forcing the distributed path there is a multi-hour tail.
  • Scoring and golden distribute when the row-frame estimate (n_rows × 512 bytes) exceeds the budget. The same data therefore distributes scoring on a 48 GB-driver cluster but keeps it in-memory on a 256 GB box.

Overrides

distributed_routing.* config pins and the legacy GOLDENMATCH_DISTRIBUTED_CLUSTERING_THRESHOLD env var are honored but linted. CLUSTERING_THRESHOLD=0 forces distributed WCC regardless of the projection; at scale the linter flags it ERROR and the runtime refuses unless allow_slow_path=true.