The 2.0 major is a small, clean break: four removed legacy paths, and the pipeline’s output is unchanged. The real difference between a 1.0 install and a 2.0 install is the entire 1.x feature accumulation below. You are not upgrading into a rewrite; you are upgrading into three months of additive capability with a tidy semver line drawn under it.
What actually breaks in 2.0
Four items, each with a 1.x deprecation runway. Full steps in Migrating to v2.GOLDENMATCH_IDENTITY_ID_SCHEME+ the legacy:hash:lookup bridge are gone. Fingerprintable records resolve to a single canonical:h1:id. If you persist an identity DB, rungoldenmatch identity migrate-ids --path <db>before upgrading. Un-fingerprintable rows keep their:hash:id.GOLDENMATCH_CLUSTER_FRAMES_OUTis gone. The Arrow frames-out clustering path is the only path now (it has been the default and output-equivalent since 1.x). Publicbuild_clustersstill works as a frames-backed adapter.RunHistory.cheapest_healthy()removed. Usepick_committed()._scale_aware_backend(internal shim) removed. The publicauto_configure_df/dedupe_dfAPI routes through the v3 planner.
Where 1.0 started (2026-03-23)
GoldenMatch 1.0 was already a complete, production-stable entity-resolution toolkit, not a preview:- A frozen public API:
gm.dedupe(),gm.match(),gm.pprl_link(),gm.evaluate()with typed results, 96 public exports, 21 CLI commands, a config YAML schema, a REST API + client, MCP tools, and Jupyter_repr_html_()display. - Accuracy as a first-class gate:
goldenmatch evaluate --min-f1 0.90exits non-zero in CI, plusgoldenmatch labelfor interactive ground-truth building. - A Ray distributed backend for large workloads, LLM-assisted scoring, PPRL (privacy-preserving record linkage), domain packs, and match explanations.
What landed across 1.x
Zero-config and the introspective planner
The headline posture of the suite, “find duplicates in 30 seconds with no config,” hardened across 1.x. The v3 controller/planner replaced single-threshold heuristics with a rule table over runtime + complexity profiles that picks the execution plan (backend, scale path, confidence handling) for you. At large row counts it refuses to silently guess, raising rather than committing a low-confidence config. You set nothing for correctness; the planner chooses.Identity Graph (v1.15)
dedupe() returns run-local clusters whose IDs mean nothing across runs. v1.15 (2026-05-12) added a durable, queryable Identity Graph: stable entity_ids (UUIDv7) that survive re-runs, evidence edges that record why two records linked, an append-only event log, and the same JSON view from seven surfaces (Python, CLI, REST, MCP, A2A, SQL, web UI). Off by default; opt in with an identity: config block. See Identity Graph.
Scale: single-box at scale, then distributed 100M
- Arrow-native groundwork (v1.25, 2026-06-01): columnar pair-stream and two-frame cluster representations, plus optional Rust/Arrow-C kernels behind the
goldenmatch[native]wheel, with the pure-Python + Polars pipeline kept as the byte-for-byte reference. - Distributed 100M (v1.26, 2026-06-04): the Phase-5 streaming pipeline (
GOLDENMATCH_DISTRIBUTED_PIPELINE=2) ran a full 100,000,000-row dedupe end to end in ~213s with the driver peaking at 0.30 GB RSS, by removing every driver-side collect. Later validated on a real 5-node GCP cluster (100M in ~9.2 min, clusters recovered exactly).
Native acceleration, signed off
The optional native kernel matured into a signed-off set (clustering, block_scoring, pairs, featurize, hashing) that runs automatically under the default GOLDENMATCH_NATIVE=auto when the wheel is installed, each cleared to bit/ULP parity against the Python reference so flipping native on or off never changes output. See Tuning & opt-ins.
Probabilistic Fellegi-Sunter that beats Splink
Thetype: probabilistic matchkey grew from a scorer into a Splink-class Fellegi-Sunter engine: supervised m estimation, a match-weight waterfall, calibration, and accuracy analysis. By v1.29/1.30 (2026-06-09), zero-training FS auto-config beat hand-rolled, expert-tuned Splink on every dataset Splink scores, on one shared evaluator (historical_50k pairwise F1 0.778 vs 0.757, febrl3 0.991 vs 0.965), made reproducible by an EM training-pair determinism fix.
New surfaces and engines
- SQL-native UDFs for DuckDB, Postgres, and DataFusion (identity resolve/view/history/conflicts/list as in-database functions).
- An opt-in Sail (Spark Connect) distributed tier as a buildable, Sail-native alternative to Ray (
goldenmatch[sail]). - Opt-in WASM acceleration for the TypeScript port, so the edge-safe pure-TS packages can optionally reach the same Rust
*-corekernels.
Which 1.x are you coming from?
- From a recent 1.2x (>= 1.25): the upgrade is essentially just the migration steps. Your identity ids are already
:h1:; you mostly remove two now-ignored env vars. - From an older 1.x (< 1.25) with a persisted identity DB: run
goldenmatch identity migrate-idsfirst, then upgrade. Your DB may still hold legacy:hash:ids that 2.0 will otherwise treat as new entities. - From 1.0 with no identity graph: there is nothing to migrate.
pip install --upgrade goldenmatchand you pick up every capability above; dedupe/match/PPRL/evaluate are unchanged.
See also
- Migrating to v2 — the step-by-step upgrade.
- v1 vs v2 at a glance — the one-screen comparison tables.
- Tuning & opt-ins — every
GOLDENMATCH_*knob in 2.0. - The full per-version detail lives in
packages/python/goldenmatch/CHANGELOG.md.