Skip to main content
Blocking reduces the comparison space from O(N^2) to O(N*B) by grouping records that share a key. GoldenMatch supports 10 strategies.

Strategy overview

Static blocking

Group records by exact value of the blocking key.
Multiple keys produce independent blocks that are unioned. Transforms are applied before grouping.

Adaptive blocking

Static blocking with automatic sub-splitting for oversized blocks. When a block exceeds max_block_size, it splits on the highest-cardinality column within the block.

Sorted neighborhood

Sliding window over records sorted by a key. Catches near-matches that differ by one character in the blocking key.

Multi-pass blocking

Run multiple blocking passes and union the results. Best recall for noisy data.

ANN hybrid blocking

New in v1.2.6. Combine multi-pass string blocking with ANN fallback for oversized blocks. When a block exceeds max_block_size and would normally be skipped, GoldenMatch embeds only the unique text values in that block and uses FAISS to create smaller sub-blocks.
How it works:
  1. Multi-pass blocking creates string-based blocks (fast, handles most data)
  2. Blocks exceeding max_block_size trigger ANN fallback instead of being skipped
  3. ANN embeds only unique text values (e.g., 61K records with 187 unique texts = seconds)
  4. FAISS finds nearest neighbors among unique texts, Union-Find creates sub-blocks
  5. Sub-blocks still exceeding max_block_size (after 10x cap) are skipped
On the Bulldozer dataset (401K rows), this recovered 363 sub-blocks from 15 oversized blocks that would otherwise be skipped, matching 949 additional records. Requires Vertex AI (GOLDENMATCH_GPU_MODE=vertex) or local sentence-transformers for embedding.

ANN blocking

Use FAISS approximate nearest-neighbor search on sentence-transformer embeddings. Requires pip install goldenmatch[embeddings].
ann_pairs is a faster variant (50—100x) that returns direct pairs instead of block groups:

SimHash (semantic) blocking

New in #1082. Bucket records by semantic near-duplication. SimHash embeds a text column, projects each embedding through num_planes random hyperplanes into a 0/1 signature, then bands the signature into LSH buckets. Records whose embeddings are cosine-near collide in a band and become candidates. This catches paraphrases that share meaning but little surface text, where the lexical lsh strategy (word/char-shingle MinHash) only catches shared shingles.
Provide either num_bands or threshold (not both required; num_bands wins if you set both). More bands means looser blocking (higher recall, more candidates); fewer bands is tighter. The default embedder is the zero-config in-house model, so SimHash works without external credentials; set model to use a configured provider. For a text corpus (a column of long free text), auto-config picks simhash automatically when an embedder is reachable, and falls back to the lexical lsh strategy when it is not. So dedupe_df(corpus) selects the semantic near-dup path with no config when embeddings are available. SimHash buckets dense embeddings; the lexical lsh strategy (MinHash) buckets sparse shingle sets. Use SimHash for semantic paraphrase, lsh for lexical near-dups.
High-recall corpus dedup without writing a blocking config. For large text corpora where you want near-duplicate recall at throughput scale, the throughput tier (dedupe_df(..., throughput=0.95)) picks lsh or simhash automatically based on embedder availability, then confirms candidate pairs by sketch distance — no BlockingConfig required. See Throughput tier.

Canopy blocking

TF-IDF-based canopy clustering with loose and tight thresholds.

Learned blocking

Data-driven predicate selection via a two-pass approach: sample pairs, train predicates, apply to full data. Achieves 96.9% F1 matching hand-tuned static blocking on DBLP-ACM.
Cache the learned rules to skip re-training on subsequent runs.

MinHash/LSH blocking

New in #1081. Probabilistic sketching for near-duplicate text. Each record’s text column is shingled (char- or word-grams), MinHashed into a signature, and the signature is split into LSH bands; records that share at least one band bucket become candidates. This is the set-similarity (Jaccard) path — the right tool for document/corpus-scale near-duplicate detection, where keyed-predicate strategies (static/multi-pass) are a poor fit.
A lower threshold yields more bands (higher recall, more candidate pairs); a higher one yields fewer (more precision, fewer pairs).
LSH measures lexical overlap, not semantic similarity. It excels at near-duplicates that share words/characters (boilerplate, reposts, lightly edited copies). It does not catch paraphrases that mean the same thing in different words — on the Quora Question Pairs benchmark (semantic duplicates) it recovers only ~21% of labeled pairs at threshold: 0.5, versus ~98% on a lexical-near-duplicate set. For semantic matching, use the ann strategy (embedding nearest-neighbor).
Empty / whitespace-only text rows have nothing to sketch and are skipped. The underlying kernel (goldenmatch.core.sketch) is shared byte-for-byte across Python, the optional native extension, and the TypeScript port. Recall is measured by an always-on synthetic gate plus a Quora Question Pairs benchmark (bench-lsh-recall.yml).

Token blocking

New in #2488. An inverted index over the tokens of a free-text column. Each record is indexed under every token it contains, so a record belongs to many blocks and two records become candidates when they share any surviving token. This is the one strategy here that is not a single-key scheme. Every keyed strategy (static, adaptive, multi_pass, …) puts a record in one block per pass, so a true pair that disagrees on the derived key is lost before scoring. On free-text product titles that is most pairs: on Amazon-Google the key zero-config commits reaches only 7.15% of the ground-truth pairs.
Cost is controlled by document-frequency pruning, not a block-size cap. A token in D records forms a block of D and contributes D(D-1)/2 pairs, so the expensive tokens are the frequent ones — and frequency is exactly what makes a token non-discriminative. Dropping them removes nearly all the cost and almost none of the recall. When max_df is unset it is derived as max_df_ratio * n clamped to [10, 1000], so the cap stays bounded at any frame size. Measured on Amazon-Google (4589 records, 1300 ground-truth pairs):
Token blocking complements lsh; it does not replace it. MinHash/LSH estimates Jaccard similarity over the whole token set, so it needs the two strings to overlap substantially. Cross-vendor product titles routinely share two or three discriminative tokens out of fifteen — strong evidence, low Jaccard — which is why LSH peaks at 55.85% on the table above. Use lsh for near-duplicate documents, token for short free text where only a few tokens agree.
Blocking auto-suggest will not choose token by default. Committing it automatically took Amazon-Google’s end-to-end F1 from 0.1014 to 0.0000 — the failing auto-config subprofile flips to blocking and the committed config yields no candidate pairs. That is undiagnosed (#2488), so the strategy is available from an explicit config as shown above, while auto-suggest stays behind GOLDENMATCH_TOKEN_BLOCKING=1.

Auto-select

Let GoldenMatch pick the best blocking key by histogram analysis:
The analyzer scores each key by block count, max block size, estimated comparisons, and recall. Use the CLI to see suggestions:

Performance impact

Blocking key choice dominates fuzzy matching performance. A coarse key (e.g., state) creates huge blocks and slow scoring. A fine key (e.g., email) misses near-duplicates. Rules of thumb:
  • Target max block size under 1,000 records
  • Use multi_pass for best recall, adaptive for best speed
  • Use learned to auto-discover optimal predicates
  • Use ann_pairs for semantic/product matching