Strategy overview
Static blocking
Group records by exact value of the blocking key.Adaptive blocking
Static blocking with automatic sub-splitting for oversized blocks. When a block exceedsmax_block_size, it splits on the highest-cardinality column within the block.
Sorted neighborhood
Sliding window over records sorted by a key. Catches near-matches that differ by one character in the blocking key.Multi-pass blocking
Run multiple blocking passes and union the results. Best recall for noisy data.ANN hybrid blocking
New in v1.2.6. Combine multi-pass string blocking with ANN fallback for oversized blocks. When a block exceedsmax_block_size and would normally be skipped, GoldenMatch embeds only the unique text values in that block and uses FAISS to create smaller sub-blocks.
- Multi-pass blocking creates string-based blocks (fast, handles most data)
- Blocks exceeding
max_block_sizetrigger ANN fallback instead of being skipped - ANN embeds only unique text values (e.g., 61K records with 187 unique texts = seconds)
- FAISS finds nearest neighbors among unique texts, Union-Find creates sub-blocks
- Sub-blocks still exceeding
max_block_size(after 10x cap) are skipped
GOLDENMATCH_GPU_MODE=vertex) or local sentence-transformers for embedding.
ANN blocking
Use FAISS approximate nearest-neighbor search on sentence-transformer embeddings. Requirespip install goldenmatch[embeddings].
ann_pairs is a faster variant (50—100x) that returns direct pairs instead of block groups:
SimHash (semantic) blocking
New in #1082. Bucket records by semantic near-duplication. SimHash embeds a text column, projects each embedding throughnum_planes random hyperplanes into a 0/1 signature, then bands the signature into LSH buckets. Records whose embeddings are cosine-near collide in a band and become candidates. This catches paraphrases that share meaning but little surface text, where the lexical lsh strategy (word/char-shingle MinHash) only catches shared shingles.
num_bands or threshold (not both required; num_bands wins if you set both). More bands means looser blocking (higher recall, more candidates); fewer bands is tighter. The default embedder is the zero-config in-house model, so SimHash works without external credentials; set model to use a configured provider.
For a text corpus (a column of long free text), auto-config picks simhash automatically when an embedder is reachable, and falls back to the lexical lsh strategy when it is not. So dedupe_df(corpus) selects the semantic near-dup path with no config when embeddings are available. SimHash buckets dense embeddings; the lexical lsh strategy (MinHash) buckets sparse shingle sets. Use SimHash for semantic paraphrase, lsh for lexical near-dups.
High-recall corpus dedup without writing a blocking config. For large text corpora where you want
near-duplicate recall at throughput scale, the throughput tier (
dedupe_df(..., throughput=0.95)) picks
lsh or simhash automatically based on embedder availability, then confirms candidate pairs by sketch
distance — no BlockingConfig required. See Throughput tier.Canopy blocking
TF-IDF-based canopy clustering with loose and tight thresholds.Learned blocking
Data-driven predicate selection via a two-pass approach: sample pairs, train predicates, apply to full data. Achieves 96.9% F1 matching hand-tuned static blocking on DBLP-ACM.MinHash/LSH blocking
New in #1081. Probabilistic sketching for near-duplicate text. Each record’s text column is shingled (char- or word-grams), MinHashed into a signature, and the signature is split into LSH bands; records that share at least one band bucket become candidates. This is the set-similarity (Jaccard) path — the right tool for document/corpus-scale near-duplicate detection, where keyed-predicate strategies (static/multi-pass) are a poor fit.threshold yields more bands (higher recall, more candidate pairs); a
higher one yields fewer (more precision, fewer pairs).
LSH measures lexical overlap, not semantic similarity. It excels at
near-duplicates that share words/characters (boilerplate, reposts, lightly
edited copies). It does not catch paraphrases that mean the same thing in
different words — on the Quora Question Pairs benchmark (semantic duplicates)
it recovers only ~21% of labeled pairs at
threshold: 0.5, versus ~98% on a
lexical-near-duplicate set. For semantic matching, use the ann strategy
(embedding nearest-neighbor).goldenmatch.core.sketch) is shared byte-for-byte across
Python, the optional
native extension, and the TypeScript port. Recall is measured by an always-on
synthetic gate plus a Quora Question Pairs benchmark (bench-lsh-recall.yml).
Auto-select
Let GoldenMatch pick the best blocking key by histogram analysis:Performance impact
Blocking key choice dominates fuzzy matching performance. A coarse key (e.g.,state) creates huge blocks and slow scoring. A fine key (e.g., email) misses near-duplicates.
Rules of thumb:
- Target max block size under 1,000 records
- Use
multi_passfor best recall,adaptivefor best speed - Use
learnedto auto-discover optimal predicates - Use
ann_pairsfor semantic/product matching