Candidate-relative ranks under blocker changes
Retained-pair and blocker-selection results from the Savant entity-matching benchmark.
Abstract
One of Savant's entity-matching representations uses the empirical rank of a similarity measurement rather than the measurement alone. Because those ranks are computed from the current candidate set, changing the blocker can change a retained pair's coordinates even when its raw measurements have not changed.
We measured this on public entity-matching benchmarks by rerunning the blocker and rescoring the intersection of labeled pairs present in both runs. With the final fuzzy-candidate cap set to 2, between 4.4% and 31.0% of retained decisions changed across five benchmark families. On DBLP-ACM, pair completeness remained 0.998 and 25.2% of decisions on retained pairs changed. Source-fixed ranks were unchanged on those pairs.
1. Background and setup
Blocking and filtering are standard parts of entity resolution because the full Cartesian pair space is usually too large to score. Conditioning the linkage problem on the pairs that pass an indexing or blocking rule is already well understood; Murray treats this explicitly for Fellegi-Sunter linkage [2], and Papadakis et al. survey the broader blocking and filtering literature [1]. Recent work such as SMBench also evaluates filters together with the candidate sets that downstream matchers actually receive [3].
The issue here arose from Savant's empirical-rank representation. The first signal was an earlier candidate-thinning experiment, selection_confounding. All observed positive edges were kept and about half of the negatives were removed. WDC rank-F1 rose from 0.374 to 0.435; DBLP-ACM fell from 0.791 to 0.588. That experiment was useful as a warning but not as an isolation, since thinning changed both which pairs existed and the distribution used to calculate the ranks.
The later blocker-observation-channel run scores only the labeled-pair intersection of a source candidate graph and a graph produced by a modified blocker. Semantic and lexical measurements are taken from the source run and reused, as are the fitted matcher and threshold. What changes in the local-rank arm is the empirical CDF used by the rank transform. We also score two controls on the same pair intersection: the raw measurements, and ranks calculated against the source CDF in both arms.
Development data are used for fitting and calibration data for thresholds and experimental selection. Holdout labels are used for scoring; labels are not used to construct the CDFs. The signed run completed on DBLP-ACM, WDC, Abt-Buy, Walmart-Amazon, Amazon-Google, iTunes-Amazon, Beer, and Fodors-Zagats. DBLP-Scholar completed the earlier cap-only run but not the longer signed run, which repeatedly failed at the workspace execution boundary.
2. Retained-pair results
We first reduced the final fuzzy-candidate cap to 8, 4, and 2. At cap 2:
| Family | Pair completeness | Rank representations changed | Decisions changed |
|---|---|---|---|
| DBLP-ACM | 0.998 | 100% | 25.2% |
| DBLP-Scholar | 0.994 | 100% | 26.0% |
| Abt-Buy | 0.874 | 100% | 31.0% |
| WDC | 0.606 | 100% | 22.2% |
| Walmart-Amazon | 0.990 | 100% | 4.4% |
The progression with tighter caps is visible on several datasets. DBLP-ACM changes 5.4% of retained decisions at cap 8, 14.6% at cap 4, and 25.2% at cap 2. DBLP-Scholar gives 3.8%, 11.7%, and 26.0%; Abt-Buy gives 9.2%, 18.6%, and 31.0%.
For interpretation, DBLP-ACM is more useful than WDC in the cap-2 run. WDC's pair completeness has already fallen to 0.606, so the blocker is doing two large things at once: dropping support and moving the local rank reference. DBLP-ACM retains almost every labeled match and still changes one quarter of the decisions made on pairs present in both runs.
Across the control arms, feature displacement, probability displacement, and decision changes were zero on the retained pairs. The local-rank arm crossed the fixed threshold for the rates reported above.
Conditional F1 at cap 2 was:
| Family | Source-reference F1 | Modified-reference F1 | Δ F1 |
|---|---|---|---|
| DBLP-ACM | 0.893 | 0.775 | -0.118 |
| WDC | 0.403 | 0.050 | -0.353 |
| Abt-Buy | — | — | -0.288 |
| Walmart-Amazon | — | — | -0.036 |
Outside the final-cap series, WDC includes positive conditional-F1 changes: approximately +0.040 for fuzzy-filter-100pct, +0.039 for sparse-query-tokens-4, and +0.025 for sparse-candidate-cap-1. Local ranks therefore remain part of the representation set used in Savant experiments.
At WDC cap 2, lexical-only substitution changes conditional F1 by -0.346, compared with -0.353 when all target references are substituted. On DBLP-ACM, moving the semantic and lexical references together reproduces the full result to within a few thousandths. Abt-Buy cap 4 gives +0.017 for semantic-only and +0.033 for lexical-only, whereas their joint substitution gives the full -0.037 change. Walmart-Amazon cap 2 closes only after the value reference is included with semantic and lexical.
3. What the blocker did to the observed non-match population
A separate benchmark, blocker-selection-geometry, records semantic and lexical measurements under a broad blocker and then applies narrower blocker policies to those already measured pairs. It is not a second feature-extraction pass; the measurements in the following tables are frozen before selection.
| WDC survival set | Spearman rho, labeled non-matches |
|---|---|
| Broad reference | +0.280 |
| 100% fuzzy blocks + cap 16 | +0.268 |
| Default 80% fuzzy blocks + cap 16 | +0.191 |
| 40% fuzzy blocks + cap 16 | +0.115 |
| DBLP-ACM survival set | Spearman rho, labeled non-matches |
|---|---|
| Broad reference | +0.260 |
| 100% fuzzy blocks + cap 16 | +0.274 |
| Default 80% fuzzy blocks + cap 16 | +0.268 |
| 40% fuzzy blocks + cap 16 | +0.213 |
On WDC, most of the movement appears when the per-record fuzzy-block filter is tightened. Altering only the final top-k changes this particular statistic much less. The same cap reduction also has a very different recall cost across the two datasets: WDC goes from 12,195 candidates and 0.932 pair completeness at cap 16 to 4,884 and 0.606 at cap 2, while DBLP-ACM goes from 57,851 candidates and about 1.000 pair completeness to 13,709 and 0.998.
4. Additional experiments
Selection metadata had already been added to candidate edges: pinning, pre-cap rank, candidate pressure, and one-sided admission. Adding these variables to the rank baseline changed held-out F1 by +0.037 on WDC, +0.079 on DBLP-ACM, +0.051 on Abt-Buy, +0.003 on Walmart-Amazon, -0.025 on Beer, and -0.006 on iTunes-Amazon. We also tried a direct fuzzy-block-censoring scalar after the WDC selection runs. Relative to that provenance model its deltas were -0.008, +0.004, +0.003, and 0.000 on WDC, DBLP-ACM, Abt-Buy, and Beer. We did not keep that scalar in the matcher.
The fixed-reference comparison is more consequential for the implementation:
| Dataset / representation | Default F1 | Cap-2 F1 |
|---|---|---|
| DBLP-ACM local rank | 0.906 | 0.803 |
| DBLP-ACM fixed rank | 0.872 | 0.882 |
| WDC local rank | 0.459 | 0.504 |
| WDC fixed rank | 0.372 | 0.433 |
At DBLP-ACM cap 2, fixed rank scores 0.882 conditional F1 versus 0.803 for local rank. On WDC, the corresponding values are 0.433 and 0.504. A later model exposed fixed ranks together with local-minus-fixed displacement variables and interaction terms. On the unseen Abt-Buy cap-2 run it scored 0.504 conditional F1; the fixed-rank model scored 0.686. Clamping the displacement variables to ranges observed during development and calibration left the factorized score at 0.504.
Reference storage
Our first storage experiment sampled candidate rows from the reference. Some calibration-selected samples looked acceptable and still exceeded the holdout F1 tolerance on Abt-Buy and Walmart-Amazon, despite decision agreement above 99%.
The rank transform consumes independent marginal empirical CDFs and never uses joint candidate rows. The next implementation stored run-length marginal CDFs directly, followed by optional quantization. Calibration selected the smallest tested summaries with at least 0.99 worst decision agreement and no more than 0.005 absolute F1 difference from the exact reference.
| Family | Exact CDF points | Selected points | Holdout agreement | Max |Δ F1| |
|---|---|---|---|---|
| DBLP-ACM | 104,775 | 592 | 0.9907 | 0.0023 |
| WDC | 69,514 | 649 | 0.9958 | 0.0020 |
| Abt-Buy | 35,213 | 729 | 0.9945 | 0.0036 |
| Walmart-Amazon | 132,527 | 627 | 0.9942 | 0.0032 |
I would not put much weight on the individual signed F1 magnitudes for Beer or Fodors-Zagats because both have few positive examples. DBLP-Scholar is simply missing from the longer signed panel. The retained-pair effect measured here is also a property of representations that use the surrounding candidate population; a pairwise feature vector that never looks at other candidates does not acquire this particular rank-reference dependency.
5. Implementation
The production change is limited to the fixed-reference path. When such a matcher is constructed, Savant stores the empirical reference as part of matcher state. EntityMatcherAuthority includes the reference CDF, semantic sensor identity, source provenance, feature contract, fitted model, and threshold. Exact CDFs are the default and quantization is explicit. Local ranks remain available.
References
- Papadakis, Skoutas, Thanos & Palpanas. Blocking and Filtering Techniques for Entity Resolution: A Survey. ACM Computing Surveys 53(2), 2020.
- Murray. Probabilistic Record Linkage and Deduplication after Indexing, Blocking, and Filtering. Journal of Privacy and Confidentiality 7(1), 2016.
- Astappiev, Neuhof, Fisichella & Papadakis. SMBench: No-code benchmarking of learning-based entity matching. Information Systems 139, 2026.
- Ruan, Shi & Bauernhansl. Fine-tuning large language models with contrastive margin ranking loss for selective entity matching in product data integration. Advanced Engineering Informatics 67, 2025.
Benchmark values in this note come from the Savant research harness. Development data are used for fitting, calibration data for threshold and experiment selection, and holdout data for scoring.