Fragmenting is naive and inefficient - blockers should be able to rely on the entlet's entire underlying data structure to determine similarity and block accordingly (which should remove the need for blocking entirely).
The blocker should cluster by ER fields by default but permit additional fields to be considered for blocking only
Custom fingerprinting will be needed to handle nested structures.
"v2" (make another issue) should permit training the blocker with entlet pairs of good/bad matches.
This implies a rework of post-blocking stages' data structures.
Fragmenting is naive and inefficient - blockers should be able to rely on the entlet's entire underlying data structure to determine similarity and block accordingly (which should remove the need for blocking entirely).
The blocker should cluster by ER fields by default but permit additional fields to be considered for blocking only
Custom fingerprinting will be needed to handle nested structures.
"v2" (make another issue) should permit training the blocker with entlet pairs of good/bad matches.
This implies a rework of post-blocking stages' data structures.