Skip hub nodes and include block nodes in sampling#60
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR introduces two major improvements to the sampling algorithm:
Ability to skip hub nodes: Hub nodes (e.g., exchange or mixer nodes) are characterized by a highly significant in/out-degree. Because two nodes connected through a hub are not necessarily related, including a hub node and its subsequent connections can artificially inflate and mislead the sampled neighborhood. To address this, we added an option to define a traversal stopping criterion: the algorithm will no longer expand beyond a node if its in/out-degree exceeds a specified threshold.
Automatic block node inclusion: In domains like finance,
Blocknodes provide crucial macroeconomic context for a transaction. Pairing a transaction (Tx) node with its respectiveBlocknode yields richer features for downstream learning tasks. This PR adds a configuration option ensuring that everyTxnode sampled into a neighborhood automatically brings its correspondingBlocknode with it.Additionally, this PR renames the sampling algorithm from
Forest FiretoPanorama. Our implementation has evolved far beyond the originalForest Fireconcept, and updating the name prevents confusion moving forward.