Semantic Similarity for Code Benchmark
The dataset contains examples of base code that represent know algorithms implemented in serveral programming languages. The base code contains all the symbol names and comments are in some language like English, Spanish, etc. The modified version of the base code are exactly the same base code where all the symbols and comments where translated to some other language different.
For contribution to this project, please follow the documentation in the CONTRIBUTING.md file.