This repository was archived by the owner on Oct 9, 2023. It is now read-only.
Repository navigation
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Snaplet makes use of Copycat in order to turn Personally Identifiable Information input data from your production database into output data that resembles the original value, yet does not allow the original value to be inferred.
If you had a large database though, collisions in these output values became quite likely - in other words, it would be likely that two different input values in your database would share the same output value returned by Copycat. For example, if you had a table with 77,000 rows in it, and you were using
copycat.uuid()for a particular column, there was about a 50% chance of two rows sharing the same output value for that column.In this release, we're using a newer version of copycat (0.6.0) that should make collisions significantly less likely: under the hood, copycat is now using md5 alone for hashing.
Of course, this still depends on the data type. For example, for
copycat.uuid(), collisions are significantly less likely than forcopycat.firstName(), simply because the range of output values is larger.You can expect some more updates ahead with more details about the new collision probabilities for copycat.
All reactions