Gr

Greenmask 0.2 – Database anonymization tool

Hacker News

Greenmask 0.2 – Database anonymization tool

Hi! My name is Vadim, and I’m the developer of Greenmask ( https://github.com/GreenmaskIO/greenmask ). Today Greenmask is almost 1 year and recently we published one of the most significant release with new features: https://github.com/GreenmaskIO/greenmask/releases/tag/v0.2.0 , as well as a new website at https://greenmask.io . Before I describe Greenmask’s features, I want to share the story of how and why I started implementing it. Everyone strives to have their staging environment resemble production as closely as possible because it plays a critical role in ensuring quality and improving metrics like time to delivery. To achieve this, many teams have started migrating databases and data from production to staging environments. Obviously this requires anonymizing the data, and for this people use either custom scripts or existing anonymization software. Having worked as a database engineer for 8 years, I frequently struggled with routine tasks like setting up development environments—this was a common request. Initially, I used custom scripts to handle this, but things became increasingly complex as the number of services grew, especially with the rise of microservices architecture. When I began exploring tools to solve this issue, I listed my key expectations for such software: documentation; type safety (the tool should validate any changes to the data); streaming (I want the ability to stream the data while transformations are being applied); consistency (transformations must maintain constraints, functional dependencies, and more); reliability; customizability; interactivity and usability; simplicity. I found a few options, but none fully met my expectations. Two interesting tools I discovered were pganonymizer and replibyte. I liked the architecture of Replibyte, but when I tried it, it failed due to architectural limitations. With these thoughts in mind, I began developing Greenmask in mid-2023. My goal was to create a tool that meets all of these requirements, based on the design principles I laid out. Here are some key highlights: * It is a single utility - Greenmask delegates the schema dump to vendor utilities and takes responsibility only for data dumps and transformations. * Database Subset ( https://docs.greenmask.io/latest/database_subset ) - specify the subset condition and scale down size. We did a deep research in graph algorithms and now we can subset almost any complexity of database. * Database type safety - it uses the DBMS driver to decode and encode data into real types (such as int, float, etc.) in the stream. This guarantees consistency and almost eliminates the chance of corrupted dumps. * Deterministic engine ( https://docs.greenmask.io/latest/built_in_transformers/trans... ) - generate data using the hash engine that produces consistent output for the same input. * Dynamic parameters for transformers ( https://docs.greenmask.io/latest/built_in_transformers/dynam... ) - imagine having created_at and updated_at dates with functional dependencies. Dynamic parameters ensure these dates are generated correctly. We are actively maintaining the current project and continuously improving it—our public roadmap at https://github.com/orgs/GreenmaskIO/projects/6 . Soon, we will release a Python library along with transformation collections to help users develop their own complex transformations and integrate them with any service. We have plans to support more database engines, with MySQL being the next one, and we are working on tools which will integrate seamlessly with your CI/CD systems. To get started, we’ve prepared a playground for you that can be easily set up using Docker Compose: https://docs.greenmask.io/latest/playground/ I’d love to hear any questions or feedback from you!

Share card

Actual performance

94points
19comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: created, started, para · Missing: supports, reddit linkedin, podcasting
90%90% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: user, dock, new · Missing: mac, agents, macos
78%78% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, existing, io · Missing: https docs, excited, just released
73%73% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Strong signals: soon, users · Missing: plus, platform, intuitive
51%51% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: users, para · Missing: mobile apps, ios, personal
40%40% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Up
Uphold – Tool for programmatically verifying database backups67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Uphold – Tool for programmatically verifying database backups

Hacker News89
GraphMyDB
GraphMyDB33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

database-tool

Product Hunt
DatabaseFileRecovery SQLite Database Recovery Tool
DatabaseFileRecovery SQLite Database Recovery Tool53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Repairs and recovers Corrupt SQLite Database

Indie Hackerscommitment-full-time
Su
SummaDB, a hierarchical database that syncs with PouchDB60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SummaDB, a hierarchical database that syncs with PouchDB

Hacker News5
No
Noms – The versioned, forkable, syncable database68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Noms – The versioned, forkable, syncable database

Hacker News43
A
A reactive Database61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A reactive Database

Hacker News1
A
A reactive Database61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A reactive Database

Hacker News1
Re
Relations in a NoSQL database68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Relations in a NoSQL database

Hacker News6
Di
Digestable Ingredient Database68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Digestable Ingredient Database

Hacker News2
Ne
NebulaDB, my first attempt at a database63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

NebulaDB, my first attempt at a database

Hacker News30