Da

Data Caterer – Data testing tool for any data source

Hacker News

Data Caterer – Data testing tool for any data source

Hi everyone, I'm Peter. A Software Engineer who has been around the data space for the majority of my career. Data Catering is a tool I've been itching to try and implement as I've spent lots of time debugging data issues, generating manual data to test jobs/services or replicating production-like data flows. The usual way to solve this problem is to stream/load a subset of production data, mask it and store it in a pre-prod environment. This is okay but can increase the risk of personal data leaking as pre-prod environments may not have the same security measures in place (Optus data leak as an example https://en.wikipedia.org/wiki/2022_Optus_data_breach ), having network connections between prod and non-prod being opened, and it may not cover all data scenarios (different permutations/combinations of values, unknown unknowns). Another approach is to generate data yourself. There are a number of existing tools that can aid with data generation but lack the ability to clean up the generated data, generate both batch and event data, define relationships between datasets (for example, an account can have multiple transactions associated with it, or foreach account create event, the same account should exist in a CSV file), or validating data. With this in mind and following along the ideas of having a tool that can be run anywhere, ability to automatically discover, is metadata driven, generate and validate data, and giving users the ability to customise how it is run, Data Caterer was created. Main features include: - Metadata discovery - Batch or event data generation - Maintain referential integrity across any dataset - Create custom data generation scenarios - Clean up generated data - Validate data - Suggest data validations In terms of technical details, it is a Spark based application where users can use the Scala/Java API or YAML files to interface with it. It is a single docker image that can be run following the quick start found here ( https://data.catering/get-started/docker/ ). If you want to find out more details, check out the main website here ( https://data.catering/ ). I believe there is room for improvement in the data testing area as we should look to be more proactive in detecting and resolving bugs related to data quality or system integration. Any feedback is appreciated.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: created, started, ios · Missing: supports, reddit linkedin, podcasting
92%92% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: user, dock, single · Missing: mac, agents, macos
65%65% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, existing, ide · Missing: https docs, excited, just released
58%58% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: interface, users · Missing: plus, platform, intuitive
40%40% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: ios, personal, users · Missing: mobile apps, entrepreneurs, apps
37%37% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
10%10% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Da
Data Catering – Data generation and testing tool55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Data Catering – Data generation and testing tool

Hacker News1
To
Tool to sanitize data from Java heap dumps37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tool to sanitize data from Java heap dumps

Hacker News2
I
I made a data agent for hypothesis testing25%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made a data agent for hypothesis testing

Hacker News3
Op
Open-source Data Anonymization tool nxs-data-anonymizer62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source Data Anonymization tool nxs-data-anonymizer

Hacker News4
Cs
CsCheck 0.9.0 – regression testing without data files37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CsCheck 0.9.0 – regression testing without data files

Hacker News2
We
We added an Oink data importer for Cheers37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We added an Oink data importer for Cheers

Hacker News9
Te
Techcrunch data 2005-201250%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Techcrunch data 2005-2012

Hacker News2
Ma
Mambocollector – Statsd Data collector for MySQL44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mambocollector – Statsd Data collector for MySQL

Hacker News2
Mu
Munge your data with TXR50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Munge your data with TXR

Hacker News2
Ag
AgriCatch – Data aggregation on Django43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AgriCatch – Data aggregation on Django

Hacker News4