No

Norma – build good datasets (using an objective)

Hacker News

Norma – build good datasets (using an objective)

My team has worked for F500s, startups, and everything in between. In every case, we found it almost impossible to assemble an ideal dataset for training models. In real-world systems, the information you actually need is scattered across 30–300+ tables, stored in different warehouses, parquets, CSVs, and legacy DBs that nobody fully understands anymore. We realized the real job isn’t ETL (too wide), or feature engineering (too narrow) it’s constructing the ideal representation of the problem so downstream models can actually learn something meaningful. So we built Norma, an optimization-first data platform. It does the things every ML team wishes their stack would do: 1. Unity Catalog integration that works out of the box - connect a warehouse, instantly browse tables with lineage, schemas, and metadata. 2. A unified SQL/Python pipeline engine - both languages execute in the same memory buffer (via DuckDB), so no more glue code or brittle data hops. 3. An AI assistant for transformations - ask for a feature, a join, an explanation, a visualization (generates pipeline steps). 4. Multi-bandit 5-fold cross-validation - fast, automatic evaluation of transformed datasets with xgboost. 5. Visual lineage + shared datasets - every step is inspectable, reproducible, and sharable across teams. That’s what we have today. We’re still building: - Automatic leakage detection (timestamp violations, post-outcome signals, unsafe joins) - Relevant table discovery (find the tables that actually matter for predicting your target) - Relevant row selection (especially for PFN-style models with row limits) - Automated feature representation (scaling, encoding, aggregation, embeddings) - AutoGluon + TabPFN integration (train strong models on normalized, optimized datasets) - Differential privacy guardrails for LLM usage inside your data workflows We’re trying to build the equivalent of a representation compiler: raw warehouse → optimal feature space → any model or BI tool. If you’ve ever lost days hunting through a schema, debugging leakage, redoing feature pipelines, or trying to understand why a model plateaus even though your data is “fine,” I’d genuinely love your feedback. We’re still working closely with teams to refine our features and capabilities, and we’d love to share a private beta with your team. Please join the waitlist! Happy to answer anything here.

Share card

Actual performance

3points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, visual · Missing: mac, agents, macos
92%92% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
75%75% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, ide, pipe · Missing: https docs, excited, just released
50%50% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform · Missing: plus, intuitive, reviews
34%34% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
32%32% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr, training · Missing: mrr, revenue, profit
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ca
Cached Datasets53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cached Datasets

Hacker News4
El
Elbi – good on the go57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Elbi – good on the go

Hacker News5
ES
ES-6/next/whatever. Whats in it, whats good, and whats shit53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ES-6/next/whatever. Whats in it, whats good, and whats shit

Hacker News6
Go
Good-Enough Golfers44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Good-Enough Golfers

Hacker News1
Ar
Are You a Good Estimator?44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Are You a Good Estimator?

Hacker News15
Co
Cotopaxi – Gear for Good44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cotopaxi – Gear for Good

Hacker News59
CodeFaster
CodeFaster40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Get good

Indie Hackers4content
Good Twinkle Lights
Good Twinkle Lights41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Most twinkle lights suck. These are good.

Indie Hackers1$350/moadvertising
Horseshoe Good Luck Necklace with Initial & Births
Horseshoe Good Luck Necklace with Initial & Births14%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Horseshoe Good Luck Necklace with Initial & Birthstone Charm

Indie Hackers
I
I made this tool for navigating pandas datasets50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made this tool for navigating pandas datasets

Hacker News20