Se

SemHash – Fast Semantic Text Deduplication for Cleaner Datasets

Hacker News

SemHash – Fast Semantic Text Deduplication for Cleaner Datasets

We’ve just open-sourced SemHash, a lightweight package for semantic text deduplication. It lets you effortlessly clean up your datasets and avoid pitfalls caused by duplicate samples in semantic search, RAG, and machine learning. Main Features: - Fast and hardware friendly: Deduplicate datasets with millions of records in minutes, on a CPU. - Flexible: Works on single or multiple datasets (e.g., train/test deduplication), and multi-column data (e.g., Question-Answering datasets). - Lightweight: Minimal dependencies (largest is NumPy). - Explainable: Easily inspect duplicates and what caused them, and view the lowest similarity duplicates to adjust the threshold based on your dataset. We found that text deduplication is more complex than it appears, so we built SemHash to simplify the process. Duplicate samples can skew model training, reduce generalization, and cause train-test leakage—leading to unreliable results. Techniques like minhash handle exact or near-exact duplicates, but semantic deduplication also catches semantically redundant samples, which we believe is an important aspect of deduplication. Furthermore, it’s not trivial to see why something was removed with minhash, which we also believe is important. We already found some interesting results on some well known datasets in our benchmarks which are included in the repo. We are curious to hear your feedback! Do you currently deduplicate your datasets before training, and what techniques do you use?

Share card

Actual performance

6points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, single · Missing: agents, macos, agent
74%74% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: io · Missing: https docs, excited, just released
72%72% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
59%59% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: friendly · Missing: plus, platform, intuitive
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
34%34% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
23%23% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Se
Semantic Segmentation Datasets61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Semantic Segmentation Datasets

Hacker News1
Se
Semantic Text62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Semantic Text

Hacker News1
Ca
Cached Datasets53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cached Datasets

Hacker News4
Sc
SchemaVer for semantic versioning of schemas64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SchemaVer for semantic versioning of schemas

Hacker News1
Ultrasonic Cleaner
Ultrasonic Cleaner15%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ultrasonic Cleaner

Indie Hackers
Te
Tensorpack a CLI tool for semantic discovery across datasets52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tensorpack a CLI tool for semantic discovery across datasets

Hacker News1
I
I made this tool for navigating pandas datasets50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made this tool for navigating pandas datasets

Hacker News20
Ge
Geckoboard Datasets API58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Geckoboard Datasets API

Hacker News1
CleanGeek
CleanGeek11%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Windows Cleaner With No Registry Cleaner

Indie Hackers
Se
Semantic Text Analysis as a service57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Semantic Text Analysis as a service

Hacker News5