Se

SemHash – Semantic Text Deduplication, Outlier Filtering and Sampling

Hacker News

SemHash – Semantic Text Deduplication, Outlier Filtering and Sampling

We’ve just released SemHash v0.3.0, a major rework of our open-source text pre-processing library. We’ve added two new functionalities: outlier filtering & representative sampling. The core API has been reworked to make sure all of these features can be used together in an intuitive way. Our new features use the existing approximate nearest neighbors index that we already used for semantic deduplication, so they can be ran very quickly after building the index on your dataset. The core package can now be used for: - Semantic Deduplication: Remove semantic duplicates from your dataset. This can prevent train/test set overlap in classification tasks, or prevent duplicate samples in RAG/semantic search. - Outlier Filtering: Surface and filter the most anomalous samples from your dataset. This can help with automated removal of low quality data, or data that should not be in your dataset. - Representative Sampling: Select the most central and diverse examples using Maximal Marginal Relevance. This can help you quickly explore and understand a dataset, or even build a small, diverse, high quality dataset, for example for LLM finetuning. We’ve designed these features in the same way as our semantic deduplication: CPU friendly, lightweight, and explainable. We hope these features help you create cleaner datasets, or simply understand your data better. We’re curious to hear your feedback, and whether there are any other features you think would improve SemHash further!

Share card

Actual performance

7points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: new, tasks, using · Missing: mac, agents, macos
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Hacker NewsStrong engagement from HN community · Strong signals: just released, exist, existing · Missing: https docs, excited, lua
64%64% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Strong signals: intuitive, friendly · Missing: plus, platform, reviews
56%56% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
53%53% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
44%44% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: margin · Missing: arr, mrr, revenue
12%12% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Se
Semantic Text62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Semantic Text

Hacker News1
Hi
Hierarchical Filtering on Elasticsearch59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hierarchical Filtering on Elasticsearch

Hacker News7
Sc
SchemaVer for semantic versioning of schemas64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SchemaVer for semantic versioning of schemas

Hacker News1
A
A vector database with semantic SQL-like filtering56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A vector database with semantic SQL-like filtering

Hacker News5
Gl
Globset for Nix Source Filtering68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Globset for Nix Source Filtering

Hacker News7
Re
Restoring Selections in ContentEditable DIVs while Filtering Content35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Restoring Selections in ContentEditable DIVs while Filtering Content

Hacker News2
IM
IMQuickSearch – filtering your NSArrays of NSObjects like a boss.44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

IMQuickSearch – filtering your NSArrays of NSObjects like a boss.

Hacker News3
Lo
Logdy v0.14 – Semantic log filtering now available49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Logdy v0.14 – Semantic log filtering now available

Hacker News1
Se
Semantic Text Analysis as a service57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Semantic Text Analysis as a service

Hacker News5
Se
Semantic Image Rabbit Hole60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Semantic Image Rabbit Hole

Hacker News3