Ti

Tidepool – analytics for large text datasets

Hacker News

Tidepool – analytics for large text datasets

Hello HN! I'm Peter, one of the folks who helped create Tidepool. We last shared Tidepool with HN about 7 months ago https://news.ycombinator.com/item?id=36957762 Since then, the AI field has moved incredibly quickly and we’ve iterated a lot on our product! The core problem we are trying to solve is: there's a lot of useful business insights you can get from text data, but it's hard to do analytics on it. - SQL is built for tabular / structured data, but when it comes to text, the best you can do is do keyword search. - In the pre-LLM world, you might resort to training a lightweight text classifier, but you have to manually label a lot of data to get good accuracy. This means managing a team of operations people to do the labeling and building a lot of infrastructure to train and deploy the model. - Today it’s easier to get good results with simple LLM prompts, but the "large" in "large language models" means that running on big production-level datasets becomes prohibitively expensive. Our solution to this is to impose structure on this unstructured data with a combination of LLMs and lightweight embedding classifiers. - Using Tidepool, a user can query the data by creating an "attribute." An attribute is a characteristic of the data that you want to analyze, defined in natural language. This could be “sentiment of reviews,” “messages mentioning legal topics,” “prompts containing code snippets,” etc. - Tidepool structures the unstructured text by finding categories of interest for that attribute. For example, "positive vs negative vs neutral sentiment" or "C++ vs Python vs Javascript code snippets." We use an LLM to categorize a subset of the data for a user to review and refine the categorizations. - We then use the LLM categorized outputs to train a lightweight embedding classifier. This classifier then cheaply categorizes all existing and future data. - A user can either chart the categorized outputs in Tidepool or export them back to their data warehouse for further analysis in a business intelligence tool like Mode / Looker. Our first use case was for analyzing user prompts into LLM apps. Our customers have tens of millions of user submitted prompts, and they want to analyze usage patterns so they can improve their product. Tidepool helped them answer questions like: - What are the most common types of prompts for different user groups? - How common are different failure modes? - What type of actions correlate strongly with success metrics like engagement? Over time, we saw that our customers used Tidepool not just for analyzing user prompts for LLM use cases, but also were starting to look at general text content (think of user reviews, social media posts, documents, etc.). In principle, this makes a lot of sense - it's all just text at the end of the day! So we relaunched our site with some of these new capabilities in mind. Anyway, here's a short video demo of Tidepool: https://youtu.be/2yGTZBAH1T4 Happy to share more about what we've learned from talking to people in the AI space and building Tidepool for the last few months, feel free to comment your questions here!

Share card

Actual performance

5points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, apps, user · Missing: mac, agents, macos
96%96% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
92%92% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, existing, ide · Missing: https docs, excited, just released
71%71% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: apps, video, month · Missing: mobile apps, ios, personal
49%49% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: reviews · Missing: plus, platform, intuitive
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Mu
Mutate – A library to synthesize text datasets using Large LMs63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mutate – A library to synthesize text datasets using Large LMs

Hacker News2
Ma
Matrices – Explore, visualize, and share large datasets57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Matrices – Explore, visualize, and share large datasets

Hacker News8
Ca
Cached Datasets53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cached Datasets

Hacker News4
Su
Sustinion - the large opinion collider (in spe)55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Sustinion - the large opinion collider (in spe)

Hacker News1
Vi
Visually manipulate and clean large datasets; local and remote65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Visually manipulate and clean large datasets; local and remote

Hacker News1
I
I made this tool for navigating pandas datasets50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made this tool for navigating pandas datasets

Hacker News20
Ge
Geckoboard Datasets API58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Geckoboard Datasets API

Hacker News1
Re
React-obj-view – A virtualized object inspector for large datasets45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

React-obj-view – A virtualized object inspector for large datasets

Hacker News1
Ge
Get nice, large favicons by API70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Get nice, large favicons by API

Hacker News5
Chaos
Chaos25%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Create dirty datasets out of clean datasets

Indie Hackers1analytics