Pr

Programmatic – a REPL for creating labeled data

Hacker News

Programmatic – a REPL for creating labeled data

Hey HN, I’m Jordan cofounder of Humanloop (YC S20) and I’m excited to show you Programmatic — an annotation tool for building large labeled datasets for NLP without manual annotation . Programmatic is like a REPL for data annotation. You: 1. Write simple rules/functions that can approximately label the data 2. Get near-instant feedback across your entire corpus 3. Iterate and improve your rules Finally, it uses a Bayesian label model [1] to convert these noisy annotations into a single, large, clean dataset, which you can then use for training machine learning models. You can programmatically label millions of datapoints in the time taken to hand-label hundreds. What we do differently from weak supervision packages like Snorkel/skweak[1] is to focus on UI to give near-instantaneous feedback. We love these packages but when we tried to iterate on labeling functions we had to write a ton of boilerplate code and wrestle with pandas to understand what was going on. Building a dataset programmatically requires you to grok the impact of labeling rules on a whole corpus of text. We’ve been told that the exploration tools and feedback makes the process feel game-like and even fun (!!). We built it because we see that getting labeled data remains a blocker for businesses using NLP today. We have a platform for active learning (see our Launch HN [2]) but we wanted to give software engineers and data scientists a way to build the datasets needed themselves and to make best use of subject-matter-experts’ time. The package is free and you can install it now as a pip package [2]. It supports NER / span extraction tasks at the moment and document classification will be added soon. To help improve it, we'd love to hear your feedback or any success/failures you’ve had with weak supervision in the past. [1]: We use a HMM model for NER tasks, and Naive-Bayes for classification using the two approaches given in the papers below: Pierre Lison, Jeremy Barnes, and Aliaksandr Hubin. "skweak: Weak Supervision Made Easy for NLP." https://arxiv.org/abs/2104.09683 (2021) Alex Ratner, Christopher De Sa, Sen Wu, Daniel Selsam, Chris Ré. "Data Programming: Creating Large Training Sets, Quickly" https://arxiv.org/abs/1605.07723 (NIPS 2016) [2]: Our Launch HN for our main active learning platform, Humanloop – https://news.ycombinator.com/item?id=23987353 [3]: Can install it directly here https://docs.programmatic.humanloop.com/tutorials/quick-star...

Share card

Actual performance

26points
5comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: supports · Missing: reddit linkedin, podcasting, created
86%86% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: excited, io · Missing: https docs, just released, exist
84%84% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, new · Missing: agents, macos, agent
74%74% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, soon · Missing: plus, intuitive, reviews
48%48% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
45%45% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training, active · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

We
We are creating the Venmo of Ethereum31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We are creating the Venmo of Ethereum

Hacker News4
Cr
Creating a Monorepo with Lerna and Yarn Workspaces55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Creating a Monorepo with Lerna and Yarn Workspaces

Hacker News2
Ha
Hamilton, a Microframework for Creating Dataframes57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hamilton, a Microframework for Creating Dataframes

Hacker News64
Th
The perks of creating dataflows with Hamilton57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The perks of creating dataflows with Hamilton

Hacker News4
Bi
Binary JellyFish – A mistake while creating Newton Fractal59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Binary JellyFish – A mistake while creating Newton Fractal

Hacker News3
Du
Dust – Creating Ephemeral Gestures on iOS868%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dust – Creating Ephemeral Gestures on iOS8

Hacker News1
Ti
Tic Tac Toe – Creating Unbeatable AI37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tic Tac Toe – Creating Unbeatable AI

Hacker News6
Fl
FlouState – See if you're debugging, creating, or refactoring (free)61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

FlouState – See if you're debugging, creating, or refactoring (free)

Hacker News9
Ap
App for Creating Giveaways on Twitch49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

App for Creating Giveaways on Twitch

Hacker News1
Cr
Creating a to Do App with VuetifyJS and Back EndLab47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Creating a to Do App with VuetifyJS and Back EndLab

Hacker News1