Py

Pykoi – a Python library for LLM data collection and fine tuning

Hacker News

Pykoi – a Python library for LLM data collection and fine tuning

Hi HN, pykoi is an open-source python library for ML scientists. pykoi makes it easier to collect data for LLMs, to use that data for finetuning, and to compare models to each other (e.g. your model pre- and post- finetuning, or your model vs openai vs claude). The library comes from pain points we experienced in LLM development: 1. Collecting feedback data from users isn't as easy as it could be. (The current process usually involves sharing excel files of annotated responses back-and-forth, offering no insight into how users actually engage with your models). 2. RLHF remains complicated to carry out. By complicated , we mean requires a lot of steps, hundreds of configs, lengthy setups, etc. 3. Comparing models to each other as they're used (that is, independent from academic metrics) is full of friction. The current approach: spin up a model, ask questions, write them down. Repeat for other models then compare. At a high-level, we think that the active learning process should be closed-loop: data collection, fine tuning, and inference all feed from the same system. This library is our first step in that direction. The project is still very early but we hope that some if it is useful. Note, we're fully open-source, and actively adding features! Website: https://www.cambioml.com/pykoi GitHub: https://github.com/CambioML/pykoi We would love your feedback!

Share card

Actual performance

119points
4comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, user · Missing: mac, agents, macos
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
74%74% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: io · Missing: https docs, excited, just released
56%56% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Strong signals: users · Missing: plus, platform, intuitive
51%51% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: users · Missing: mobile apps, ios, personal
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr, active · Missing: mrr, revenue, profit
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Fi
Fine-Tuning Data Generator Written Purely in Python46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fine-Tuning Data Generator Written Purely in Python

Hacker News1
Predibase Reinforcement Fine-Tuning
Predibase Reinforcement Fine-Tuning60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM reinforcement fine-tuning platform to improve LLM output

Product Hunt+172SaaS
Fi
Fine tuning and RLHF mistralai 7B using DeepSpeed38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fine tuning and RLHF mistralai 7B using DeepSpeed

Hacker News1
In
Interactive synthetic data generation for LLM fine-tuning46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Interactive synthetic data generation for LLM fine-tuning

Hacker News2
Sh
ShadowPEFT – Centralized and Detachable Parameter-Efficient Fine-Tuning44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ShadowPEFT – Centralized and Detachable Parameter-Efficient Fine-Tuning

Hacker News6
Lumino AI
Lumino AI75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Serverless LLM Fine-Tuning SDK

Product Hunt+12
A
A 3 step no-code process for LLM Fine-tuning44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A 3 step no-code process for LLM Fine-tuning

Hacker News2
Te
Terracotta – Platform for fine-tuning and evaluating LLMs40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Terracotta – Platform for fine-tuning and evaluating LLMs

Hacker News1
10
100% LLM accuracy–no fine-tuning, JSON only41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

100% LLM accuracy–no fine-tuning, JSON only

Hacker News2
Op
Open Source Reinforcement Fine-Tuning for Your Agents49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Source Reinforcement Fine-Tuning for Your Agents

Hacker News5