We

We wrote a book on LLM system evals with a bear and fox

Hacker News

We wrote a book on LLM system evals with a bear and fox

Hey all! @sridatta and I wrote a book/zine called Forest Friends on system evals for LLM-driven apps. But it's a bit more whimsical, a bit more visual, and very much inspired by the meme of LLMs being a shoggoth polished into a smiley face with RLHF. LLM system evals are important as companies move past the flashy AI demos to reliable production apps. System evals keep coming up as the answer for what you "should do", but it's not exactly a standard part of the software engineering toolkit. So we pulled from @sridatta's seven years as a research engineer at Google, plus a ton of best practices from around the web, and we wrote a zine that could take you from zero to eval in a way that’s fun to read. To be clear, model evals and system evals are two different things. The former compares different models, and the latter is a metric on how well you're servicing your customer queries. When you create a system eval, you're essentially defining what "good" looks like for your system. Lots of people use "vibes-based evals" (LGTM@K). It's a good place to start and will get you further than you think. But at some point, you need system evals as you get more users and more diversity in queries. To quote: Garry Tan says "Don’t rawdog your prompts! Write evals!" https://x.com/garrytan/status/1842210665550983409 Swyx says "Production AI Engineering starts with Evals" https://x.com/latentspacepod/status/1844870676202783126 How did we end up doing a zine? Coming off collaborating on the Technium Podcast, we wanted to tackle a topic where we had deep expertise while also exercising our entrepreneurial muscles. We came up with writing a zine, inspired by Julia Evans' Wizard Zine and Sailor Mercury's Bubble Zine. Originally, we were shooting for 30 pages, but ended up with 60 pages. This was also an experiment in image generation for a product. Initially, I created illustrations by hand. Midway, I decided to switch to using Midjourney to make our deadline. I needed to generate scenes with consistent characters in a specific style set in a specific architecture, and it turned out to be hard. Initially, I would be generating images 8 to 10 hours a day. Eventually, I got better at predicting what would work, and generating a suitable image dropped to 1.5 hours. Rest assured, however, all the text is human-generated and hand-edited. The issue has been well received so far. Here are some quotes from early readers: "Thanks for this resource! It provides a comprehensive introduction to LLM evaluation systems that rings true to my daily work as an AI engineer—all in less than an hour of reading and with minimal jargon. I’ll be recommending this to my team." "I was a fan of The Poignant Guide to Ruby many years ago, so it’s great to see a playfulness brought to the world of LLMs. I’m building an evals platform that makes it as easy as possible for any developer to get started with evals. This edition has been great to make sure we get the basics and terminology right." "Here's an engaging intro to evals by @sridatta and @iamwil. They've clearly put a lot of care and effort into it, where the content is well organized with plenty of illustrations throughout. Across 60 pages, they explain model vs. system evals, vibe checks and property-based tests, designing eval criteria, aligning LLM evaluators, how to measure alignment via various metrics, how to analyze evals to improve our system, and more. Now I can just direct folks to [the zine] instead of having to write it myself haha" The zine is available now. There's a preview if you want to check the vibe. https://forestfriends.tech/assets/preview.pdf I'd love to hear any feedback on the first issue, or what other topics you'd like to see tackled in later issues. If you have questions about the process of making the zine, I'd be happy to answer those also. Here's the link: https://forestfriends.tech Here's where to buy: https://issue1.forestfriends.tech/

Share card

Actual performance

11points
2comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: created, started · Missing: supports, reddit linkedin, podcasting
95%95% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, google, apps · Missing: mac, agents, macos
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
63%63% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: apps, google, users · Missing: mobile apps, ios, personal
51%51% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: plus, platform, users · Missing: intuitive, reviews, host
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

We
We wrote a book on system evals68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We wrote a book on system evals

Hacker News3
Zi
Zine on LLM Evals44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Zine on LLM Evals

Hacker News1
Forest Friends Zine
Forest Friends Zine38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A zine about building LLM evals with animals and shoggoths

Indie Hackerscommitment-side-project
Stax
Stax61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Move your LLM evals from vibes to data

Product Hunt+178A/B Testing
Ag
Agent/LLM observability for tracing, cost, evals, and debugging69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Agent/LLM observability for tracing, cost, evals, and debugging

Hacker News1
Ev
Evals Skills52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Evals Skills

Hacker News4
Py
Pytest-evals – Simple LLM apps evaluation using pytest58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pytest-evals – Simple LLM apps evaluation using pytest

Hacker News13
Cl
Claude Code skills for building LLM evals50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Claude Code skills for building LLM evals

Hacker News2
My
My first ever e-book, Hacking Obamacare65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

My first ever e-book, Hacking Obamacare

Hacker News6
I
I Wrote a Book on ReactiveCocoa75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I Wrote a Book on ReactiveCocoa

Hacker News1