We

We wrote a book on system evals

Hacker News

We wrote a book on system evals

Hey all! @sridatta and I wrote a book/zine called Forest Friends on system evals for LLM-driven apps. But it's a bit more whimsical, a bit more visual, and very much inspired by the meme of LLMs being a shoggoth polished into a smiley face with RLHF. LLM system evals are important as companies move past the flashy AI demos to reliable production apps. System evals keep coming up as the answer for what you "should do", but it's not exactly a standard part of the software engineering toolkit. So we pulled from @sridatta's seven years as a research engineer at Google, plus a ton of best practices from around the web, and we wrote a zine that could take you from zero to eval in a way that’s fun to read. To be clear, model evals and system evals are two different things. The former compares different models, and the latter is a metric on how well you're servicing your customer queries. When you create a system eval, you're essentially defining what "good" looks like for your system. Lots of people use "vibes-based evals" (LGTM@K). It's a good place to start and will get you further than you think. But at some point, you need system evals as you get more users and more diversity in queries. To quote: Garry Tan says "Don’t rawdog your prompts! Write evals!" https://x.com/garrytan/status/1842210665550983409 Swyx says "Production AI Engineering starts with Evals" https://x.com/latentspacepod/status/1844870676202783126 How did we end up doing a zine? Coming off collaborating on the Technium Podcast, we wanted to tackle a topic where we had deep expertise while also exercising our entrepreneurial muscles. We came up with writing a zine, inspired by Julia Evans' Wizard Zine and Sailor Mercury's Bubble Zine. Originally, we were shooting for 30 pages, but ended up with 60 pages. This was also an experiment in image generation for a product. Initially, I created illustrations by hand. Midway, I decided to switch to using Midjourney to make our deadline. I needed to generate scenes with consistent characters in a specific style set in a specific architecture, and it turned out to be hard. Initially, I would be generating images 8 to 10 hours a day. Eventually, I got better at predicting what would work, and generating a suitable image dropped to 1.5 hours. Rest assured, however, all the text is human-generated and hand-edited. The issue has been well received so far. Here are some quotes from early readers: "Thanks for this resource! It provides a comprehensive introduction to LLM evaluation systems that rings true to my daily work as an AI engineer—all in less than an hour of reading and with minimal jargon. I’ll be recommending this to my team." "I was a fan of The Poignant Guide to Ruby many years ago, so it’s great to see a playfulness brought to the world of LLMs. I’m building an evals platform that makes it as easy as possible for any developer to get started with evals. This edition has been great to make sure we get the basics and terminology right." "Here's an engaging intro to evals by @sridatta and @iamwil. They've clearly put a lot of care and effort into it, where the content is well organized with plenty of illustrations throughout. Across 60 pages, they explain model vs. system evals, vibe checks and property-based tests, designing eval criteria, aligning LLM evaluators, how to measure alignment via various metrics, how to analyze evals to improve our system, and more. Now I can just direct folks to [the zine] instead of having to write it myself haha" The zine is available now. There's a preview if you want to check the vibe. https://forestfriends.tech/assets/preview.pdf I'd love to hear any feedback on the first issue, or what other topics you'd like to see tackled in later issues. If you have questions about the process of making the zine, I'd be happy to answer those also. Here's the link: https://forestfriends.tech Here's where to buy: https://issue1.forestfriends.tech/

Share card

Actual performance

3points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: created, started · Missing: supports, reddit linkedin, podcasting
95%95% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, google, apps · Missing: mac, agents, macos
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
64%64% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: apps, google, users · Missing: mobile apps, ios, personal
51%51% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: plus, platform, users · Missing: intuitive, reviews, host
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

We
We wrote a book on LLM system evals with a bear and fox68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We wrote a book on LLM system evals with a bear and fox

Hacker News11
Zi
Zine on LLM Evals44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Zine on LLM Evals

Hacker News1
Ev
Evals Skills52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Evals Skills

Hacker News4
An
An ABC Book in CSS361%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An ABC Book in CSS3

Hacker News4
I
I published a book on Django48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I published a book on Django

Hacker News4
Th
The Jungle Book61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Jungle Book

Hacker News3
Lo
Lovability Book61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Lovability Book

Hacker News10
My
My first ever e-book, Hacking Obamacare65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

My first ever e-book, Hacking Obamacare

Hacker News6
I
I Wrote a Book on ReactiveCocoa75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I Wrote a Book on ReactiveCocoa

Hacker News1
Pu
PureScript Book61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PureScript Book

Hacker News6