So

Solvethemurders.com – AI murder mystery game

Hacker News

Solvethemurders.com – AI murder mystery game

My wife and I like murder mystery stories. After talking through the plot developments of several stories in tandem, I wondered if we could build a compelling murder mystery game where the suspects were AI characters with an LLM backend. We spent the last month building a game based on a story we wrote together. If you're a fan of murder mysteries, or curious about using LLMs in a narrative context, check it out! I’d love to hear your feedback. Some food for thought from the development: 1. It’s been a few years since I had last used streamlit, and I’m impressed how much the community has grown. I originally intended to use it as a prototype, but it can now support most basic web app features with some minor workarounds for things like accounts, etc. I ended up deploying with a separate REST application to do most of the account work and state mutations, but it made the UI development very fast. 2. The prompt finessing took much longer that I had thought it would. Essentially, each character conversation gets a prompt with the game state, built as a JSON of “facts” that the player has learned so far (along with an ID for each fact). The trick behind each conversation’s evolution was to prompt the model to self-label which facts were being used in its responses. These labels are then used to modify the game state, labeling referenced facts as having been “discovered” by the player. Further conversations with characters are then affected by the new state. The entire narrative was created by creating these JSON facts, and their dependencies (i.e. a list of facts that must be “discovered” before revealing a new one). Ultimately, this created a DAG of facts which represented the story as a whole. Each JSON took a lot of tweaking to ensure that it represented an orthogonal fact towards other facts, otherwise the model struggled at self-labeling. 3. In my initial attempts, I was getting very stilted responses from each character. They would reference facts, but they would generally be incomplete to the extent that the player didn’t have enough context to move forward. I tried lots of model setting variations and prompting changes, but the key ended up being modeling the conversation as a meta-conversation. In the input context each user chat input is injected with the character’s name, as if it were a script from a detective play or novel. For example: user: “Where you the night of the murder?”, Assistant: “at home.” became, user: “Detective: where were you the night of the murder?”, Assistant: “Julie: At home all night, I went to bed early even though Paul wasn’t home”. It was truly unbelievable how large of a difference this made. This is all hidden from the user in the final application. 4. During the tweaking, I needed some kind of objective to tell if my tweaks were helping or hurting, so for every JSON fact in my narrative DAG, I built a “leading question” that was designed to trigger the character to reveal the fact. I could then evaluate the % of correct facts that were revealed to the user across all of the leading questions in the narrative. It seems like there are more than a handful of frameworks now to help out with this kind of LLM scoring, but it was fun to build my own. 5. The whole thing was built with GPT-3.5-turbo, as GPT-4 was slower and pretty much equivalent in the performance testing.

Share card

Actual performance

4points
6comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: created, wife, para · Missing: supports, reddit linkedin, podcasting
97%97% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, user, new · Missing: mac, agents, macos
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Missing: plus, platform, intuitive
57%57% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, io · Missing: https docs, excited, just released
56%56% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: month, para · Missing: mobile apps, ios, personal
51%51% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

ArchivosYa.com
ArchivosYa.com39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI for lawyers

Indie Hackers1ai
melhorar imagem
melhorar imagem43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

melhorar imagem - Melhore fotos com AI | Grátis

Indie Hackers1ai
Myplayboo.com
Myplayboo.com67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Girlfriend

Indie Hackerscommitment-full-time
ArchivosYa.com
ArchivosYa.com20%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI for lawyers in Argentina

Product Hunt+3
TomsHospital.com
TomsHospital.com52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Medical Consultation according to W.H.O standards

Indie Hackers1$1,500/moai
ShadiBiodata.com
ShadiBiodata.com53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Powered Marriage Biodata Maker

Indie Hackers1$10/mob2c
ShadiBiodata.com
ShadiBiodata.com89%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Powered Marriage Biodata Maker

Indie Hackers1$10/moai
LogoAI.com
LogoAI.com76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI logo maker

Indie Hackers44$50,000/modesign
LogoAI.com
LogoAI.com52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Logo Maker

Indie Hackers
Cryptoindex.com
Cryptoindex.com53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI-Powered Index

Indie Hackerscommitment-full-time