Me

Memory System Hitting 80.1% Accuracy on LoCoMo (Built in 4.5 Months)

Hacker News

Memory System Hitting 80.1% Accuracy on LoCoMo (Built in 4.5 Months)

I’ve been working on an independent memory-retrieval architecture for agent systems. I don’t have a CS background — previously worked climbing cell towers and doing handyman jobs — but I spent the last 4.5 months building a hybrid memory system from scratch. The system combines FAISS, BM25, and a symbolic ranking layer (MCA). Answers are generated with GPT-4o-mini at temperature 0. The focus is determinism, transparency, and reproducibility rather than model size. On the official LoCoMo benchmark (1,540 questions), the system reaches 80.1% average accuracy. To my knowledge, that’s above the publicly reported results for existing agent-memory stacks using small models. Latency is ~2.5 seconds, and cost is ~$0.10 per 1M tokens. Memory is fully isolated and local, which makes it usable for offline or enterprise applications. Repository (code + full reproducible benchmarking): https://github.com/vac-architector/VAC-Memory-System Happy to answer technical questions, discuss the architecture, or hear critiques.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agent, model, models · Missing: mac, agents, macos
81%81% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
76%76% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, existing, io · Missing: https docs, excited, just released
66%66% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: month, answers · Missing: mobile apps, ios, personal
47%47% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
37%37% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

2
2 guys -- Learned to code in 3 months and built this81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

2 guys -- Learned to code in 3 months and built this

Hacker News27
An
An enclosure for my homebrew Calculon/80 microcomputer46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An enclosure for my homebrew Calculon/80 microcomputer

Hacker News4
No
November 2020: Four Months of Notado48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

November 2020: Four Months of Notado

Hacker News10
Mi
Millionshort, 3 months later66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Millionshort, 3 months later

Hacker News22
Unsloth
Unsloth61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Finetune LLMs 2x faster, 80% less memory

Product Hunt+241Open Source
We
We've built FinWise over the last 7 months53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We've built FinWise over the last 7 months

Hacker News2
Wh
What I built 6 months into learning to code74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

What I built 6 months into learning to code

Hacker News29
A
A New Game I Built in Two Months55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A New Game I Built in Two Months

Hacker News3
A
A new Bluebook implementation of the Smalltalk-80 VM53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A new Bluebook implementation of the Smalltalk-80 VM

Hacker News11
AI
AI Search Benchmark: Perplexity Pro Reigns Supreme with 80% Accuracy27%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Search Benchmark: Perplexity Pro Reigns Supreme with 80% Accuracy

Hacker News2