I

I compressed 10k PDFs into a 1.4GB video for LLM memory

Hacker News

I compressed 10k PDFs into a 1.4GB video for LLM memory

While building a Retrieval-Augmented Generation (RAG) system, I was frustrated by my vector database consuming 8GB RAM just to search my own PDFs. After incurring $150 in cloud costs, I had an unconventional idea: what if I encoded my documents into video frames? The concept sounded absurd—storing text in video? But modern video codecs have been optimized for compression over decades. So, I converted text into QR codes, then encoded those as video frames, letting H.264/H.265 handle the compression. The results were surprising. 10,000 PDFs compressed down to a 1.4GB video file. Search latency was around 900ms compared to Pinecone’s 820ms—about 10% slower. However, RAM usage dropped from over 8GB to just 200MB, and it operates entirely offline without API keys or monthly fees. Technically, each document chunk is encoded into QR codes, which become video frames. Video compression handles redundancy between similar documents effectively. Search works by decoding relevant frame ranges based on a lightweight index. You get a vector database that’s just a video file you can copy anywhere. GitHub: https://github.com/Olow304/memvid

Share card

Actual performance

61points
23comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Hacker NewsStrong engagement from HN community · Strong signals: ide, 000, io · Missing: https docs, excited, just released
71%71% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Product HuntOn track for Day 1 leaderboard · Strong signals: coding, code · Missing: mac, agents, macos
59%59% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: video, month, monthly · Missing: mobile apps, ios, personal
53%53% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
50%50% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
18%18% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
4%4% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Bo
Bookmarklet for Previewing References in PDFs54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Bookmarklet for Previewing References in PDFs

Hacker News9
PD
PDFy – an Imgur for PDFs71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PDFy – an Imgur for PDFs

Hacker News2
Me
MemX – Shared memory for LLM agents49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

MemX – Shared memory for LLM agents

Hacker News1
Re
Redis-LLM – Redis module integrates LLM with Redis45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Redis-LLM – Redis module integrates LLM with Redis

Hacker News2
Li
LitLLM the Spiciest LLM Wrapper34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LitLLM the Spiciest LLM Wrapper

Hacker News1
LL
LLM Reasonsers46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Reasonsers

Hacker News2
Re
Resilient LLM23%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Resilient LLM

Hacker News1
He
Hegelion – Force your LLM to argue with itself before answering58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hegelion – Force your LLM to argue with itself before answering

Hacker News1
I
I Stopped Hoping My LLM Would Cooperate49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I Stopped Hoping My LLM Would Cooperate

Hacker News3
LLM Hotkey
LLM Hotkey39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.
TrustMRROther