LLMem, a read through cache for OpenAI chat completions
LLMem, a read through cache for OpenAI chat completions
When building a system around OAI, I found myself sending the same request multiple times as part of developing/testing some other part of the system. On top of wasting money in this way, I was also throwing away potentially useful later training data to specialize a smaller LLM for my use case. I’m hosting an open server atm since I hit it from various different networks for my projects, or you easily enough run it as a local service.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
I made an OpenAI-powered tamagotchi that reacts to what you read online
Cataloging the NYRB and LRB with OpenAI Embeddings
Analyzing a snapshot isolation read anomaly
Shit I Read
Read less, do more with SummarizeBot
Read less and do more with SummarizeBot
Read-Only PasteBin
I ported danmaz74 "HN: Mark All Read" to Firefox
Show HN : Figure out how much more you have to read
URLTodo – The treatment for "Read It Never" syndrome