Epsilla – Open-source vector database with low query latency
Epsilla – Open-source vector database with low query latency
Hey HN! We are building Epsilla ( https://github.com/epsilla-cloud/vectordb ), an open-source, self-hostable vector database for semantic similarity search that specializes in low query latency. When do we need a vector database? For example, GPT-3.5 has a 16k context window limit. If we want to let it answer a question about a 300 page book, we cannot put the whole book content into the context. We have to choose the sections of the book that are most relevant to the question. Vector database is specialized at ranking and picking the most relevant content from a large pool of documents based on their semantic similarity. Most vector databases utilize hierarchical navigational small world (HNSW) for indexing the vectors for high precision vector search, and its latency significantly degrades when the precision target is higher than 95%. At a previous company, we worked on building the parallel graph traversal engine. We realized that the bottleneck of HNSW performance is because there are too many sequential traversal steps that don't fully leverage multi-core CPU computation resources. After some research, we found that there are algorithms such as SpeedANN that are targeting this problem, which is not leveraged by industry yet. So we built the Epsilla vector database to turn the research into a production system. With Epsilla, we shoot for 10x lower vector search latency compared to HNSW based vector databases. We did an initial benchmark against the top open source vector databases: https://medium.com/@richard_50832/benchmarking-epsilla-with-... We provide a Docker image for you to install Epsilla backend locally, and provide a Python client and a JavaScript client to connect and interact with it. Quickstart: docker pull epsilla/vectordb docker run --pull=always -d -p 8888:8888 epsilla/vectordb pip install pyepsilla git clone https://github.com/epsilla-cloud/epsilla-python-client.git cd examples python hello_epsilla.py We just started a month ago. We'd love to hear what you think, and more importantly, what you wish to see in the future. We are thinking about a serverless vector database on cloud with a consumption based pricing model, and we are eager to get your feedback.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
stationary_vector: A Parallelizable, Low-Latency C++ Vector
VectorAdmin – An open-source vector database management system
Open Source Edge Proxy for Low Latency Distributed Authorization
OctaneDB – Fast, Open-Source Vector Database for Python
Crux, an open-source bitemporal Datalog database
OSV, Database for open source vulnerabilities
Undb open source nocode database
An open-source OLAP Multidimensional Database
Open Source Database Schemas
HelixDB, Open-Source Hybrid Graph-Vector Database