US

USearch Images demo in 200 lines of Python

Hacker News

USearch Images demo in 200 lines of Python

Hey everyone! I am excited to share updates on four of my & my teams' open-source projects that take large-scale search systems to the next level: USearch, UForm, UCall, and StringZilla. These projects are designed to work seamlessly together, end-to-end—covering everything from indexing and AI to storage and networking. And yeah, they're optimized for x86 AVX2/512 and Arm NEON/SVE hardware. USearch [1]: Think of it as Meta FAISS on steroids. It's now quicker, supports clustering of any granularity, and offers multi-index lookups. Plus, it's got more native bindings than probably all other vector search engines combined: C++, C, Python, Java, JavaScript, Rust, Obj-C, Swift, C#, GoLang, and even slightly outdated bindings for Wolfram. Need to refresh that last one! UForm v2 [2]: Imagine a much smaller OpenAI CLIP but more efficient and trained on balanced multilingual datasets, with equal exposure to languages from English, Chinese, and Hindi to Arabic, Hebrew, and Armenian. UForm now supports 21 languages, is so tiny that you can run it in the browser, and outputs small 256-dimensional embeddings. Perfect for rapid image and video searches. It's already available on Hugging-Face as "unum-cloud/uform-vl-multilingual-v2". UCall [3]: It started as a FastAPI alternative focusing on JSON-RPC (instead of REST protocols), offering 70x the bandwidth and 1/50th the latency. It was good but not enough, so we've added REST and TLS support, broadening its appeal. I've merged that code, and it is yet to be tested. Early benchmarks suggest that we still hit the same 150'000-250'000 requests/s on a single CPU core in Python by reusing HTTPS connections. StringZilla [4]: This project lets you sift through multi-gigabyte or terabyte strings with minimal use of RAM and maximal use of SIMD and SWAR techniques. All these projects are engineered for scalability and efficiency, even on tight budgets. Our demo, for instance, works on hundreds of gigabytes of images using just a few gigabytes of RAM and no GPUs for AI inference. That is a toy example with a small, noisy dataset, and I look forward to showing a much larger setup. Interestingly, even this tiny setup illustrates issues common to UForm and much larger OpenAI CLIP models - the quality of Multi-Modal alignment [5]. It also shows how different/accurate the search results are across different languages. Synthetic benchmarks suggest massive improvements for some low-resource languages (like Armenian and Hebrew) and more popular ones (like Hindi and Arabic) [6]. Still, when we look at visual demos like this, I can see a long road ahead for us and the broader industry, making LLMs Multi-Modal in 2024 :) All of the projects and the demo code are available under an Apache license, so feel free to use them in your commercial projects :) PS: The demo looks much nicer with just Unsplash dataset of 25'000 images, but it's less representative of modern AI datasets, too small, and may not be the best way to honestly show our current weaknesses. The second dataset - Conceptual Captions - is much noisier, and quite ugly. [1]: https://github.com/unum-cloud/usearch [2]: https://github.com/unum-cloud/uform [3]: https://github.com/unum-cloud/ucall [4]: https://github.com/ashvardanian/stringzilla [5]: https://ashvardanian.com/posts/searching-stable-marriages/#e... [6]: https://www.unum.cloud/blog/2023-08-17-uform-graphcore

Share card

Actual performance

9points
2comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, started · Missing: reddit linkedin, podcasting, created
96%96% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, visual · Missing: mac, agents, macos
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: excited, ide, 000 · Missing: https docs, just released, exist
67%67% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: video, way · Missing: mobile apps, ios, personal
52%52% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: plus, efficient · Missing: platform, intuitive, reviews
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Py
PyProjectPlanner, budgeting and timelines in 200 lines of Python45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PyProjectPlanner, budgeting and timelines in 200 lines of Python

Hacker News2
Si
Simulating Infectious Disease in 200 lines of C60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Simulating Infectious Disease in 200 lines of C

Hacker News5
Ha
Hangman in 3 lines of Python57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hangman in 3 lines of Python

Hacker News58
A
A very simple Lisp interpreter in 250 lines of literate Python70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A very simple Lisp interpreter in 250 lines of literate Python

Hacker News5
Re
Retirement Calculator in 8 lines of Python52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Retirement Calculator in 8 lines of Python

Hacker News1
Qu
Querying 50k Gameservers in 100 lines of Python46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Querying 50k Gameservers in 100 lines of Python

Hacker News1
Bu
Build python ASTs in 502 lines of python without using modules65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Build python ASTs in 502 lines of python without using modules

Hacker News2
Fu
FuzzyFinder – in 10 lines of Python53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

FuzzyFinder – in 10 lines of Python

Hacker News10
El
ElasticSearch for Images – demo with shoe images62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ElasticSearch for Images – demo with shoe images

Hacker News5
ti
tiling windows for MacOSX in less then 200 lines of code55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

tiling windows for MacOSX in less then 200 lines of code

Hacker News8