Platform to serve ML/DL models at low latency
Platform to serve ML/DL models at low latency
I've always had a hard time deploying ML models into production. The traditional approach is to use Flask, Gunicorn with Nginx. This requires a lot of setup time. Also, inferring model with flask is slow and requires custom code for caching and batching. Scaling in multiple machines is also hard. We have created panini.ai as a solution. https://www.panini.ai/ is a platform to serve ML/DL models at low latency and makes it super easy to deploy AI models in the cloud. Once, deployed in the cloud it will provide you with an API key to infer the model. Our backend is written in C++, which provides very low latency during model inference and the model is stored in Kubernetes so, it is scalable to multiple nodes. We take care of caching and batching inputs during model inference. I have also created a YouTube tutorial on how to use panini: https://www.youtube.com/watch?v=tCz-fi_NheE&t= Please let me know what you guys think. If you have any questions, please email me at avin@panini.ai
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Panini AI – A platform to serve ML/DL models at low latency
Low-latency jamming over the internet
NoraSector – low-latency WebRTC police/fire/EMS scanner (Seattle)
Efemarai – Visualizing and debugging ML models
stationary_vector: A Parallelizable, Low-Latency C++ Vector
VulcanSQL – Serve high-concurrency, low-latency API from OLAP
Global Low Latency APIs
Low Latency Replication from Postgres to ClickHouse Using PeerDB
An ultra low-latency WebRTC radio
Loriot.io – low-latency network as a service