Continuous-eval – Granular evaluation of GenAI pipelines
Continuous-eval – Granular evaluation of GenAI pipelines
Hi HN - we are the creators of “continuous-eval”, an open-source tool to test and evaluate generative AI apps. "Continuous-eval" came from our efforts to measure, validate and improve the reliability of a finance AI copilot we were developing for banks. End-to-end evaluation was not enough for us. We wanted to have granular evaluations that help pinpoint the bottlenecks and identify what / how to improve. We’ve since developed more metrics and made the framework more flexible so it can evaluate components like agent tool use, code change, retrieval steps, etc. Let us know what you think of our approach to GenAI App evaluation.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Buildbox Continuous Delivery Pipelines
Automated GenAI evaluation that works
Continuous Fuzzing for Go
Bencher – Continuous Benchmarking
Rues an Expression Evaluation Sidecar
Pipelines
Pipelines
GenAi Yule Lights
Shadergarden: Create reloadable graphical pipelines with Lisp and GLSL
Create continuous development pipelines for static site generators