GenderBench – Evaluation suite for gender biases in LLMs
GenderBench – Evaluation suite for gender biases in LLMs
GenderBench is an open-source evaluation benchmark that measures gender biases in large language models. This is my attempt to decompose this pretty complex and difficult topic into interpretable measures. My goal was to systematize the evaluation of unfair behavior in LLMs and help other developer and researchers do their own tests. What is linked here is the report that is generated from GenderBench logs that quantifies how LLMs behave in various situations when gender can be considered. Links: Repository - https://github.com/matus-pikuliak/genderbench Report - https://genderbench.readthedocs.io/latest/_static/reports/ge...
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Rues an Expression Evaluation Sidecar
DeepEval – Evaluation and Unit Testing for LLMs
Open Evaluation
Ajoft HRMS Suite
GPTCache – Redis for LLMs
prompttest – pytest for LLMs
Visualizing arithmetic and logical expression evaluation in Swift
An Empirical Evaluation of Linear Probing Algorithms
Convert VHDL to Verilog using GHDL (+ first evaluation)
Prompt versioning, testing, and evaluation