Ge

GenderBench – Evaluation suite for gender biases in LLMs

Hacker News

GenderBench – Evaluation suite for gender biases in LLMs

GenderBench is an open-source evaluation benchmark that measures gender biases in large language models. This is my attempt to decompose this pretty complex and difficult topic into interpretable measures. My goal was to systematize the evaluation of unfair behavior in LLMs and help other developer and researchers do their own tests. What is linked here is the report that is generated from GenderBench logs that quantifies how LLMs behave in various situations when gender can be considered. Links: Repository - https://github.com/matus-pikuliak/genderbench Report - https://genderbench.readthedocs.io/latest/_static/reports/ge...

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, open · Missing: mac, agents, macos
73%73% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
63%63% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
46%46% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
38%38% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Indie HackersIH features products with proven revenue · Missing: supports, reddit linkedin, podcasting
33%33% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
3%3% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ru
Rues an Expression Evaluation Sidecar56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rues an Expression Evaluation Sidecar

Hacker News1
De
DeepEval – Evaluation and Unit Testing for LLMs43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DeepEval – Evaluation and Unit Testing for LLMs

Hacker News18
Op
Open Evaluation53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Evaluation

Hacker News3
Aj
Ajoft HRMS Suite31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ajoft HRMS Suite

Hacker News1
GP
GPTCache – Redis for LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPTCache – Redis for LLMs

Hacker News7
pr
prompttest – pytest for LLMs34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

prompttest – pytest for LLMs

Hacker News2
Vi
Visualizing arithmetic and logical expression evaluation in Swift61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Visualizing arithmetic and logical expression evaluation in Swift

Hacker News3
An
An Empirical Evaluation of Linear Probing Algorithms37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An Empirical Evaluation of Linear Probing Algorithms

Hacker News14
Co
Convert VHDL to Verilog using GHDL (+ first evaluation)51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Convert VHDL to Verilog using GHDL (+ first evaluation)

Hacker News2
Flapico
Flapico57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Prompt versioning, testing, and evaluation

Product Hunt+149Developer Tools