Te

Testing AI for Legal Document Classification

Hacker News

Testing AI for Legal Document Classification

Hey HN! I recently had the chance to experiment with ChatGPT to review legal documents for a client. They wanted to speed up the process of categorizing legal documents and wondered if AI could help. I've been playing with ChatGPT for a few years and following all the nuances and challenges associated with LLMs, so I advised them to tread carefully. They agreed to a systematic test case to find the best approach and to keep law students and professors in the loop for the final QC. I carefully developed and revised a pre-prompt to accompany each document via ChatGPT's Assistant tool. This prompt included relevant definitions for elections, document types, and descriptions for each tag. After the results were returned, I analyzed and summarized the findings and provided the data to the client for careful review. The results were pretty promising! They reviewed the same documents and came to the same conclusion as the AI the vast majority of the time! Additionally, the tags provided were accurate enough to help speed up the legal evaluation. tl;dr - here are some other high-level results: 1. Both budget (4o-mini) and full-size (4o) models do a reasonable job determining whether a document is related to elections. 2. The newer, larger models perform adequately for the more complicated task of assigning tags to documents. 3. The mini model struggled to return accurate or relevant document topic tags. 4. By evaluating each document multiple times, we could check for "concurrence" of the AI's evaluation. 5. Cost is a factor, given the length of some of these documents. I estimated the backlog of documents at ~150M tokens. If we evaluate each document multiple times, the costs really add up!

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
74%74% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Missing: plus, platform, intuitive
51%51% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, ide, io · Missing: https docs, excited, just released
42%42% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
26%26% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
12%12% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ag
Agentsnap – Snapshot testing for AI agents18%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Agentsnap – Snapshot testing for AI agents

Hacker News5
Un
Understudy: Scenario Testing for AI Agents20%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Understudy: Scenario Testing for AI Agents

Hacker News4
Mi
Mida.so – automated A/B testing with AI28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mida.so – automated A/B testing with AI

Hacker News2
Contractize
Contractize20%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Legal document automation

Indie Hackers
AutoContract
AutoContract53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

“Just launched AutoContract – AI legal document generator"

Indie Hackers1ai
ZeroLeaks
ZeroLeaks81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Security testing for AI agents

Product Hunt+8
PageTest.AI
PageTest.AI56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Smarter website content testing with AI

Indie Hackers1ai
PageTest.AI
PageTest.AI64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Smarter website content testing with AI

Product Hunt+147User Experience
SignedSorted
SignedSorted31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI-Drafted Legal Agreements

Indie Hackerscommitment-side-project
Janus
Janus85%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Simulation testing for AI agents

Product Hunt+268Analytics