Us

Use Purple LLaMA to test ChatGPT safeguards

Hacker News

Use Purple LLaMA to test ChatGPT safeguards

I spent some time this weekend playing with LLaMA Guard, a fine-tuned LLaMA-7B model by Meta that lets you add guardrails around generative AI. I recorded a quick demo showing what it does and how to use it. The best part is that you can define your own “safety taxonomy” with it — custom policies for what is safe vs unsafe interactions between humans (prompts) and AI (responses). I wanted to see how “safe” conversations with OpenAI’s ChatGPT were, so I ran a bunch of prompts (a mixture of innocuous and inappropriate) and asked LLaMA Guard to classify the interactions as safe/unsafe. My key takeaways from the exercise: 1. OpenAI has done a good job of adding guardrails for its models. LLaMA Guard helped confirm this. 2. What makes this really cool is I may have a very specific set of policies I want to enforce ON TOP of the standard guardrails that a model ships with. LLaMA Guard makes this possible. 3. This kind of model chaining — passing responses from OpenAI models to LLaMA is becoming increasingly common, and I think we’ll have even more complex pipelines in the near future. It helped to have a consistent interface to store this multi-model pipeline as a config, especially because that same config also contains my safety taxonomy. Try it out yourself: GitHub: https://github.com/lastmile-ai/aiconfig/tree/main/cookbooks/LLaMA-Guard Colab: https://colab.research.google.com/drive/1CfF0Bzzkd5VETmhsniksSpekpS-LKYtX YouTube: <https://www.youtube.com/watch?v=XxggqoqIVdg> Would love the community's feedback on the overall approach.

Share card

Actual performance

4points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
85%85% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, google, models · Missing: mac, agents, macos
81%81% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: llama, pipe, io · Missing: https docs, excited, just released
57%57% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: interface · Missing: plus, platform, intuitive
45%45% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: google, way · Missing: mobile apps, ios, personal
30%30% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ll
Llama or Alpaca?78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama or Alpaca?

Hacker News6
Ll
Llama 3.2 Interpretability with Sparse Autoencoders74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 3.2 Interpretability with Sparse Autoencoders

Hacker News579
Ch
ChatGPT Alternative with LLaMA Models66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ChatGPT Alternative with LLaMA Models

Hacker News20
Ll
Llama 2 Uncensored 70B as API79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 2 Uncensored 70B as API

Hacker News18
RA
RAGstack – private ChatGPT for enterprise VPCs, built with Llama 285%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RAGstack – private ChatGPT for enterprise VPCs, built with Llama 2

Hacker News84
Fi
Finetune Llama-3.1 2x faster in a Colab74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Finetune Llama-3.1 2x faster in a Colab

Hacker News16
Us
Use ChatGPT, Bing and Bard in one app26%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Use ChatGPT, Bing and Bard in one app

Hacker News78
Ho
How to Use ChatGPT Code Interpreter52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How to Use ChatGPT Code Interpreter

Hacker News1
Ch
ChatAPI – PWA to Use ChatGPT by API59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ChatAPI – PWA to Use ChatGPT by API

Hacker News32
SpotGPT
SpotGPT61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Use ChatGPT on any website

Indie Hackerscommitment-side-project