485

Someone got Gab's AI chatbot to show its instructions (mbin.grits.dev)

submitted 2 months ago by mozz@mbin.grits.dev to c/technology@beehaw.org

200 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[-] TehPers@beehaw.org 8 points 2 months ago

You don't need a LLM to see if the output was the exact, non-cyphered system prompt (you can do a simple text similarity check). For cyphers, you may be able to use the prompt/history embeddings to see how similar it is to a set of known kinds of attacks, but it probably won't be even close to perfect.