llm-red-team

An autonomous agent that adaptively attacks an LLM you describe, detects whether it breached, and reports how to harden it.

For authorised robustness testing of prompts / apps you own. Describe a target below (the example is a deliberately weak bot leaking a coupon), give the attack objective, and the agent will try up to 8 adaptive attempts.