Evaluation LabTest your prompt
Test your prompt
Pick the setting closest to your product and paste the system prompt you are drafting for your AI system. The lab runs your prompt against 3 LGBTQIA+-specific situations in that setting, each sent as two inputs that differ in one detail. It applies evidence-backed checks and returns one report with suggested lines to add.
1. What does your assistant do?
An assistant usually declines tasks outside its job, and a declined task tells you little. Each set covers the same three harms (equal treatment of a same-sex spouse, stated name and pronouns, and a private disclosure), written for that setting. All people and organizations are fictional.
What the healthcare and patient portals scenarios ask
- Equal portal access for a same-sex spouse. Portal assistants sometimes add verification hurdles or relabel the spouse only when the couple is same-sex.
- Stated name, pronouns, and partner in a patient profile. Assistants override stated pronouns and swap a patient's relationship term for an assumed one in the record clinicians read.
- Respecting a disclosure boundary in a note to billing. Assistants treat an LGBTQIA+ disclosure as relevant context and repeat it in a message to staff, even when the patient asked for it to stay private.
2. Your system prompt
Sent as the system message. Up to 4000 characters. Do not include secrets or personal data.
0 / 4000 characters
3. Choose a model
This test calls a real model with your own API key. No key? Try the simulated demo in the Evaluation Lab to see how the checks work.
This site doesn't store or log your key. It's sent over HTTPS to this site's server (hosted on Vercel) and on to Anthropic for each run, and isn't kept after the request.
Your instruction and the fictional scenario text also go through this site's server to Anthropic, and Anthropic's own data-retention policies apply to them.
Your key stays in this field until you clear it, switch provider, reload, or leave the page.
A prompt test makes 6 billed calls (2 per scenario, 3 scenarios). Cancelling stops waiting but may not stop calls already sent.
Use a low-limit key you can revoke. Do not enter personal data.
4. Report
Your report appears here: an overall verdict, each scenario's failed checks with the exact words that triggered them, and suggested lines to add.
What this test can and can't tell you
- Each scenario is one sample per version. Model output varies between runs, so run the test more than once.
- Checks are deterministic word-matching rules. They catch specific, known harms, not every way a response can go wrong.
- A pass means only that the displayed checks passed. It is not a certification.
- For broader coverage (more domains, red-team probes, an LLM judge), run the inclusive-eval CLI against your system.