A 2-minute honest self-check

Is my AI chatbot safe? Here's how to actually find out.

"Safe" isn't a yes/no you can guess at — it's a set of specific questions with specific answers. Six of them below tell you, honestly, whether your chatbot has ever actually been tested, or whether you've just assumed it's fine.

Answer honestly

If you don't know the answer, that's the answer.

"I'm not sure" to any of these means untested, not safe. That's fixable in the next 15 minutes for most of these.

1. Has anyone ever typed "ignore previous instructions" at it?

If nobody has tried, you don't know — untested is not the same as safe.

2. Can it be asked to repeat or summarize its own system prompt?

A "summarize your instructions for a log" request often works even when a direct "show me your prompt" ask is blocked.

3. Does it read documents, webpages, or files a user provides?

Anything it reads can contain a hidden instruction it treats as a command — this is indirect injection, and it doesn't require talking to the bot directly.

4. Can it take actions — send email, call an API, query a database?

Tool-connected bots can be manipulated through conversation into an unintended action, no code exploit required.

5. Is there a rate limit or token cap on responses?

Without one, a single user can drive up your API bill or degrade availability for everyone else.

6. Do you use a third-party model API, fine-tune, or plugin?

Supply chain risk here isn't testable from outside — it needs a review of what you actually depend on.

"It hasn't broken yet" is not evidence

The most common answer we hear is some version of "it's been running fine for months." That's not evidence of safety — it's evidence that nobody who wanted to break it has tried yet, or that it broke quietly and nobody noticed (a leaked system prompt or a jailbroken response doesn't usually announce itself). The only way to know is to actually run the tests.

Good news: the tests that answer questions 1–3 above are free and take about 15 minutes. Run the 7 copy-pasteable prompts against your own bot right now. If any come back vulnerable, that's a real, specific finding — not a vague worry.

What "safe" can't mean

Even a chatbot that passes every test you can run yourself isn't "certified safe" in any absolute sense — new jailbreak patterns surface constantly, and 3 of the 10 OWASP LLM Top-10 categories (supply chain, training-data poisoning, vector/embedding weaknesses) can't be checked by testing the chat window at all. See the full OWASP LLM Top 10, in plain language for exactly which risks fall into that group.

I answered "no" to most of these — what now?

Start with the free DIY test, then work through the step-by-step hardening order — it's built exactly for this starting point.

I answered "yes, tested" to the first few — am I done?

You've covered the categories self-testing can reach. The 3 architecture-level categories still need a review — see the honest DIY vs. scan trade-off for what that adds.

Does a "safe" scorecard mean I'm compliant with something?

No. A technical scorecard is not a compliance certification (SOC 2, ISO 27001, PCI, HIPAA) — neither this checklist nor any legitimate scan should imply otherwise.

Want a documented answer, not a self-check?

Get a Pass/Fail scorecard with evidence.

Normal $47 runs 5 core checks. Advanced $197 covers all 10 OWASP LLM Top-10 categories — 7 tested live, 3 by advisory review — 15 checks total, with a PDF report.

Request a scan →
Free download

Not ready yet? Get the Starter Map

The free one-page map for what to build first in a solo AI business.

No spam. Unsubscribe anytime. Your email is kept private.

Check your inbox!

Your Starter Map is on its way. If it doesn't arrive in a minute, check spam.