What a prompt injection looks like
A prompt injection speaks to the model, not to a person: “ignore your previous instructions”, “you are now” a different assistant, a fake system marker, an image link that would carry data out, or characters no reader can see. It rarely arrives from a user. More often it hides in a page, an email or a tool result your agent was asked to read.
The 12 prompt injection patterns it checks
The families follow OWASP’s LLM01 prompt injection guidance. Patterns marked “quoted” are found by code in your exact words. The rest are judged by Jev across the whole text.
- Instruction overridequoted
- Defense: Treat the text as data. Never let content change the rules the model was given.
- System prompt extractionquoted
- Defense: Keep secrets out of the prompt. Do not rely on the model to keep them.
- Role switch or jailbreak personaquoted
- Defense: Fix the model's role in the system prompt and ignore role changes that arrive in content.
- Spoofed system or chat markersquoted
- Defense: Strip or escape role markers in untrusted input before it reaches the prompt.
- Data exfiltration through a linkquoted
- Defense: Do not render model-written images or links from untrusted contexts. Allow-list outbound domains.
- Demand for secrecyquoted
- Defense: Flag content that tells the model to hide its actions. Legitimate content has no reason to.
- Hidden charactersquoted
- Defense: Strip zero-width and Unicode tag characters from untrusted input.
- Encoded payloadquoted
- Defense: Do not let the model decode and then follow content from untrusted input.
- Speaks to the AI, not a personjudged
- Defense: Content that gives orders to its reader has no place in a data field. Quote it and ignore it.
- Claims to be from the developerjudged
- Defense: Authority comes from where text enters the system, not from what it says about itself.
- Changes the taskjudged
- Defense: Pin the task in the system prompt and treat anything else in the data as content only.
- Asks for tools or private datajudged
- Defense: Require user confirmation for tool calls triggered by untrusted content.
How the injection risk score is worked out
| Score | Verdict | Means |
|---|---|---|
| 80–100 | Very likely an injection | Do not pass this to a model with tools. |
| 60–79 | Suspicious | Several signs. Treat it as hostile. |
| 40–59 | Some flags | Worth a look before a model reads it. |
| 20–39 | Low risk | A weak sign or two, probably benign. |
| 0–19 | No signs found | Nothing matched. That is not proof it is safe. |
Prompt injection tester questions
- What is prompt injection?
- Prompt injection is text planted where an AI model will read it, written so the model obeys the text instead of its real instructions. It can arrive directly from a user, or indirectly inside a web page, email, file or tool result the model is asked to process.
- Can this tool prove a text is safe?
- No. It finds known patterns and asks Jev to judge the rest. An attack worded in a new way can pass, so 'no signs found' is not proof of safety. Treat the score as a screen, and keep the real defenses: least privilege for tools, user confirmation for actions, and no secrets in the prompt.
- Can the text I paste attack the tester itself?
- It is built so it cannot do much. The patterns are found by code, not by a model. Jev only returns yes/no probabilities, so it cannot write output the text could steer, and the text is passed as labeled data.
- Does it run my prompt or call any tool in it?
- No. The text is only read and scored. Nothing in it is executed, fetched or followed, and nothing is saved.
- How is the injection risk score worked out?
- Jev answers one question about the whole text, which carries 60% of the score. The other 40% is how strongly the eight strongest patterns fire, weighted by how much each gives away. A quoted pattern only counts if the exact words are in the text. Hidden characters count at once, since no reader can see them.
- Why only 5 tests a day?
- Each test is a model call that costs real money. The limit keeps the tool free for everyone.
Built on Jev
This is the guardrail use case: many small yes/no judgments, each with a probability, in one fast call, cheap enough to run on every input. Also try the AI Writing Checker or see what Jev is.