Skip to content
Made with Jev

Jev agent prompts

Copy-paste prompts that wire TypeSafe Jev into coding agents: model routing, token cost, tool and budget gates, and quality checks. Use cases come first, then setup. Each prompt tells your agent to read the TypeSafe docs before it writes code, so it does not invent fields. Keep your API key in an environment variable. Never paste it into chat or commit it.

Quick install

  • npx skills add typesafe-ai/skills --skill typesafe-ai
  • claude plugin marketplace add typesafe-ai/skills

24 prompts

Pick model and reasoning effort per task

Jev scores difficulty and precision, a table in code maps that to a model lane, and the model stays put while the task is the same, so the cache survives.

Claude CodeCursorCodex+3 more
markdown
Pick the model and reasoning effort for each task in this project with Jev, instead of sending everything to the biggest model.

If the goal is routing turns inside the Claude Code or Codex CLI, check https://github.com/gargpratyush/jev-router first (open source, MIT). It wraps the real CLI and lets Jev pick the model per turn. If it fits, set it up and stop here. If not, say why and continue.

1. List the models and effort levels this project can call, with prices from config or the provider's page. Define lanes, cheapest first, such as SMALL_LOW, SMALL_HIGH, LARGE_LOW, LARGE_HIGH. Use this project's real model names.
2. Before each task, ask Jev in one request, with only the state the task needs (see https://docs.typesafe.ai/primitives):
   - Score: How hard is this task? Write a 4-level rubric for this project's work, with one example per level.
   - Noul: Does this task need exact precision (code that must compile, money, legal text, writes to data)?
   - Noul: Is this a different task from the previous turn? Send the previous task too.
3. Map the answers to a lane in code, in one table. Needing precision raises the lane one step. Low confidence picks the higher lane. Describe each option by what it can do; a bare label carries nothing.
4. Keep the same model for the session while the task has not changed. Switching models loses the prompt cache, so switch only when the lane changes.
5. Set a deadline on the Jev call. A late answer counts as no answer, and the default lane takes the task.
6. Let the user force a lane.
7. Run it in shadow first: Jev picks a lane while the current setup still does the work. Log the disagreements and show them to me.
8. Then run 30 real tasks through the router and through the current setup. Report lane, cost, latency, and every task where the cheaper lane failed a check. Do not make it the default until I approve.

Source: Made with Jev · Same idea as Jev Router, for your own providers

Build one production Jev decision, end to end

Community

Turns your agent into a Jev engineer: read the official docs, pick ONE decision the agent makes hundreds of times, and ship runnable code with an escalation rule, failure handling and a cost comparison. Turn on web access first.

Claude CodeCursorCodex+3 more
Prompt
You are an expert AI AGENT ENGINEER specializing in TypeSafe AI's Jev. Don't explain Jev to me. Build with it and hand me something I can run.

First read the official docs at docs.typesafe.ai (Python SDK: typesafe-sdk) and use only APIs that exist there. Verify before you write code.

Then pick ONE production use case where an agent makes the same bounded decision hundreds of times: routing, next-action choice, action gating, filtering, or deciding when to escalate to a stronger LLM.

Explain what problem it solves and why Jev, not a full LLM, should make that decision.

Then build it and give me:
1. Architecture: input > state > Jev decision > action > new state. Mark what Jev handles vs. what needs Claude.
2. The Jev decision: exact state in, allowed options, structured output.
3. Full runnable code. Simplest stack, keys in env vars, install steps.
4. The agent loop: observe, decide, execute, update, repeat.
5. The escalation rule for when Claude takes over from Jev.
6. Failure handling: low confidence, failed action, changed environment, invalid action, stuck loop.
7. A realistic example with 3+ decision cycles showing state and decision.

Prioritize a working system over a flashy demo. Explain the code in plain English.

End with WHY JEV: full LLM for every decision vs. Jev + Claude, compared on cost, latency, consistency and scale. Then a section titled THE ONE-SENTENCE IDEA.

NEGATIVE PROMPT
- Don't invent APIs, SDK methods, config options or benchmark numbers. Don't fake results.
- Don't force Jev into open-ended reasoning or text generation.
- Don't give a toy demo or a vague explanation.
- Don't give several use cases. ONE.
- Don't use pseudo-code. Everything must run.

Source: @mikenevermiss · Transcribed from the image in the post

Audit your agent, install one typed decision, prove it

Community

One paste: the agent maps where a big model does bounded work, ranks the candidates, wires in one Jev decision, tests it in shadow mode, and keeps it only if the numbers move.

Claude CodeCursorCodex+3 more
Prompt
Audit the agent environment you are running in right now. Find where open-ended model calls are doing bounded work, then wire in one typed decision and prove it helped. Work through every step. Do not stop after handing me a plan.

1. READ WHAT YOU CAN SEE
Name three places in this project where a large model is called only to choose, score, classify, verify, retry, route or ask for approval. One short example each, quoted from what is actually in front of you. Do not claim access to files or chats you cannot open.

2. NAME THE ENVIRONMENT
State which agent you are: Claude Code, Codex, Grok Bot, Cursor, another coding agent, or a plain chat with no tools. Report:
  AGENT, PROJECT PATH, OS, CAPABILITIES (terminal, files, network, persistent instructions).
If you cannot modify my system, say so and stop at step 4 with a handoff a capable agent could run.

3. MAP THE WASTE
Trace one real pass: request, context, model, tools, result, review. For each candidate decision, give frequency per day, context size, latency, what a wrong answer costs, and whether a person reviews it.

4. RANK THE CANDIDATES
One table: DECISION, CURRENT PATH, PRIMITIVE, EXPECTED VALUE, RISK, IMPLEMENTATION COST, PRIORITY. Keep the three highest. Reject anything that still needs open writing, deep reasoning, or a rule code can compute exactly.

5. READ THE LIVE DOCS
Read the current documentation before writing any call: the index first, then only the pages this task needs. Preserve my existing configuration. Never put an API key in chat, logs, screenshots, source control or generated files. Ask me to set it privately.

6. BUILD ONE CONTRACT
For the top candidate write one decision contract: the state, the exact question, the answer type, the options or criteria, the confidence routing, the fallback, the human-review boundary, the data that must never be sent, and one realistic test fixture.

7. TEST IT THREE WAYS
Run in shadow mode first, acting on nothing. Then run three cases: one obvious, one ambiguous, one that must escalate. Show me the answers and what the code did with them. An SDK import, a mocked response or a successful raw request is not a working integration.

8. MAKE IT THE DEFAULT
Add a short project-scoped rule to the instruction file this agent actually reads, preserving what is already there. The rule should say: before expensive research, before repeating a failed approach, before loading several tools, before choosing between materially different routes, and before a consequential action, ask the bounded question first. Irreversible actions stay behind my confirmation.

9. PROVE IT
Report: what you found, what you changed, files touched, tests run, whether a live call succeeded, the estimated saving, what still needs evaluation, and how to roll it back in one command. Never call this done without evidence that the agent used the decision during a real task.

10. MEASURE, THEN KEEP OR REVERT
Run five real tasks on the old path and the new one. Compare tokens, latency, cost, retries and how often I had to step in. Keep the contract only if it wins without hiding failures. If the numbers do not move, roll it back rather than defending it.

Source: @zodchiii · Transcribed from the image in the post

Find where you are wasting tokens

Audit every LLM call in a project, mark the ones that only make a decision, and rank what a typed Jev question would save. No code changes until you pick.

Claude CodeCursorCodex+3 more
markdown
Audit this project for tokens spent on decisions instead of on writing.

Many LLM calls only answer yes or no, pick one option, or give a rating, then the code throws the prose away. A decision model like Jev answers those as typed values: Noul (probability of yes), Choice (one label from a list) and Score (a level on a rubric). Read these first:
- https://docs.typesafe.ai/concepts/use-case-map
- https://docs.typesafe.ai/primitives

1. Find every LLM call. For each, record: file and function, model, what the prompt is for, how the output is used (parsed to a boolean, matched to a label, turned into a number, or shown to a person as text), and how often it runs (per request, per item, per loop turn). Use logs or tests to estimate volume. Write "unknown" when you cannot.
2. Mark each call:
   - DECISION: the code only uses a boolean, a label or a number from the output.
   - MIXED: it decides, then also writes text.
   - GENERATION: the text is the product.
3. For each DECISION call, write the Jev question that would replace it: the type, the instruction, the options or rubric, and the state it needs. Send only the state that question needs.
4. Look for other waste: the same context re-sent every loop turn, a frontier model on a classify step, retries with nothing changed, long system prompts on small jobs, agent loops with no stop condition.
5. Rank the findings by estimated tokens saved per day, in a table. Mark every estimate as an estimate and show how you got it.

Do not change code yet. When I pick items, do one at a time behind a flag, run old and new on the same real examples, and report the agreement rate and the cost before and after.

Source: Made with Jev · Uses TypeSafe's use-case map and question types

Cheap model first, Jev checks, expensive model only when flagged

A verified cascade: a small model answers, Jev asks narrow yes/no questions about that answer, and only flagged items go to the big model.

Claude CodeCursorCodex+3 more
markdown
Add a cheap-first cascade to [THE LLM CALL I NAME] and check every cheap answer with Jev before it is used.

The pattern comes from TypeSafe's cascade cookbook: a cheap model answers first, Jev checks that answer with narrow yes/no questions, and only flagged items go to the expensive model.
- https://docs.typesafe.ai/cookbooks/sde_cascade
- https://docs.typesafe.ai/primitives/noul

1. Read both pages. Use the field names from the docs, not from memory.
2. Keep the current expensive call as the fallback. Put a cheap model in front of it.
3. Write Noul questions where a high probability means "escalate". Start from the cookbook's and adapt them to this task:
   - Is this value unsupported by, or absent from, the source?
   - Does the source fail to report the thing this field asks for?
   - Is the empty result wrong, because the source does contain the information?
   Send Jev the source and the cheap answer as state. Nothing else.
4. Escalate when ANY flag is over the threshold. Use the max, not the average, so one confident red flag is enough. Put the threshold in one constant; the cookbook starts at 0.7.
5. Log per item: cheap answer, each flag's probability, escalated or not, final answer, cost.
6. Build a labeled set of at least 50 real items. If none exists, ask me to label them. Report escalation rate, accuracy and cost for cheap-only, cascade and expensive-only. Show the accuracy at several thresholds before you pick one.

If cascade accuracy is clearly below expensive-only, do not ship it. Tell me the gap in numbers.

Gate your agent's tool calls: allow, ask or block

Before each tool call, Jev checks scope, data exposure and how hard it is to undo. Thresholds and the always-ask list live in code, not in the prompt.

Claude CodeCursorCodex+3 more
markdown
Put a Jev gate in front of this agent's tool calls. Each call is allowed, sent to me for approval, or blocked.

Read first:
- https://docs.typesafe.ai/cookbooks/llm_guardrails
- https://docs.typesafe.ai/cookbooks/function_calling
- https://docs.typesafe.ai/primitives/noul
- https://docs.typesafe.ai/primitives/score

1. List every tool this agent can call. Tag what each can do: read-only, writes local files, calls an external API, sends a message, spends money, deletes, deploys or publishes. Show me the list and wait for my corrections.
2. Read-only tools skip the gate. Irreversible ones (delete, pay, publish, deploy, message another person) always need my approval, whatever Jev says. Put that rule in code.
3. For everything else, before the call, send Jev the state: my request, the pending tool name and arguments, and the last few steps. Ask in one request:
   - Noul: Does this tool call go beyond what the user asked for?
   - Noul: Could this call change or expose data the user did not mention?
   - Score: How hard is this call to undo? Rubric: 0 trivial, 1 easy, 2 hard, 3 impossible.
4. Route with two thresholds per question, kept in one config: at or above the action threshold, block; between the review and action thresholds, ask me; below review, allow. A high undo score can turn an "ask" into a "block".
5. Log every decision: tool, arguments with secrets redacted, probabilities, route.
6. Test with 10 calls that should pass and 10 that should be stopped, written from this project's real tools. Show me which ones got through.

Fail closed. If the Jev call errors or times out, ask me instead of running the tool. Read TYPESAFE_API_KEY from the environment and never print it.

Compact long agent sessions with Jev, safely

Cut context by keeping or dropping whole tool results instead of rewriting them in a summary. Pinned messages, retrievable originals, a task-boundary check, and a replay that proves the agent still finishes the job.

Claude CodeCursorCodex+3 more
markdown
Add Jev-based compaction to this agent's long sessions, and prove the agent still finishes its work afterwards.

Jev compaction does not write a summary. It keeps the original words in order and only drops or truncates. For each tool call, Jev answers two questions: should the call stay, and should its result stay verbatim. The risk is the opposite of a summary: a whole result can go that the agent needed later.

Read first:
- https://docs.typesafe.ai/primitives/noul
- https://docs.typesafe.ai/concepts/state

Check these open-source projects first (both MIT). If one fits this agent, use it and keep steps 7 and 8 as the test:
- https://github.com/tamaratran/fast-jev-compaction: per-call keep or drop, with pinned messages and a token budget
- https://github.com/kunchenguid/compact-adviser: asks Jev whether the session is at a safe task boundary

1. Show me how this agent compacts today (summary, truncation, nothing) and where the context limit is hit. Find 3 real long sessions I can replay. Keep the originals.
2. Build the state from the conversation, oldest first. Replace each tool result with a short note (tool name, size, "omitted"), and cut long tool inputs. Jev takes up to 32K tokens of state per request; split the questions across requests when needed.
3. Pin what must never go: the first message and the newest few turns. Pinned calls are never sent to Jev.
4. For every unpinned tool call, ask Jev in parallel:
   - Noul: Is this tool call still needed for the current task?
   - Noul: Must its result stay word for word, rather than as a short note?
   Send the current task with the state, so the questions are about this task. Keep the thresholds in one constant.
5. Compact only at a task boundary. Before compacting, ask: Noul, is the session at a point where the previous task is finished and the next one has not started? If not, wait.
6. Keep every dropped result retrievable by ID, so the agent can fetch it again instead of re-running the command.
7. Watch for loops. If the agent repeats a step it already did, log it as a compaction miss and show me the dropped result it needed.
8. Replay the 3 sessions with and without compaction. Report tokens before and after, cache-write cost if the provider reports it, loops, and whether the agent finished the same task.

Tokens removed is the easy number. Task success after compaction is the one that matters. Do not turn it on by default until the replays pass.

Source: Made with Jev · Built from our Jev compaction guide and the plugins it covers

Build a second brain in Obsidian with Jev

One inbox for saved posts, notes and ideas. Jev files each item under your own projects, flags duplicates, and scores which notes matter when you start a task, so your writing agent starts with your material.

Claude CodeCursorCodex+3 more
markdown
Build me a second brain in my Obsidian vault, with Jev making the sorting and relevance decisions. Start with ONE project and the material I already collected for it.

Read first:
- https://docs.typesafe.ai/primitives (Choice, Noul, Score)
- https://docs.typesafe.ai/cookbooks/semantic_find
- https://docs.typesafe.ai/cookbooks/rerank_typesafe

1. Ask me for the vault path and the one project to start with. Create an Inbox folder. Saved posts, articles, meeting notes and ideas land there as markdown files.
2. Ask me for my active projects and interests, with one sentence each on what belongs there. Those sentences are the Choice criteria. Do not invent categories.
3. For each new inbox item, write frontmatter first: source URL, date saved, type. Never drop the original link or date; I must always be able to check where a note came from.
4. Sort: one Jev Choice, the item as state, my projects as options, plus "none of these". Below the confidence threshold (one constant), leave the item in Inbox and list it for me instead of guessing.
5. Dedupe: search the vault for the most similar notes (plain text search or embeddings, whatever this setup already has). Then ask Jev for each close match:
   - Noul: Is the new item a duplicate of this note?
   - Noul: Does the new item add something this note does not have?
   Duplicate and adds nothing: link it to the existing note and archive it. Adds something: keep both and link them.
6. Retrieve: when I start a task, search the vault for candidates, then ask Jev to Score each one: how useful is this note for the task? 0 not useful, 1 background, 2 useful, 3 must read. Send only the task and the note as state.
7. Hand the top notes (score 2 or more, capped) to the writing agent before it starts, with their source links. Example: for a proposal, bring back what the customer said on the call and the case study that answers their objection.
8. Before sending notes to the Jev API, show me what leaves my machine and let me exclude folders. Read TYPESAFE_API_KEY from the environment and never print it.
9. Test on my next real task. Show me what it brought back and what it missed, then tune the thresholds. Do not add a second project until I approve the first.

Source: @VibeMarketer_ · Adapted from the setup in the post

Budget gate before expensive runs

Hard limits in code, then Jev decides if a browser run, long loop or big-context call is worth it, or if a cheaper path answers the same question.

Claude CodeCursorCodex+3 more
markdown
Add a budget gate in front of this agent's expensive operations: browser sessions, long agent loops, batch jobs, big-context calls and paid APIs.

One rule first: numbers stay in code. Jev reads instructions literally and is not a calculator. Code computes cost, counts and dates. Jev answers only the judgment questions.

If the expensive part is a browser agent, look at https://github.com/browser-use/jev-ultrafast first (open source, MIT). It already has Jev choose each browser action. The gate below still decides whether the run should start at all.

1. Find the expensive operations. For each, write how to estimate its cost before it runs (tokens x price, pages x time, calls x rate). Take prices from config or the provider's page. Do not invent them.
2. Add hard limits in code: per-run budget, per-day budget, max loop turns, max retries. Over a hard limit, stop. No model is asked.
3. Under the limits, before the run, ask Jev in one request (see https://docs.typesafe.ai/primitives):
   - Choice: What is enough for this request? CHEAP_PATH (search, cached data, a small model), FULL_RUN, or ASK_USER. Describe each option in one sentence for this project.
   - Noul: Was this same task already done in this session? Send the recent task log as state.
   - Score: How much would the answer change if we skip the expensive step? 0 not at all, 1 a little, 2 a lot, 3 it cannot be answered without it.
4. Act only on confident answers. Keep the confidence threshold in one constant. Below it, take the cheaper path and tell the user what was skipped.
5. Log estimated and actual cost per run, and how often each route fired.

After a day of use, report spend before and after, runs skipped, and any skipped run that should have gone ahead.

Source: Made with Jev · Hard limits in code, judgment in Jev

Retry, change approach or escalate after a failure

Stop agents from retrying the same broken thing. Code collects the evidence, Jev picks the next move, and low confidence goes to you.

Claude CodeCursorCodex+3 more
markdown
Teach this agent to decide what to do after a failure, instead of retrying the same thing.

If the workers are Codex or OpenCode, check https://github.com/thruwire/foreman first (open source, MIT). It supervises coding workers with Jev: done or not, stuck, off track, needs a human. If it covers this, set it up instead and tell me what it does not cover.

Read first:
- https://docs.typesafe.ai/primitives/choice
- https://docs.typesafe.ai/patterns/confidence-routing

1. Find where this agent handles failures: failed tool calls, failing tests, errors, timeouts, rejected outputs. Show me what it does now.
2. Collect the evidence that code gets for free: error text, exit code, test output, attempt count, and what changed since the last attempt (git diff or argument diff). Never ask Jev something software can tell you.
3. Hard rules in code first. Same error and nothing changed since the last attempt: never retry the same way. Attempt count over the limit: stop and report.
4. Otherwise ask Jev one Choice, with the evidence as state:
   - RETRY_SAME: a transient error (network, rate limit, flaky test)
   - CHANGE_APPROACH: the approach is wrong; the error repeats or points at the design
   - ESCALATE_MODEL: the task is harder than the current model or effort can handle
   - ASK_USER: missing information, a permission, or a decision only the user can make
   - STOP: the task cannot be done as asked
   Write one clear sentence per option for this project.
5. Below the confidence threshold, kept in one constant, choose ASK_USER.
6. Log every failure with the choice, the confidence and the outcome. After 20 failures, show me how often each choice led to success.

Source: Made with Jev · Uses TypeSafe's confidence-routing pattern

Quality gate before anything gets published

A scorer script checks each draft against the brief, flags facts that are not in the sources and catches leaked secrets. The agent cannot publish around it.

Claude CodeCursorCodex+3 more
markdown
Add a quality gate that scores drafts before this agent publishes anything: posts, emails, docs, release notes.

1. Write a scorer script in this project's language. It takes a draft, plus the brief and sources if there are any, and asks Jev in one request (field names from https://docs.typesafe.ai/primitives):
   - Score: Does the draft do what the brief asks? Rubric 0 to 3; describe each level.
   - Noul: Does the draft state facts, numbers or quotes that are not in the brief or the sources? Send the sources as state.
   - Noul: Does the draft contain a secret, a private email address or internal-only information?
2. Keep the thresholds in the script, in one place. Exit 0 on pass and 1 on fail, and print JSON with every answer.
3. Add this to the agent's instruction file (CLAUDE.md, AGENTS.md or the equivalent):
   - Never publish a draft without running the scorer.
   - On exit 1, print the JSON, rewrite the draft and score it again.
   - After three failures, stop and show me all three drafts with their scores.
   - Do not edit the thresholds and do not publish around a failing score.
4. Test with 5 good drafts and 5 drafts you make bad on purpose. Show me the scores.

Source: Made with Jev · The exit-code pattern is from OpenTweet's Claude Code guide

Filter RAG passages and check citations

Score each retrieved passage before it enters the prompt, drop the noise, and check every cited claim against its passage after the answer is written.

Claude CodeCursorCodex+3 more
markdown
Filter retrieved passages with Jev before they go into the prompt, then check the citations in the answer.

Read first:
- https://docs.typesafe.ai/cookbooks/classifying_rag_passages
- https://docs.typesafe.ai/cookbooks/citation_check

1. After retrieval, ask Jev about every passage in one request: Score, how directly does this passage answer the question? 0 off-topic, 1 related, 2 partly answers it, 3 answers it.
2. Keep passages at or above a threshold kept in one constant, and cap how many go in. If none pass, tell the user the sources do not cover it instead of answering.
3. After the answer is written, for each cited claim ask: Noul, is this claim unsupported by the passage it cites? Remove or flag the claims over the threshold.
4. Test on 30 real questions. Report prompt tokens per answer before and after, and answer quality by this project's existing checks.

LLM → JEV → AGENT install prompt

Community

The viral setup prompt from @0xCodila: install Jev into your agent, check the environment, persist instructions and prove it works. We cut the step that cloned the author's own repo, so it installs from official TypeSafe sources only.

Claude CodeCursorCodex+1 more
Prompt
I want you to install Jev by TypeSafe AI into the AI agent or LLM environment I am using right now, then make it available for future tasks in this project. I should be able to create jev-based tasks always. You should call Jev when it tactically could improve the work, without requiring me to mention Jev.

Complete the setup and verify it: work through the steps below after giving me a plan.

**1. Understand my work**

Read the conversation and project context you can actually access. Identify three specific situations in my work where Jev could help decide what to do before expensive agent work begins. Give one short example for each. Do not claim access to other chats or files you cannot see.

**2. Check this environment**

Before running any installation command, identify the exact environment you are currently running in: Claude Code, Codex, Grok Bot, Muse, Cursor, another coding agent, or an ordinary LLM chat.

Check this environment:
- AGENT: detected agent name / unavailable
- PROJECT: detected project path / unavailable
- OS: uname -s / equivalent / unavailable
- CAPABILITIES: git / python3 / npm / ~/.codex/ / ~/.cline/ / persistent instructions / unavailable

**3. Use verified sources**

Use only these sources for install steps:
- https://docs.typesafe.ai/introduction/quickstart
- https://docs.typesafe.ai/agent-skill
- https://docs.typesafe.ai/primitives

The official TypeSafe skill teaches agents how to call Jev.

**4. Install and test**

If you have terminal and file access, install the official TypeSafe skill. In Claude Code:

claude plugin marketplace add typesafe-ai/skills
claude plugin install typesafe@typesafe-ai

In other agents:

npx skills add typesafe-ai/skills --skill typesafe-ai

Then guide me to create an API key at https://console.typesafe.ai/settings/keys. Tell me exactly how to set TYPESAFE_API_KEY in this environment (I should not paste it into the chat).

Jev returns typed Choice, Score, and Noul answers. Make one small test call and use the returned value exactly as the docs describe it; do not invent fields.

**5. Make future use natural**

Add a concise, project-scoped introduction to the persistent instruction file supported by the detected agent: CLAUDE.md for Claude Code, AGENTS.md for Codex, or the equivalent for this agent. The instruction must tell the agent:

- Jev is a decision model. Call it for consequential decisions (model routing, tool selection, retry, escalation), not for writing code.
- Authenticate: read TYPESAFE_API_KEY from the environment. Never print, expose, log, or commit this key.

Connect this instruction to a tool invocation this agent actually performs. Do not claim that doing so alone guarantees Jev will run.

**6. Prove it works**

Test three tasks in this environment: a simple question that should skip Jev; a substantial task with a real routing decision; a repeated failure where Jev could recommend changing approach.

Return a compact report of what was installed, what tests passed, and what still needs action. Give three personalized prompts to try next. Never mark the integration complete without evidence that tests passed.

Source: @0xCodila · Edited: official sources only

Install TypeSafe skill (any agent)

Official TypeSafe skill installation prompt. Works with Claude Code, Codex, Cursor, and most coding agents.

Claude CodeCursorCodex+2 more
Prompt
Install the TypeSafe skill. If you're in Claude Code, run `claude plugin marketplace add typesafe-ai/skills`, then `claude plugin install typesafe@typesafe-ai`. If you're in another agent, run `npx skills add typesafe-ai/skills --skill typesafe-ai` and select your agent. Use one installation method. You can read the skill directly at https://github.com/typesafe-ai/skills/blob/main/skills/typesafe-ai/SKILL.md (raw: https://raw.githubusercontent.com/typesafe-ai/skills/main/skills/typesafe-ai/SKILL.md). Then use the TypeSafe skill when working on this project.

Source: TypeSafe AI · MIT license

Find opportunities for Jev

Explore your codebase for places where intelligent judgment could replace fragile parsing or complex conditionals.

Claude CodeCursorCodex
Prompt
Using the TypeSafe skill, explore the project and find opportunities for using intelligent judgement to stand in for complex parsing or other fragile code.

Source: TypeSafe AI

Run Jev experiments

Run experiments with your TypeSafe API key and propose improvements based on results.

Claude CodeCursorCodex
Prompt
Using the TypeSafe skill, run some experiments using the TypeSafe API key that I've exported to TYPESAFE_API_KEY. Propose changes based on the most promising results.

Source: TypeSafe AI

Analyze against TypeSafe cookbooks

Check your code against TypeSafe cookbooks to find refactoring opportunities.

Claude CodeCursorCodex
Prompt
Using the TypeSafe skill, analyze my code and see if there are any applicable cookbooks (https://console.typesafe.ai/docs/cookbooks) that show how I could refactor my code to be less fragile or complex.

Source: TypeSafe AI

Install the Jev MCP server

One server, every client. Paste the prompt into any MCP-capable agent, or use the config for yours. Keep TYPESAFE_API_KEY in the server's env: some clients strip inherited variables.

Claude CodeCursorCodex+3 more
Prompt
Install the Jev MCP server for me. The package is @jkudish/jev-mcp on npm and the server command is `npx -y @jkudish/jev-mcp`; register it as an MCP server with your client. Check whether TYPESAFE_API_KEY is already set in the server environment; if not, walk me through setting it up without pasting the key into the chat (I can create one at console.typesafe.ai/settings/keys). When it's registered, ask if I'd like to try a claim verification, and when we do, show me the verdicts and cost. Full instructions: https://github.com/jkudish/jev-mcp#readme

Source: jkudish · MIT license

Claude Code preflight plugin

Install Jev preflight plugin for Claude Code. Requires v2.1.257+.

Claude Code
bash
claude plugin marketplace add muse0509/jev-preflight
claude plugin install jev-preflight@jev-preflight
claude plugin enable jev-preflight@jev-preflight

# Then open /plugin and set typesafe_api_key

# Requires Claude Code 2.1.257+ · Git · Bash · TypeSafe API access

Source: @muse_jp_sol

Cline jev-browser plugin

Official Cline jev-browser plugin setup. Requires Vercel AI Gateway key.

Cline
json
# 1. Customize → Marketplace → Plugins → search "jev-browser"
# 2. Save Vercel AI Gateway key to ~/.cline/plugins/cline-jev-browser.config.json:

{"gateway": {"apiKey": "..."}}

# 3. Restart Cline and ask for a browser task

Source: @cline · Official plugin

Codex → Luna → Jev orchestration

Multi-model coding orchestration: Luna writes code, Jev makes decisions, deterministic tools verify. Put in project root AGENTS.md or ~/.codex/AGENTS.md.

Codex
markdown
# CODEX → LUNA → JEV
Multi-model coding orchestration for reliable, efficient software engineering.

## 01 Default multi-model coding orchestration
For every software-engineering request, automatically operate as an architect and orchestrator.
Goal: complete work reliably using the minimum necessary model and reasoning level.

## 02 Jev decision model
Jev is available as a decision/control model.
Authentication: read TYPESAFE_API_KEY from the environment. Never print, expose, log, or commit this key.

Use Jev ONLY for bounded decisions:
- task complexity classification
- reasoning-level routing
- continue vs stop
- retry decisions
- escalation decisions
- completion assessment

Do NOT use Jev for:
- writing code
- generating patches
- architecture design
- anything deterministic tooling can answer

## 03 Initial routing
Before implementation, classify the task:
- SMALL → Luna, low reasoning
- MEDIUM → Luna, medium reasoning
- HIGH → Luna, high reasoning
- ESCALATE → Sol, high reasoning

Always use the lowest sufficient lane.

## 04 Deterministic verification (critical rule)
Prefer deterministic evidence over AI judgment whenever possible.
- Tests → run the test runner
- Compilation → run the compiler
- Types → run the type checker
- Lint → run the linter
- Changed code → run git diff

Never ask Jev to determine something software can determine.

## 05 Agent loop
After each implementation cycle:
1. Inspect the actual diff
2. Run relevant deterministic checks
3. Gather concise evidence
4. Use Jev for any remaining judgment
5. Choose: CONTINUE / RETRY / VERIFY / ESCALATE / COMPLETE

## 06 Escalation path
Luna Low → Luna Medium → Luna High → Sol High

Escalate only when evidence warrants it:
- repeated attempts fail
- tests keep failing
- security-sensitive code changed
- architectural uncertainty remains
- Jev confidence is below threshold

Do not escalate because a stronger model is available.

## 07 Completion
Only report completion when:
- requested behavior is implemented
- deterministic checks pass
- diff matches requested scope
- no unresolved failures exist

Never hide failed verification.

Source: @UT_Codex · Public X post

Grok Bot usage router with Jev

Community

Ask Jev before expensive browser runs. Noul guard for irreversible actions. Paste into Grok Bot memory/instructions.

Grok Bot
markdown
# Grok Bot + Jev decision layer

Before running browser automation or expensive tool calls, use Jev to decide:

## Usage routing
Before a browser run:
1. Use Jev Choice to classify: SIMPLE_SEARCH | NEEDS_BROWSER | SKIP
2. If SIMPLE_SEARCH: use lightweight tools
3. If NEEDS_BROWSER: proceed with automation
4. If SKIP: explain why and ask for clarification

## Research decisions
After gathering data:
1. Use Jev Score to rate completeness on a 0-3 rubric (0 nothing useful, 1 gaps, 2 mostly there, 3 enough to answer)
2. If confidence < 0.8: escalate to human
3. If score < 2: continue research
4. If score >= 2: proceed to synthesis

## Action guard (Noul)
Before irreversible actions (post, delete, purchase, merge):
1. Use Jev Noul: "Is this action safe to run without human approval?"
2. Noul returns a probability of yes from 0 to 1
3. If it is below 0.9: require human approval
4. Keep the 0.9 threshold in code, not in this prompt

## Key rules
- Jev answers bounded questions; Grok writes and reasons
- Never use Jev for code generation or open-ended tasks
- Always check Jev confidence; escalate when uncertain
- Read TYPESAFE_API_KEY from environment; never expose it

Source: Community synthesis · Based on Grok Bot + Jev builds

Muse + Jev setup

Community

Shadow mode: Jev decides proceed vs skip before tool execution. From viral @0xCodila workflow.

Muse
markdown
# Muse → Jev → Execution

Use Jev as a decision layer before tool execution:

## Shadow mode workflow
1. Muse proposes an action or tool call
2. Before execution, ask Jev: "Should this action proceed?"
   - Context: current task, proposed action, risk level
   - Return: Choice(PROCEED | SKIP | REVIEW)
3. Route based on Jev decision:
   - PROCEED: execute immediately
   - SKIP: explain reasoning, propose alternative
   - REVIEW: show me the plan, wait for approval

## When to invoke Jev
- Before browser automation
- Before file modifications in sensitive directories
- Before API calls with side effects
- Before multi-step workflows
- When previous attempts failed

## Jev setup
- Install: `npx skills add typesafe-ai/skills --skill typesafe-ai`
- Export TYPESAFE_API_KEY (get at console.typesafe.ai/settings/keys)
- Never paste API key in chat

## Key principle
Jev makes bounded decisions; Muse does creative work. Use Jev for "should I" not "how do I".

Source: @0xCodila · Muse + Jev execution flow

OpenAI Dots connection prompt

Community

Frame Jev as paste-into-agent instructions for decision-making before agent actions. For Dots or similar OpenAI agents.

Dots
markdown
# Decision layer for AI agent actions

Before taking consequential actions, pause and decide:

## When to decide
- Before sending messages or emails
- Before creating, modifying, or deleting content
- Before purchasing or financial transactions  
- Before scheduling or committing to plans
- Before sharing information externally

## Decision protocol
For each action above:
1. Describe the proposed action clearly
2. List potential consequences (positive and negative)
3. Assess confidence in the appropriateness of this action
4. If confidence is high (>0.9): propose the action and explain
5. If confidence is medium (0.7-0.9): show options and ask
6. If confidence is low (<0.7): explain uncertainty and defer

## Integration option
If your agent supports external decision APIs:
- TypeSafe Jev provides typed Choice/Score/Noul endpoints
- Get API key at console.typesafe.ai/settings/keys
- Use for bounded decisions only (not for generating content)
- Check confidence scores; escalate uncertain decisions

## Core principle
An agent that decides before acting is more reliable than one that acts immediately. Build in a thinking step.

Source: Made with Jev · Pattern for decision-aware agents; no official OpenAI Dots API assumed

Built for Made with Jev. Spot a prompt we should add? Share it and ship a PR.