Can Jev read images?
No. Jev accepts text and structured state, so a picture has to become text before the decision. OCR, a vision-language model, a detector or an image encoder does that stage; Jev decides over what it produces. This page is the routes, the limits each one publishes and where they break, and the 8 builds in this directory that already give a text-only model eyes.
Updated 2 Oct 2026 · by Made with Jev

In short
- Jev has no image input. Everything below is a way to turn the picture into the text Jev already takes.
- OCR is the smallest useful first step for anything with writing in it, and it reads visible text and nothing else.
- A “free API” is a capped allowance: 1,000 Cloud Vision units a month, 25,000 OCR.space requests, 5,000 Azure transactions. None of them is unlimited throughput.
- A vision-language model writes a richer record, but it can also answer the question directly. That is the baseline Jev has to beat, and sometimes it does not.
- Scores do not stack. An OCR score, a detector score and Jev’s confidence are different numbers, so a confidently misread image can be routed just as confidently.
What Jev accepts, and what it does not
The current documentation lists Jev’s input as text and structured state, with no native image, audio or video support, and explicitly recommends preprocessing anything non-text into text or structured fields. Customer fine-tuning and LoRA adaptation are not offered either: you can build an adapter around Jev, but you cannot upload a vision adapter into its hosted weights.
So Jev answers bounded questions about text you hand it — a choice among options, a score against a scale, or a yes/no probability. It does not describe the picture itself. The division is clean: the visual stage reads the image, and Jev decides how the evidence affects the workflow. The routes below follow keno’s walk-through, which measured the limits on each one; the first call itself is on how to use Jev and the price on the pricing page.
If you would rather skip the extra stage, there is a decision model that takes pictures directly. Cloudflare’s Clef reads text, JSON, images or video, returns a probability for each option, and has an API compatible with Jev. The weights are Apache 2.0. How it compares is on Jev vs Cloudflare Clef, and who should pick which is on the Clef alternatives page.
Jev OCR: reading the text in an image
For screenshots, scans, receipts and photographed notices, OCR is often the smallest useful first step. Several of the options are local and free to run.
| Option | What it contributes | Where it fits |
|---|---|---|
| Tesseract | Recognised text, plus TSV and hOCR | A straightforward local baseline |
| EasyOCR | Text, bounding boxes and scores; CPU mode | A multilingual prototype |
| PaddleOCR | OCR and document-processing tools | Documents where layout and structure matter |
| Apple Vision | Text observations from images | Native apps on supported Apple platforms |
| Google ML Kit | On-device text recognition | Android and iOS apps, offline |
Keep the context around the extracted text, because a bare list of words loses it: which attachment it came from, whether recognition was incomplete, and which field or region held it. Without that, a price and a discount read the same, and so do a warning and a button label.
OCR has a clear boundary. It reads visible writing, so it will not explain a text-free product photograph or tell you what is happening in a street scene. Those need the stages further down.
Hosted OCR and vision APIs: what a free tier really means
“Free API” usually means a limited allowance, and each provider counts something different — requests, pages, features or tokens.
| Service | What can become Jev input | Published free access |
|---|---|---|
| Google Cloud Vision | OCR text, image labels, detected objects | First 1,000 units of each feature a month; billing required |
| Gemini Developer API | Descriptions or structured visual observations | Selected models have a free tier; limits vary by model and project |
| OCR.space | Recognised text, optional word positions | 25,000 requests/month, 500/day/IP, 1 MB/file, 3 PDF pages |
| Azure Vision F0 | OCR and supported image analysis | 5,000 transactions/month, 20/minute; not a long-term default |
| Hugging Face Inference | Outputs from supported image models | $0.10 monthly credits; an experiment allowance, not throughput |
Two footnotes worth keeping. Azure’s Image Analysis 4.0 is deprecated and scheduled to retire on 25 September 2028, so it is not a version to build new work on. And the data rules differ: Cloud Vision says it does not use submitted content to train its models, while Gemini’s unpaid-service terms generally allow product-improvement use and human review, with regional exceptions. Check the applicable terms before sending private images.
Gemini is the interesting one for this page, because image understanding includes captioning and visual question answering. It can supply richer evidence than OCR alone — and it might also solve the final task directly, which is exactly why the last section here tests whether Jev adds anything.
Local vision-language models
When the question is about objects, relationships or layout, a vision-language model can produce an evidence record from the image. One downloadable option is Qwen3-VL, including a 2B Instruct checkpoint. Hardware needs depend on the checkpoint, the quantization, the image resolution, the context and the runtime, so a model-size label is not a memory guarantee.
Ask the visual stage for specific observations — readable text, visible damage, occlusion, uncertain details — rather than a general caption, because a caption can omit the exact detail a later decision needs. A richer record costs more to produce and transmit, so build its fields around the downstream question and keep uncertainty in rather than replacing missing evidence with a confident guess.
One caveat that catches people out: local processing does not make the whole workflow local. Anything you send onward to Jev leaves that local stage.
Detectors and embeddings: objects, positions and similar images
A detector returns a bounded vocabulary of objects and their locations. YOLOX is one open-source example with small model variants. For an inventory workflow it might identify boxes and equipment, if those classes are supported or it was trained for them, and code can count the detections and compute the spatial relationships. Jev then evaluates those observations against a written handling policy — which avoids asking it to do geometry from a list of coordinates. If the final task is simply “did the detector find a box?”, a rule may be enough and Jev is not needed.
An image encoder is the other local route. Google’s SigLIP 2 supports image-text retrieval, zero-shot classification and image embeddings, and the linked checkpoint is Apache-licensed; CLIP is the established reference for the family. Two useful patterns: compare a photo with candidate descriptions and pass the names and scores onward, or retrieve similar human-labelled examples and give Jev their labels and context. In both, do the vector comparison outside Jev and pass it the labels and scores, not the raw vector.
And read the similarity straight: a cosine similarity of 0.83 is not an 83% chance of being correct.
Colour, histograms and the limit of a few points
A few colour samples might flag a dark photo or sort packaging by colour. They will not reliably tell you which object is in the picture. A local image-processing stage can compute brightness, colour proportions, edge density or a coarse spatial grid — OpenCV’s histogram tools describe distributions of intensity or colour — and those are useful for detecting nearly blank scans, sorting standard packaging and flagging dark photos.
The boundaries are structural. A global histogram throws away spatial arrangement, different objects can share similar colours, and sampling a few points can miss small objects entirely. Keeping regions preserves some layout but still supplies no object semantics.
TypeSafe documents one more limitation here, and it is specific: Jev performs better with semantic colour names than with hex values, and cannot reliably judge closeness between RGB or hex values. So do the conversion in code and give Jev meaningful fields or buckets — an experimental input might read “image quality: dark, dominant colour: red, text detection: incomplete” rather than a hex code.
Check what your app already knows before you look at the picture
Sometimes the image is only a presentation of data your application already has. A document may carry a text layer, a webpage may expose useful DOM or accessibility text, a product asset may have verified catalogue metadata, and a barcode may encode an identifier that resolves to a structured record. Passing that data straight to Jev avoids unnecessary visual inference.
Filenames and unverified alt text are especially weak evidence, because they may be stale or simply wrong. For repeated, unchanged assets, cache an extraction keyed to the image and the extractor version, and invalidate it when the image changes. If you substitute a matching reference image’s label, say that it is retrieval from known material rather than recognition of new content.
The cascade: OCR, detector, vision model, human
The methods above can form a cascade — existing text, then OCR, then a detector or encoder, then a vision-language model, then human review. That is a design pattern, not a required sequence: a scene with no writing should not have to pass through OCR first. For a support screenshot, start with the text and escalate when it is unreadable or insufficient; for a product photo, start with the visual attributes.
A conceptual evidence record, the kind of state you would hand to Jev, might look like this. It is an example, not a measured model response or an SDK schema.
{
"attachment_type": "support_screenshot",
"extraction_method": "local_ocr",
"observed_text": ["Payment failed", "Update payment method"],
"customer_message": "I cannot renew my subscription",
"limitations": ["small footer text unreadable"],
"available_routes": ["billing", "account_access", "other", "manual_review"]
}Ask one narrow question per decision, keep observations separate from instructions, and retain an explicit fallback for missing evidence. Jev’s state documentation supports objects with descriptive fields, so the record above can carry the uncertainty through to the decision instead of hiding it.
Why a confident score does not repair missing evidence
There are several different numbers in a pipeline like this: an OCR recognition score, a detector score, an embedding similarity, Jev’s answer probability and Jev’s choice confidence. They are not interchangeable. TypeSafe describes confidence as a statistic derived from the answer distribution, not an independent visual verification.
So if the first stage confidently misreads an error message, the next stage can confidently route the wrong issue, and multiplying two scores together does not produce an end-to-end correctness probability. The fix is the same one every other page here gives for confidence: label a set of cases yourself, see where the confidence falls when the pipeline is right against when it is wrong, and set the threshold there. For a binary event, cases assigned a probability near 0.8 should contain that event about 80% of the time. That is what a calibration curve checks, and it is the only way to know whether a threshold means anything.
Give Jev eyes with Obsidian, Claude and MCP
The second half of this idea is not about pixels at all: it is about retrieval. Kev Builds Apps shows the setup in a full walkthrough. Install the Model Context Protocol server plugin in Obsidian, exit restricted mode, enable the plugin and copy the API key it generates. Then, from Claude Code, tell the agent to connect to your vault, name the graph, and pass the key. Jev reaches the vault through the MCP server.
With the connection up, the vault becomes the visual memory. Put media and data in it — images, logos, videos, PDFs, CSV — and have the agent store the media metadata and structured information in the graph. After that, ask Jev to retrieve a specific asset, like a logo file or a clip, or to answer a correlation question across the stored metadata, and it searches the graph and returns the referenced file. The directory build for this is MikeLembo’s jev-graph-search, which adds a Jev-powered search to Obsidian.
Albiona Hoti
@albicodes
I built a visual reference finder with Jev One single prompt → 100 images from Cosmos, NASA, and The Met my new rabbit hole for creative work 🌻

GitHubSearch
Jev-powered Obsidian search
Search an Obsidian vault with Jev deciding relevance.
MikeLembo
10What the image stage costs to run
The price people quote is the price of the decision stage, not the image workflow. Jev is $0.042 per million input tokens with no output-token charge; everything in front of it is a separate bill. For a sequential pipeline:
- Total latency = preparation + visual extraction + Jev + validation + network and queue time.
- Total cost = visual extraction + Jev + infrastructure + retries and review.
Parallel processing can raise throughput without lowering every item’s latency, and local inference spends hardware and time even when there is no per-call invoice. Report a cached run and a first-time extraction separately, because they cost different amounts and only one of them is honest about steady state.
When a vision model alone is enough
The result that argues against this whole page: an image-capable model like Gemini can answer the bounded question itself, so there is no guarantee that OCR plus Jev beats it on accuracy — only on cost and latency at volume. The proposal in the source article is a first experiment on consented or synthetic support screenshots with known routing labels, comparing four setups.
| Baseline | What it tests |
|---|---|
| OCR + rules | Whether semantic classification is needed at all |
| OCR + Jev | Whether Jev improves routing from extracted text |
| VLM alone | Whether one image-capable model is sufficient |
| VLM + Jev | Whether a separate decision stage improves the result |
Include blurry images, unfamiliar layouts, missing text, irrelevant attachments and cases with no valid automatic route, and split by document or template family so near-duplicates do not leak between development and evaluation. Report accuracy, false automatic actions, the manual-review rate, p50 and p95 latency, and cost per correctly completed task. Freeze the prompts and thresholds before the final test.
The builds that already do this
Each of these turns an image, a screen or a vault into the text Jev decides over, and each one publishes what it measured.
keno
@kenonews
JEV can’t see images. So how do you give it eyes? What’s possible? What can run locally? How far can you get for free? The answers (and the catches) are in the full breakdown below:

XDocuments and OCR
How to Give JEV Eyes
Fayaz Ahmed
@fayazara
Made myself a little image classifier with OCR + Jev It was able to categorise ~900 images in 40 seconds Pretty cool
Trinay Hari
@hari_trinay
Built a construction plan-set classifier with Jev. Proq turns civil and building plan sets into bills of materials using an LLM pipeline we built on GPT-4.1. Jev classified an entire 26-sheet plan set in 2.9 seconds for $0.0052. It matched GPT-4.1 and GPT-6 Astra on 100% of sheet-level classifications while running 17–21x cheaper and 5x faster than our production pipeline.
Milind S
@milindlabs
Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all Jev gets. It returns a probability across those elements and tells me the best one to click. Then it clicks, re-runs detection, and decides again. In a loop until the goal is done. ~90ms per decision. Faster than any LLM computer use I've tried. Blazing fast computer use, without any latency @typesafeai is building something really interesting
XAgents and browsers
Computer use without screenshots
Richard Meng
@richard_meng_01
Nitpicky, an AI generated photo detector powered by jev AI generated photos can be told from nits. That's why we build something to zoom into every detail: faces, fingers, characters, numbers, poses, where common senses fall apart, judged by jev
XDocuments and OCR
Nitpicky, zooming into AI images

Live siteOpen source
Laya Vision
Typed decisions about an image plus text, in one forward pass.
r33drichards
201MWhere these numbers come from
Every figure on this page is its author’s or the provider’s own: keno’s walk-through for the routes and their limits, Kev Builds Apps for the Obsidian setup, and the quickstart, the reranking guide and the tracing setup for the decision stage they feed.
Guideyoutube.com/@KevBuildsApps
Jev + Claude + Obsidian Just Changed AI Forever (Full Setup)
Kev Builds Apps wires Jev into an Obsidian vault over MCP: store the media and its metadata in the graph, then ask Jev to retrieve the logo, video or file a task needs instead of guessing.

Guidedocs.typesafe.ai
Quickstart: state, questions, typed answers
The three question types, Choice, Score and Noul, in one support-ticket example.
Guideyoutube.com/@engineerprompt
Steerable Reranking: How JEV Solves RAG
Prompt Engineering puts Jev in the reranking step of a RAG pipeline: why vector search and cosine similarity fail, how a steerable reranker works, and how it compares with LLM rerankers and cross-encoders. A Colab notebook comes with it.
Akshay 🚀
@akshay_pachaar

Guidex.com
Build a Jev Judge
Akshay Pachaar’s worked example: evaluate a refund-support agent with Jev instead of a generative judge, and record the verdicts in Opik. Separates the judge from the evaluation system, and is explicit that a 0.98 is a probability about one proposition, not a percentage of the answer.
Guidearize.com
Trace every judgment with Phoenix
Arize’s instrumentation for Jev: one line of code to trace each decision, with the integration docs.
Common questions
- Can Jev read images?
- No. Jev accepts text and structured state, not pictures. You turn the image into text first — with OCR, a vision-language model or a detector — and pass those observations to Jev as the state it decides over.
- Does Jev support images, audio or video?
- No, and it says so itself: the documentation lists Jev's input as text and structured state and recommends preprocessing anything non-text. There is no native image input and no way to upload a vision adapter into the hosted model.
- How do I give Jev an image?
- Add a stage in front of it. OCR reads the writing, a vision-language model records objects and layout, a detector returns boxes and labels, or an image encoder retrieves similar examples. Whatever that stage produces as text is what Jev receives.
- What is the cheapest way to read text from an image?
- Local OCR, because there is no per-call bill: Tesseract, EasyOCR and PaddleOCR all publish under Apache 2.0 and run on your own machine. A hosted free tier is capped — 1,000 Cloud Vision units a month, 25,000 OCR.space requests — and none of them is unlimited.
- Can a vision model replace Jev?
- Sometimes. If one image-capable model can answer your bounded question directly, Jev may add nothing, and the honest way to find out is to test that baseline against OCR plus Jev before you build anything. Jev earns its place when the visual stage has to feed many narrow decisions, or when the decision runs at a volume that makes the larger model too slow or too dear.
- Why is a confident answer still wrong?
- Because the numbers are not interchangeable. An OCR score, a detector score, an embedding similarity and Jev's confidence measure different things, and confidence is a statistic derived from the answer distribution, not an independent check of the image. If the first stage misreads the picture, Jev can route the mistake with equal certainty.
Made with Jev is independent and not affiliated with TypeSafe AI. Every figure on this page is the one its author published, linked to where it can be checked.