Blog Article

What Is Typesafe AI's Jev? A First Look at a Model That Decides Instead of Writes

A hands-on first look at Typesafe AI's Jev model. What the three output types mean, why the hot dog test returned 61% true, and what happened when we built a YouTube

I got the access email from Typesafe AI a day or two before this stream. I was at a client office all day, came home, watched some football, and called it a night. So on Friday, with client work done, I figured there was no better way to test a tool than to just explore it together live.

This is not a review. It is a first look. I had never touched the playground before I opened it on stream, and I want to keep that honesty in this post. What follows is what we learned, what surprised me, and what actually happened when we tried to build with it.

Wait, this is nothing like an LLM

The first thing you notice in the Typesafe playground is that you cannot just chat with it. Their model is called Jev, and it does not write anything. You give it some state, you ask it a typed question, and it hands back a typed answer with a confidence number.

The onboarding spells this out right away. Questions always need to be typed. That is how the model knows how to shape its answer. There are three question types, and if you learn these three you basically understand the model:

  • Null: how true is this statement? You get back a probability.
  • Choice: pick one from a set of categories you define. Each option gets a probability.
  • Score: place this on a rubric you define, like ignore, low, medium, high, critical.

That is it. No prose, no output tokens, no formatting to fight with. Riley Brown put it well in the explainer we watched on stream: because Jev cannot write at all, structure is not part of the request. It is the only thing it can return.

The hot dog test

The first walkthrough lesson asks a simple question: is a hot dog a sandwich?

I set it as a Null question and ran it. Jev came back with 61% true.

For the record, a hot dog is a hot dog. Technically you could define it as a sandwich based on how it is built, but you do not see restaurants selling a hot dog sandwich. Chat did not fully agree with me. Someone explained the 61% in Portuguese: a sandwich uses two pieces of bread, top and bottom, and a hot dog bun is one hinged piece. That kicked off a genuine debate about linguica sandwiches in Portugal and Brazil.

The real lesson is in the second half of that walkthrough. The playground told me the model was rather unsure and suggested clarifying what a sandwich is. So we added criteria to the question. A sandwich is a food where a filling such as meat, cheese, vegetables, or spread is placed between structural starch. It is not a sandwich if the food has no bread enclosing a filling, uses only a single slice, or uses a non-bread wrapper like a tortilla, wafer, or cookie.

With the criteria in place, the answer changed. That is the whole product in one demo. Vague questions get vague answers. The criteria are the product. This came back to bite us later in the stream, so keep it in mind.

The monkey selfie

The second lesson is based on the real copyright dispute over Naruto, the crested macaque who took a selfie. It shows how the state panel works. You define a subject, in this case human creative work, add baseline fact fields, and then ask typed questions against that state.

It also taught me that macaco is monkey in Portuguese, which I knew, and then forgot, and then remembered on air. Son of a. I knew that word.

Speed, cost, and a small context window

We watched Riley Brown's breakdown on stream because he had already spent time with the model. A few numbers from that video that matter for anyone thinking about building:

  • Typesafe's own benchmarks put a decision at roughly 0.4 seconds and about $0.00004 each.
  • A traditional LLM doing the same job runs closer to 3 cents and 10 seconds.
  • The context window is 64,000 input tokens. That is about 6% of a frontier model's window.

Riley ran 500 emails through it in roughly 12 to 13 seconds. A Null question for "mentions a brand deal opportunity" surfaced sponsorship pitches at 90% confidence. A Choice question categorized every email. A Score question rated importance on a custom scale that went from ignore all the way up to a level that needs a response within 30 minutes. A scam check flagged 55 of the 500.

The small context window is a design choice, not a bug. It is enough for a state panel plus your questions. It is not enough to dump your whole business into. Plan for tight, structured inputs.

One more thing I got wrong. I assumed Jev was a play on Ask Jeeves. Mr. Baker in chat pointed out it is Jevons paradox: when a resource gets more efficient, total consumption goes up, not down. I had never heard of it, so I read the Wikipedia definition on air. It fits the pitch exactly. Make decisions cheap enough and people will make a lot more of them.

My honest first take

I see the power already from the perspective of building applications or building wrappers around it. For retail users to come in and use it directly, I think you need some AI assistance. The playground is clean, but the real value shows up when a coding agent wires Jev into something you actually use.

So that is what we did.

Build 1: a YouTube feed filter

I did not hand-write a single line of Jev integration. Every build on stream used the same prompt shape: go read typesafe.ai and the docs, here is my API key, build this using Jev.

The first build went to Codex. The ask was a Chrome extension that filters course-funnel videos, AI hype men, and AI-generated slop out of the YouTube feed in real time.

Codex came back with an extension it named Signal. Separate filters for course funnels, exaggerated hype, and AI-made content. Adjustable thresholds, trusted channels, caching, and daily API limits. It also warned me up front: Jev currently accepts text only. Signal can classify hype from titles and metadata, but it cannot reliably detect undisclosed AI music or video.

That is a real limitation. Thumbnails carry a lot of the clickbait signal, and Jev cannot see them. Any visual use case needs a step that turns the image into text first.

First test was mixed. One video got tagged as likely exaggerated hype, which was right. But the Clearmud homepage showed zero filtered, and a channel that is a course funnel all day passed straight through. Part of that turned out to be my fault. I reloaded the extension folder and the API key dropped, so a chunk of the feed was never routed through Jev at all. The other part was the hot dog lesson again. Our definition of "course funnel" was not rich enough.

Then I decided to one-up it. Hide all Shorts. Hide the playables. Hide sponsored tiles. Show me raw, organic, long-form videos and nothing else.

That part worked instantly. Long form only. No more Shorts. No more games. This is OG YouTube right here. I have not been this happy scrolling YouTube in years.

Codex did not one-shot it by any means. It works on the surface, and now we need to go in and improve the judgment parts with deeper option trees and better keyword lists. But the mechanical parts were perfect on the first try.

Build 2: a model router

While Codex worked, I gave Claude a different job: a locally hosted model router that uses Jev to decide which model should answer each message. The docs have an intent routing pattern that is basically this exact thing, so Claude built on it.

The design is one Jev call per request asking five typed questions about the latest message. Task type. Difficulty on a four-level rubric. Whether a long answer is needed. Whether the stakes are high. Whether the result would be harmful. Plain code then maps those answers to a tier. Claude's summary noted Jev costs about 238 times less than a frontier model, which is why you can afford to run this on every single request.

I asked it to route to Codex, Anthropic, or Grok using my existing logins rather than raw API keys, so the only real secret on the machine was the Typesafe key.

"How's it going" got scored as trivial chat and routed to Grok. "Which frontier lab contradicts itself the most?" got scored as demanding and went to a frontier model, which opened with a caveat that it was made by Anthropic and I should discount its take accordingly. Fair enough.

This is the pattern I keep coming back to. Jev decides. Cheap deterministic code acts. You keep the expensive text-generating model out of the loop until you actually need text.

Build 3: an inbox dashboard (bonus)

As a bonus I prompted OpenClaw to research Jev and build a dashboard tab that scans, filters, and categorizes my inbox every morning, with a screenshot of Riley's dashboard as a visual reference. It is wired to my main mailbox with a few aliases. I did not demo it on stream because those are real emails and I am not showing them on a live feed.

What I would tell you to do

  • Get access. Join the waitlist at typesafe.ai. Chat also mentioned Jev is available through the Vercel AI Gateway and OpenRouter, so you may not need to wait.
  • Learn the three types first. Choice for categories, Score for a rubric, Null for how true something is. Decide which one your problem needs before you write anything.
  • Write the criteria, not just the question. If your classifier lets things through, the fix is a richer definition, not a bigger model.
  • Convert visuals to text. Jev only reads text. Describe thumbnails, frames, or audio in words before asking it to judge.
  • Let Jev decide, let code act. The router pattern is the cleanest example.
  • Pick things that repeat. Email triage every morning. Feed filtering on every scroll. Routing on every request. The economics only matter if the decision runs constantly.
  • Expect a second prompt. Both builds needed a follow-up. Test against your real feed or inbox and go back to the docs.
  • Keep keys off screen. I paused recording every time a key went in. Use OAuth logins where you can.

What happens next

The Signal extension is going to get released for free. Paste in your own Typesafe API key and go. Maybe a sensible default for people who do not have one yet. Why not? We built it in public, the community should get to use it.

We have barely scratched the surface here. There is a lot more in the docs than we touched, and I want to build something fun with Jev that nobody else is doing. Put a one in chat, or in the comments, if you want to see it launched.

Until then, touch some grass. Disconnect as best you can. Same time Monday.

Watch the full stream: https://www.youtube.com/watch?v=0U8d0aCU8Fs

Watch the full walkthrough on YouTube.

Watch on YouTube