AI Systems
Jev After One Week: Honest Answers on Speed, Cost and Accuracy
A week ago we wrote up our first three days with Jev. Since then we've put it on more work, including a live group call where it didn't act once. Here's what we know now.
One thing to keep in mind: when we give you a percentage, it's how often Jev agreed with Claude or with our app. Nobody has hand-checked Jev's answers yet, so there's no true accuracy number in this post.
How Jev works
The simplest way to explain it: Jev is an AI that never writes a word. It reads text the way a chatbot does, and then it gives you back a decision and how sure it is.
TypeSafe calls it a "System One" model, after the fast, gut-instinct kind of thinking. Claude and ChatGPT are the slow, step-by-step kind. Because Jev skips the writing, it's incredibly fast and cheap.
Here's how it works. We send it some text, say an email, a web page, or a sentence someone said on a call. We ask it one or more fixed questions about that text. It sends back an answer to each, with a confidence score. Then our own code decides: if it's very sure, act on it. If it's not, hand it to a person or to Claude.
There are three kinds of questions you can ask it. A true-or-false question, which TypeSafe calls a "Noul", like "does this page need a compliance review?", and you get back the probability it's true. A multiple choice, like "which of these five content types is this page?" And a score, as in "rate this 1 to 5." We've used the first two in real work. The score type we've only tried in setup.
In our opinion, the confidence score is the part that matters most. It's what lets you decide when to trust Jev and when to escalate.
A few practical things. You can ask several questions in one request. We ask up to nine at once, which is cheaper and faster than nine calls. You only pay for the text you send in, and the answers are free. List price is about four cents per million tokens.
TypeSafe's own test says Jev matched Claude Sonnet's score at 0.4 seconds and a fraction of a cent, versus 78 seconds and about 12 cents. But that's their test, graded against other AIs rather than people, and we haven't verified it.
What we've used it on
We've put it on seven different kinds of work in about a week. They're all the same shape: lots of small decisions where the answer comes from a fixed list.
The biggest one is a live meeting app we're building. During a real call, Jev decides whether the app should pop up a suggestion, and what kind. We ran it on 130 decisions on one call and 116 on another.
In the same app, we use it for voice commands. When you talk, it decides whether you're giving the app a command, like "type this" or "open Safari", or just talking to the people on the call. We ran 600 test sentences through that.
We also use it to check the app's own replies, making sure none of its internal instructions leak into what the user sees.
On the content side, we had it read 275 pages from our marketing knowledge base and answer four questions about each one: what type of content it is, what customer stage it's for, who the audience is, and whether it needs a compliance review. It did all 275 in about 30 seconds. We ran a stricter compliance version on 155 pages.
We used it to sort the 130 projects in the public Jev gallery so we only had to read the useful ones.
And day to day, it's become our first pass on any big pile of reading. Jev sorts it, and then Claude, or one of us, only reads the ones that matter.
What other builders are doing
There's a public gallery, madewithjev.com, with 300-plus projects on it. We've gone through it, but these are the builders' numbers, not ours.
People are using it to sort big piles. One person broke down 724 competitor ads in about 40 seconds for around 9 cents. Another sorted 1,018 research papers for 8 cents, and another sorted 500 emails for three and a half cents.
Others put it inside AI agents, where a bigger model plans and Jev picks each next click, at around a tenth to a third of a second a step. Some use it as a gatekeeper that checks every action an agent takes, or as a router that picks the cheapest model for each task. Someone even has it playing Doom at about ten moves a second.
What stands out to us is that almost every number people share is about speed or cost. Very few are measuring whether it got the answer right. That's the same gap we have.
Setup
The tool itself is simple. Most of the pain in week one was on our side, plus one thing that isn't obvious until you hit it.
A couple of small ones first. Every Jev key starts with "apikey_", and that's actually part of the key. We trimmed it off while tidying it up, and every request got rejected until we figured it out. And TypeSafe's plugin for Claude Code didn't show up until we restarted.
The bigger lesson was design. We first plugged Jev in like it was a chat AI, and it isn't. It doesn't write. It answers fixed questions. A same-day review found six mistakes in that first hookup. For example, we'd written "yes/no" where Jev expects "true/false", so every single call would have been rejected. We redesigned it three times in one day, and landed on asking nine questions in one call after reading all 18 of TypeSafe's example guides.
If you want one tip, it's this: read their cookbooks before you design, not after.
Accuracy
We can't give you an accuracy number yet. To get one, you need examples where a person has decided the right answer, and we haven't built that yet. What we can tell you is how often it agreed with Claude, and that's still telling.
On those 275 knowledge base pages, Claude had tagged them earlier, and we let Jev tag them fresh. On content type, they agreed 78% of the time overall, but 96% of the time when Jev said it was confident. That's the good story: its confidence actually meant something.
On customer stage, they agreed 76%, and 83% when Jev was confident. On audience, only about 62%. But when we dug in, that was our fault, not Jev's. Our tag followed the kind of page it was, and Jev was reading what the page actually said.
Compliance looked like 87.5%, but that number is misleading. Jev said "yes, needs review" on 98% of pages, so it wasn't really telling them apart. When we made the question stricter, it started ranking pages well, a score of about 0.93 out of 1.
So our honest read is that it's promising, and in places where Jev disagreed, it was sometimes right and we were wrong. But it's not fully proven yet.
Speed
Speed is its clearest strength. A typical answer comes back in somewhere between a sixth and a third of a second, depending on how much text we send. It read all 275 knowledge base pages in about 30 seconds.
The comparison that sold us was in our meeting app. Jev made its call in about 0.17 seconds. The main AI in the app took about 7.4 seconds for its step. That gap is the difference between something that can keep up with a live conversation and something that can't.
Cost
At our volume, it's basically free. We've spent well under a dollar across roughly 890 requests.
Here are the real numbers. Those 275 pages, about 2.6 million tokens of text, cost 11 cents. The stricter compliance run on 155 pages was 4 cents. 130 decisions on a live call cost less than one cent. The 600 voice test sentences were about 2 cents in total.
Two caveats. TypeSafe has said early pricing may be subsidized, so we'd expect it to go up at some point. And the API doesn't tell you what each answer cost, so we calculate it ourselves from token counts.
Reliability
It has never errored on us. Across roughly 1,700 answers, there were zero errors from Jev, and it served the same model version every time, so results didn't shift under us.
The only hiccups were ours. If you send Jev a badly formatted question, it rejects that one request, and our tool used to stop the whole batch when that happened. We fixed our side so it skips that item and keeps going.
The one real limit is size. Very long documents have to be split up. 39 of our 275 pages were too long and got cut off.
Where it failed
It's fast and it's cheap, but it has real blind spots, and this is probably the most useful part of this post.
The biggest one: on September 21st we let Jev actually make decisions on a live group call. On all 116 decisions, it said "not sure" and passed. It never acted once. We think our confidence bar was set too high for a busy group conversation, but the lesson is that you have to tune those thresholds on real examples, not guesses.
It takes things very literally. When someone said "open GitHub", it answered that no app was being asked for. And when someone said "Type: see you Thursday", meaning type out the words "see you Thursday", it judged the words after "type" and decided it was just normal talk.
Its confidence score isn't always reliable either. Sometimes the answers it was most sure about matched Claude less often. And if you change the wording of a question, you have to set up the confidence levels all over again.
One odd pattern: on the audience question, "both" was an option, and it picked it zero times out of 275.
And to be fair, TypeSafe is upfront about its limits. It's weak at counting, math, and comparing numbers or dates. Irrelevant text in the input makes it worse. Text inside a document can steer it. And it can't write.
Would we trust it on its own?
Not yet. Right now we run it in what we call shadow mode. Jev makes its call, but our real system still decides, and we log both side by side. We only read the cases where they disagree.
We sort its answers by confidence. When it's very sure, we act. In the middle, a person checks. When it's low, we ignore it. On our first live call, only about 8% of its decisions, 11 out of 130, were confident enough to act on.
What would change our minds is a proper test set with human-checked answers. That's the next thing we're building.
Where it fits
The way we'd put it: it's a sorter, not a writer.
It's great at tagging and sorting big piles. It's great at triage, meaning deciding whether something is worth a person's time. It's great at routing things to the right bucket or tool. And it's great as a second opinion on another system's rules.
We wouldn't use it for anything that needs writing or summarizing, for math, counting or dates, for images or audio, or for one-off, high-stakes decisions. That's what Claude is for.
The surprise value
The thing we didn't expect was how useful it is as a checker.
We ran 600 test sentences through both Jev and our app's own rules, and only read where they disagreed. That turned up real bugs in our app: normal conversation that would have been typed out as text, Safari not being recognized as an app, phrases like "that's it" not being caught as the end of a command, and spoken edits to a draft being missed. In a later check, it confirmed none of the app's last five real replies leaked internal instructions.
It also showed us our own labels were sloppy. Our audience tag was following the type of page instead of what the page said, and we'd overused an "all stages" tag. Jev moved 58 of those 229 pages to a specific stage, which tells us we need to fix our tags.
What we'd ask TypeSafe for
We'd ask for four things. First, show the cost on every answer, because right now we calculate it ourselves. Second, a confidence score on every type of answer, since some come back without one. Third, confidence levels that hold up when you reword a question, so you don't have to re-tune everything. And fourth, better handling of spoken commands, where it should read the intent and not just the words.
To be clear about what we haven't done: we've planned to use it for sorting emails, SEO pages and ad copy, but we haven't run those yet.
Our verdict
If we sum it up: it's fast, it's nearly free, and it has never errored on us. We use it as a sorter and a second opinion, with Claude or a person reading wherever it disagrees. So it's genuinely useful today, but we wouldn't let it run on its own just yet.
If you asked us to rate it, we'd say 8 out of 10, because of what it has already let us build.
And if you want one thought to take away: everyone in this space is showing off speed and cost. The company that proves accuracy on real, human-checked data is the one that wins this category.
If you want to measure it on your own data, the free skill we built for that is linked in our first Jev post.
Want this run for your business?
Most partners start with The Growth Hour.
