Back to today's topics

Verified · Sep 20, 2026

Independently verified

TypeSafe's Jev never outputs a word — typed decisions with confidence scores, a listed $42 per billion input tokens, and a launch-week API that demand briefly broke

2 sources

The company's framing, from its September 15 post: Jev is 'a new class of frontier models built to make fast, structured decisions that software can use directly', 'available today in early access', and 'While Jev gives up string generation, it's optimized for structured outputs and can't hallucinate.' The pitch line: 'Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.' TechCrunch's translation: 'a new transformer-based model, Jev, that is not a large language model (LLM). It doesn't output text, but instead produces probabilities, or what the company calls "calibrated decisions."' Who built it: Diogo Almeida, whom TechCrunch credits as an OpenAI researcher who 'helped build the chatbot and then invent reinforcement learning from human feedback (RLHF)', and who told the outlet 'We have lightning in a bottle, and yet it is not useful' about his ChatGPT-era work. The money, per the company's own table: 'Input tokens: $0.042 / MTok ($42 per billion tokens).' and 'Output tokens: FREE (too cheap to meter)' — with the company's own caveat 'We can't prove it isn't subsidized' about its own pricing. The speed claims are the company's: '70ms-500ms' end-to-end, '40x-200x faster' on System One-shaped queries, the homepage's '193.6x Faster, 444.6x Cheaper.' footnoted '*based on workflows for System One tasks' — and the blog's own concession that 'we expect that these are on the higher end of real world gains' — plus '238x Lower input price than Claude Fable 5.1'. The independent beats are TechCrunch's: 'the company briefly lost the ability to serve users from its API because demand was so high'; Vercel's Pranit Sharma got results 'five to 18 times more quickly and with greater accuracy' after replacing OpenAI's ChatGPT Luna 5.6 with Jev on a safety-command classifier; Bryo AI CTO Nikhil Mudholkar found Gemini slightly more accurate but '10 to 20 times more expensive' for classifying business emails, and called Jev's confidence scores the real difference: 'it is the only one that hands back a real probability which makes it ideal for automating workflows!!' The usage frame, from Armin Ronacher (CTO of Earendil, builders of the open source model harness Pi): 'it delegates the hallucination problem a little bit to the user' — at 50% 'maybe this is a coin toss, and I disregard it. But if it's 95%, sure, then I can do something with it.' The secrecy: 'Almeida is tight-lipped about the model's architecture, which outside observers suspect is built on top of an open-weight LLM', Almeida says it is 'trained exclusively on synthetic data using a technique he calls "reinforcement learning from calibrated decisions."' — the company's own rendering of that name is 'Reinforcement Learning for Calibrated Decisions (RLCD)'.

Why now

The launch landed Tuesday, September 15, and TechCrunch's developer read followed on Friday, September 18 — the weekend window is when tool-explainer creators can still be first on a story whose aggregation wave hasn't started. The demand signal is already public: the company's own rollout line is 'bringing developers off the waitlist as quickly as we can', and TechCrunch reports the API briefly buckled under demand, so audience curiosity about access timing is live right now. The topic also rides the automation conversation this story feeds: if a model that skips language really prices at $42 per billion input tokens with free outputs — by the company's own numbers, with its own subsidy caveat attached — then the question of which decisions in your workflow deserve a typed answer instead of a chat completion becomes a fresh explainer question no one has saturated yet. And the differentiation play is built in: every number here is a vendor figure or a single-developer test, so the creator who carries the attributions correctly is instantly the most trustworthy version of this story on the platform.

Why it is worth publishing

The most demo-able topic on today's card: a model with no chat window explains itself on screen, and the 50%/95% threshold mechanic from Ronacher's quote is a buildable mock in an afternoon. The precision play is the attribution stack — vendor benchmarks footnoted to the company's own workflows, the subsidy caveat in TypeSafe's own words, the developer numbers kept per-developer, the RLCD naming kept per-outlet — because aggregator takes will flatten all of it into a cheaper-AI-that-can't-hallucinate headline. The audience is narrower than the Gemini card's (developers and AI-tool reviewers rather than everyone), which is exactly why it pairs with the top pick instead of competing with it.

Evidence basis

Two sources, both opened this run: the TypeSafe AI blog post (byline 'Diogo Almeida, founder, TypeSafe'; post dated Sep 15, 2026; its 'available today' anchors the launch to September 15, 2026) and the typesafe.ai homepage for the benchmark block, plus TechCrunch (Tim Fernholz, 11:49 AM PDT, September 18, 2026; read in full, raw HTML fetched). Weekday-date pairs calendar-verified: Friday = September 18, 2026; today = Sunday, September 20, 2026. Numbers are numerous and all welded to their holders: $0.042/MTok ($42 per billion) input and free output (company); 70ms-500ms and 40x-200x (company); 193.6x faster and 444.6x cheaper, footnoted to the company's own workflows (company); 238x lower input price than Claude Fable 5.1 (company); five to 18 times more quickly with greater accuracy (Vercel's test, per TechCrunch); 10 to 20 times more expensive with slightly higher Gemini accuracy (Bryo AI's test, per TechCrunch); 'two orders of magnitude' (company); 'briefly' on the API outage (TechCrunch). No independent benchmark of these claims exists in any opened source.

This new AI model never writes a single word — it only returns typed decisions with a confidence score, and that's why its maker says it can't hallucinate.

Angle

Make it a 'model with no chat window' explainer. Beat one, the shape: typed decisions with calibrated confidence scores instead of text — the company's 'frontier-intelligence function call' line on screen. Beat two, why no hallucination is a design claim: outputs are defined in advance, so there is no free-form string to be wrong — keep the design mechanism attached, never a bare 'it can't hallucinate'. Beat three, the money: $42 per billion input tokens, free outputs — vendor's own numbers, with the subsidy caveat in TypeSafe's own words. Beat four, what it's for: classify, route, score, guardrail — with the developers' own test numbers, including the Gemini-accuracy counterpoint.

Format

Long-form explainer

Demo idea

A split-screen classifier mock: left side, a chat model answering a routing question in a paragraph that then has to be parsed; right side, a Jev-style typed answer — {decision, probability} — with a threshold slider below it set to Ronacher's own example points (50% 'coin toss, disregard' vs 95% 'do something with it'). Footer carries the money line — '$42 / billion input tokens · outputs free · by the company's own pricing post, which itself says it can't prove the price isn't subsidized' — and the benchmark badge reads '193.6x / 444.6x — TypeSafe's own workflows benchmark'.

Platform notes

Every vendor number ships with its holder: '193.6x faster, 444.6x cheaper' only with '*based on workflows for System One tasks' and the blog's own 'higher end of real world gains' concession; '238x lower input price than Claude Fable 5.1' only as TypeSafe's claim; the $42 price always with the company's 'We can't prove it isn't subsidized'. Developer numbers ship per-developer with TechCrunch as the relay: Vercel's five-to-18-times, Bryo's 10-to-20-times — and Bryo's finding that Gemini was slightly more accurate travels with them. 'Can't hallucinate' ships as the design claim it is (predefined outputs, typed answers), never as a general accuracy guarantee. Architecture stays contested: the company says 'a new model architecture', TechCrunch reports observers suspect an open-weight LLM base — both attributed, never settled. The launch date is Tuesday, September 15 — never 'today' in the card's own voice.

Usable claims

  • TypeSafe AI — co-founded by Diogo Almeida, whom TechCrunch credits as an OpenAI researcher who 'helped build the chatbot and then invent reinforcement learning from human feedback (RLHF)' — released Jev in early access on Tuesday, September 15, 2026: the company's post (dated Sep 15, 2026) says 'available today in early access', and TechCrunch's September 18 story dates the release only as 'This week, the company released'. TechCrunch: 'a new transformer-based model, Jev, that is not a large language model (LLM). It doesn't output text, but instead produces probabilities, or what the company calls "calibrated decisions."' The company's own words: 'While Jev gives up string generation, it's optimized for structured outputs and can't hallucinate.' and 'Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.' Pricing is the company's own: 'Input tokens: $0.042 / MTok ($42 per billion tokens).' and 'Output tokens: FREE (too cheap to meter).' — with the company's caveat 'We can't prove it isn't subsidized' about its own pricing. The homepage's '193.6x Faster, 444.6x Cheaper.' block is footnoted '*based on workflows for System One tasks', the blog says of those figures 'we expect that these are on the higher end of real world gains', and the company claims '238x Lower input price than Claude Fable 5.1'; its stated end-to-end response time is '70ms-500ms', which 'can range from 40x-200x faster' on System One-shaped queries. TechCrunch reports demand was high enough that 'the company briefly lost the ability to serve users from its API'; a Vercel engineer's test per TechCrunch got results 'five to 18 times more quickly and with greater accuracy' after replacing OpenAI's ChatGPT Luna 5.6 with Jev, while Bryo AI's CTO found Gemini slightly more accurate but '10 to 20 times more expensive' for classifying business emails.

Evidence pipeline

Breakdown

A launch this quotable is a launch this easy to misquote. Layer one, the benchmark stack: the homepage's '193.6x Faster, 444.6x Cheaper.' carries the company's own footnote '*based on workflows for System One tasks', the blog says those figures are 'on the higher end of real world gains', and the same post admits 'We can't prove it isn't subsidized' about its own pricing — a vendor that pre-emptively hedges itself has handed creators the exact attribution language to reuse, and dropping it is the difference between reporting and repeating. Layer two, the developer tests are not benchmarks: Vercel's five-to-18-times and Bryo's 10-to-20-times are single-developer results relayed by TechCrunch, and Bryo's same test found Gemini slightly more accurate — the counterpoint is part of the fact, not a footnote to it. Layer three, the design claim: 'can't hallucinate' in the company's words and TechCrunch's rendering both hang on the same mechanism — outputs defined in advance, typed answers, no free-form string — so the claim travels with its mechanism or not at all; TechCrunch's own line ('because users define the outputs in advance, it cannot hallucinate') is the template. Layer four, the contested architecture: the company pages claim 'a new model architecture' while TechCrunch reports Almeida is 'tight-lipped' and outside observers suspect a built-on-top-of-an-open-weight-LLM base — two attributed positions, no synthesis. Layer five, the naming drift as a precision badge: TypeSafe writes 'Reinforcement Learning for Calibrated Decisions (RLCD)', TechCrunch writes 'reinforcement learning from calibrated decisions' — one preposition apart, and the card that keeps them straight is the one that gets trusted when the numbers get contested. Editor's rule: vendor numbers with footnotes, developer numbers with their developer, design claims with their mechanism, and no number in this story travels without a holder.

Risks

  • Before quoting any number, open the TypeSafe blog post and homepage and read the sentence and footnote that carry it; quote vendor figures only with 'by the company's own benchmark' or 'per its own pricing post' attached, and attach the subsidy caveat whenever the price appears. Attribute every developer number to TechCrunch's report of that developer's single test, and pair the 'can't hallucinate' claim with its predefined-outputs mechanism plus Bryo's Gemini-accuracy counterpoint.

Demo ideas

  • Split-screen classifier mock: chat-model paragraph vs typed {decision, probability} answer, with Ronacher's 50%/95% threshold example as the interactive beat and the vendor-benchmark footer carrying the footnote
  • Naming-drift card as a precision flex: TypeSafe's 'Reinforcement Learning for Calibrated Decisions (RLCD)' beside TechCrunch's 'reinforcement learning from calibrated decisions' — teaching the audience to spot which outlet a claim came from