# System One Jev by TypeSafe AI

Research date: **21 September 2026 KST**. The core account uses first-party TypeSafe pages, documentation, evaluation pages, API schemas, repositories, and statements by a TypeSafe founder. A separately marked section covers four recent independent GitHub evaluations, and one training-data statement is attributed to a TechCrunch interview. I used research indexes to check for a technical paper and to distinguish unrelated uses of the acronym RLCD from TypeSafe's work.

## Bottom line

Jev is a hosted, text-in and typed-decision-out model. It does not generate prose. A request supplies one shared state and one or more `Choice`, `Score`, or `Noul` questions. Jev returns a selected closed-set value or numeric score, together with probabilities. TypeSafe says it evaluates the questions in parallel and bills only input tokens. The current documented release is `jev-1.13.0`, priced at **$0.042 per million input tokens**, with output tokens free. [Launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev) · [Models](https://docs.typesafe.ai/models) · [Introduction](https://docs.typesafe.ai/introduction)

Its strongest public evidence is for speed and price, not best-in-class accuracy. In TypeSafe's own four-workflow evaluation, Jev averages **67.8% agreement** with a model-generated reference at **$0.0004 and 0.4 seconds per case**. OpenAI Sol reaches **74.1%** at $0.0836 and 23.3 seconds; Anthropic Opus 5 reaches **73.1%** at $0.1761 and 37.8 seconds. Jev is the clear cost and latency outlier in that producer-run test, but several models are more accurate. The reference is the average of GPT-6 Astra and Claude Fable 5.1 at high reasoning, not a human or ground-truth label. [Official workflow evaluations](https://evals.typesafe.ai/) · [Launch methodology and caveats](https://typesafe.ai/blog/introducing-system-one-models-and-jev)

Four very recent independent repositories make the picture more jagged. Jev strongly beats GLiNER2.5 on small AG News and Banking77 samples, roughly ties Cohere Rerank Pro on one equal-dataset ranking aggregate, performs well on a context-rich prompt-injection test, but loses badly to Haiku 4.5 on direct phishing classification. Its calibration is good on some security tests and poor on Emotion and direct phishing. These are reproducible author-run studies, not peer-reviewed or independently replicated results.

The model itself is **not reproducible from public information**. TypeSafe has not disclosed the parameter count, computational graph, layer layout, attention or recurrence mechanism, base checkpoint, tokenizer, pretraining corpus, training-example count, reward equation, RL algorithm, optimizer, learning-rate schedule, batch sizes, compute budget, calibration procedure, inference kernel, weights, or training code. Founder Diogo Almeida said publicly that the architecture was being kept "close to the chest" and that the company had discussed writing a paper. [Founder comment](https://news.ycombinator.com/item?id=49718824)

I found no Jev or Reinforcement Learning for Calibrated Decisions technical paper, model card, checkpoint, or training repository. TypeSafe's public GitHub organization contains SDKs, an LLM comparison adapter, agent skills, infrastructure projects, and old upstream forks, but no Jev implementation. [TypeSafe GitHub organization](https://github.com/typesafe-ai) · [System One adapter](https://github.com/typesafe-ai/system-one-adapter-python) · [Python SDK](https://github.com/typesafe-ai/typesafe-sdk-python) · [JavaScript SDK](https://github.com/typesafe-ai/typesafe-sdk-js)

The practical distinction is:

- You can reproduce an **API integration** and run your own behavioral evaluation against the pinned hosted release.
- You cannot reproduce **Jev's training or inference implementation**, or verify that another implementation is architecturally equivalent, from the material TypeSafe has published.

## How to read the evidence

This report separates three kinds of claims:

- **Published fact:** TypeSafe states it directly, exposes it in the API, or publishes the result in its evaluation pages.
- **Inference:** a conclusion supported by published behavior, but not an implementation detail TypeSafe disclosed.
- **Unknown:** information required to explain or reproduce the model that is absent from the public record checked on the research date.

TypeSafe launched Jev on 15 September 2026, only six days before this review. Documentation, aliases, limits, and evaluation pages may move quickly. Pin `jev-1.13.0` rather than `jev-latest` when testing. The API response reports the resolved version. [Models and aliases](https://docs.typesafe.ai/models)

## What Jev is

### Published facts

TypeSafe calls Jev its flagship model and the first member of a new "System One" class. The intended workload is fast judgment inside software: classification, routing, ranking, moderation, risk assessment, policy decisions, and other bounded choices. The input is natural-language text, either as a string or text-bearing JSON. Image, audio, video, and binary inputs are not accepted. [System One](https://docs.typesafe.ai/concepts/system-one) · [Models](https://docs.typesafe.ai/models)

The three output primitives are:

| Primitive | Input contract | Returned value | Important semantics |
|---|---|---|---|
| `Choice` | Instructions plus up to 255 named criteria/options | Winning option, a probability for every option, and confidence | The winning option is the maximum-probability option. It is a closed set, so the model cannot invent another label. [Choice](https://docs.typesafe.ai/primitives/choice) |
| `Score` | Instructions plus 2 to 10 ordered, descriptive levels | A probability for every level, an expected numeric score, legend, and confidence | Jev evaluates each level description independently. The numeric score is a probability-weighted expected value, not a directly predicted real quantity. [Score](https://docs.typesafe.ai/primitives/score) |
| `Noul` | A yes/no proposition | Probability that the proposition is yes | It returns no separate confidence. A value near 0.5 can mean ambiguity or a genuine midpoint, and it is not a degree or fuzzy score. [Noul](https://docs.typesafe.ai/primitives/noul) |

The client-supplied question key identifies the answer in the JSON response but is not itself shown to the model. The semantic content is in the instructions and criteria. Questions in a request are evaluated independently against the same state. TypeSafe recommends decomposing broad tasks into atomic judgments and combining them in ordinary code. [Primitives](https://docs.typesafe.ai/primitives) · [How to build with TypeSafe](https://docs.typesafe.ai/concepts/how-to-build-with-system-one)

The current limits for Jev 1.13 are **64,000 tokens across the request** and **32,000 tokens for the state plus the single longest question**. The documented service limits are 250,000 tokens per second and 1,200 requests per minute, but TypeSafe warns that these are changing dynamically. English is the primary training language and the language in which accuracy is best. Other languages, including CJK scripts, are accepted but documented as less accurate. [Models](https://docs.typesafe.ai/models)

### Minimal wire contract

This is the published request shape, reduced to one example of each primitive:

```json
{
  "model": "jev-1.13.0",
  "state": "The text or text-bearing JSON to evaluate",
  "questions": {
    "route": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "billing": "Payment or subscription issues",
        "technical": "Bugs or integration problems"
      }
    },
    "severity": {
      "type": "score",
      "instructions": "How severe is the issue?",
      "criteria": [
        "Minor inconvenience",
        "Important degradation",
        "Critical outage"
      ]
    },
    "urgent": {
      "type": "noul",
      "instructions": "The issue needs immediate attention"
    }
  }
}
```

The response contains the resolved model ID, an answer keyed by each question name, and input/output token usage. Choice and Score answers expose their full discrete distributions. The hosted endpoint is `POST https://api.typesafe.ai/v1/systemone`; `GET /v1/models` lists aliases available to an authenticated account. [Quick start](https://docs.typesafe.ai/introduction/quickstart) · [API reference](https://docs.typesafe.ai/api) · [OpenAPI schema](https://api.typesafe.ai/openapi.json)

The public API does not expose a temperature, random seed, decoding budget, or model-specific sampler setting. This matters because TypeSafe's own repeated-call studies observe output variation. [API reference](https://docs.typesafe.ai/api) · [Choice consistency study](https://docs.typesafe.ai/cookbooks/consistency_choice_cookbook) · [Noul consistency study](https://docs.typesafe.ai/cookbooks/consistency_noul_cookbook)

## Architecture

### What TypeSafe has disclosed

TypeSafe makes four architectural or computational claims:

1. Jev uses a "new model architecture" and a "new parallel sampler." [Launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev)
2. It gives up arbitrary string generation. The outputs are bounded, typed decisions. TypeSafe describes this as parallel rather than autoregressive output-token generation. [Launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev) · [System One](https://docs.typesafe.ai/concepts/system-one)
3. Multiple questions against one state run independently and in parallel. Adding questions is meant to increase latency much less than making one separate generative call per question. [Introduction](https://docs.typesafe.ai/introduction) · [Parallel-questions cookbook](https://docs.typesafe.ai/cookbooks/parallel_questions)
4. Almeida agreed with a Hacker News description of Jev as "basically a zero-shot classifier" and said strings and other sequential data structures are disallowed so all outputs can be computed in parallel. He also described it as a general model for System One tasks that requires no task-specific training by the customer. [Classifier exchange](https://news.ycombinator.com/item?id=49718490) · [Parallel-output comment](https://news.ycombinator.com/item?id=49719122) · [General-model comment](https://news.ycombinator.com/item?id=49719245)

TypeSafe's launch FAQ says Jev is "neither small nor an LLM." Its AI primer nevertheless draws RLCD as a third adaptation path from pretrained language models, beside RLHF and RLVR. The diagram is a description of the training paradigm, not a disclosure of Jev's starting checkpoint or final computational graph. These statements could be compatible if TypeSafe uses language-model pretraining and then changes the interface or architecture enough that it no longer calls the deployed system an LLM. The public material does not resolve that question. [Launch post and FAQ](https://typesafe.ai/blog/introducing-system-one-models-and-jev) · [AI primer](https://docs.typesafe.ai/introduction/machine-learning-primer)

### What can reasonably be inferred

**Inference:** Jev must map a shared representation of the state and question text to a finite probability distribution over supplied criteria or levels. That is enough to describe the observable interface. It is not enough to identify the implementation. An encoder with classification heads, a candidate scorer, an energy-based model, a masked or diffusion-style model, or a custom architecture could all satisfy the same contract.

**Inference:** Removing arbitrary strings makes schema validity easy to enforce and removes the latency cost of sequential text decoding. It does not prove that every internal operation is parallel, nor that the model has no autoregressive component during training or inside an unpublished subsystem.

**Inference:** The 255-option Choice limit is a service or representation constraint. It does not disclose the dimensionality of a learned output layer. Criteria are supplied dynamically as text, so one should not assume a fixed 255-class classifier.

### What remains unknown

No first-party source checked here gives any of the following:

- total or active parameter count;
- dense versus mixture-of-experts design;
- transformer, recurrent, diffusion, state-space, or other block type;
- layer count, hidden width, head count, activation, normalization, position encoding, or context-extension method;
- tokenizer and vocabulary;
- whether state encoding is cached across questions, and where candidate/question interactions occur;
- whether probabilities come from a softmax, independent logits, an energy function, ensembles, sampling frequencies, or post-hoc calibration;
- numeric precision, quantization, serving hardware, batch size, parallelism strategy, or inference kernels;
- the meaning of "parallel sampler" in an algorithmic sense.

Do not infer the architecture from TypeSafe's public `LLaDA` or `vllm` forks. On the research date, the default branch of each fork was zero commits ahead of its upstream project. They provide no evidence that Jev is a diffusion language model or that it runs on vLLM. [TypeSafe `LLaDA` fork](https://github.com/typesafe-ai/LLaDA) · [TypeSafe `vllm` fork](https://github.com/typesafe-ai/vllm)

## Training and RLCD

### What TypeSafe says it optimized

RLCD means **Reinforcement Learning for Calibrated Decisions**. TypeSafe contrasts it with RLHF, which optimizes human preference over generated responses, and RLVR, which rewards verifiable solutions. Its stated target is a typed decision with a probability whose frequency matches empirical correctness. Across a sufficiently large group, events assigned 0.2 should occur about 20% of the time and events assigned 0.8 should occur about 80% of the time. This is a group-frequency property, not a guarantee for an individual prediction. [AI primer](https://docs.typesafe.ai/introduction/machine-learning-primer)

The launch FAQ says TypeSafe primarily considers itself a data research lab, makes all of its training data itself, does not train on customer data, and treats the details as proprietary. The model documentation separately says the same Jev weights serve every account, with no customer-specific fine-tuning or LoRA. Customers adapt behavior at request time through state, instructions, criteria, decomposition, and code. [Launch post and FAQ](https://typesafe.ai/blog/introducing-system-one-models-and-jev) · [Models: customization and data handling](https://docs.typesafe.ai/models)

A [TechCrunch interview with Almeida](https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/) goes one step further: the article reports him saying the training data is exclusively synthetic and that roughly half of TypeSafe works on data research. This is attributed secondary reporting, not a dataset release or first-party technical specification. The same interview says TypeSafe is keeping the architecture private. It does not identify the generators, teacher models, prompts, labels, filtering, or scale. The article also relays outsider speculation about a possible base model; there is no TypeSafe confirmation, so this report does not repeat that speculation as fact.

TypeSafe's broader training thesis is that task definition and data matter more than increasingly elaborate algorithms. That philosophy helps explain the product, but it is not a training recipe. [The bitterest lesson](https://typesafe.ai/blog/bitterest-lesson)

### The missing recipe

The public material does not specify:

| Reproduction requirement | Public status |
|---|---|
| Pretraining starting point | **Unknown.** No base model, checkpoint, or from-scratch declaration tied specifically to Jev. |
| Training data | **Only a high-level claim.** TypeSafe says it creates the data itself; TechCrunch attributes to Almeida the stronger statement that it is exclusively synthetic. No generators, corpus names, examples, size, language mixture, license inventory, deduplication, filtering, contamination analysis, or train/eval split are disclosed. |
| Supervised stage | **Unknown.** No SFT objective, sample format, curriculum, or mixture weights. |
| RLCD reward | **Unknown.** No reward equation, preference or outcome collection procedure, proper scoring rule, reward model, verifier, or calibration penalty. |
| RL algorithm | **Unknown.** No PPO, policy-gradient, direct optimization, rejection sampling, online/offline, or alternative algorithm is named. |
| Rollouts and labels | **Unknown.** No number of prompts, candidates, annotators, simulations, or synthetic-data generators. |
| Calibration stage | **Unknown.** No temperature scaling, isotonic regression, Platt scaling, ensemble, or end-to-end calibration procedure. |
| Optimization | **Unknown.** No optimizer, learning-rate schedule, batch size, sequence length, epochs or tokens, regularization, gradient clipping, or checkpoint selection. |
| Compute | **Unknown.** No accelerator type/count, FLOPs, training time, energy, or monetary cost. |
| Evaluation splits | **Unknown.** TypeSafe says the public workflows are not its training distribution, but it publishes no held-out RLCD calibration set or contamination audit. |
| Artifacts | **Not released.** No weights, tokenizer, config, intermediate checkpoints, training code, or inference code. |

This means "trained with RLCD" currently names an objective and product philosophy, not a reproducible algorithm. No independent team could determine whether it had reproduced RLCD or Jev from the available description.

## Probabilities, confidence, and calibration

Choice and Score expose the full probability distribution. Their `confidence` field is a statistic derived from the shape of that distribution, not a separate model judgment and not an independently estimated probability that the answer is correct. A peaked distribution gives higher confidence; a flat distribution gives lower confidence. Noul exposes only `P(yes)` because its binary probability already carries the uncertainty. [Confidence](https://docs.typesafe.ai/confidence)

The documentation's interactive three-option demo uses

```text
(3 * largest_probability - 1) / 2
```

and describes that expression as an **approximation** for the demo. For `n` options its code generalizes to `(n * p_max - 1) / (n - 1)`. TypeSafe does not state that this is the exact production formula, and it does not publish the production Score-confidence formula. [Confidence source and demo](https://docs.typesafe.ai/confidence)

Three concepts must stay separate:

- **Probability** is the model's distribution over an answer.
- **Confidence** is a summary of how concentrated that distribution is.
- **Calibration** is an empirical relationship between predicted probability and observed frequency on a labeled population.

The public evaluations report agreement/accuracy, cost, and latency. They do not publish reliability diagrams, expected calibration error, maximum calibration error, Brier score, log loss, calibration slope/intercept, or calibration by domain and confidence bucket. Therefore, TypeSafe's central calibration claim cannot be quantitatively audited from the published benchmark. A peaked distribution can also be wrong. Almeida explicitly confirmed that Jev can be confidently wrong. [Founder comment](https://news.ycombinator.com/item?id=49718780) · [Official evaluations](https://evals.typesafe.ai/)

## How good Jev is

### TypeSafe's workflow benchmark

TypeSafe evaluates four end-to-end workflows: security incident handling, agent-trace observability, invoice processing, and customer service. Every model receives the same workflow structure. It also evaluates prompt-only baselines, although Jev is a workflow-only model. The reference probability for each decision is the average of GPT-6 Astra and Claude Fable 5.1 at high thinking. Tested models use their provider's default reasoning setting. Each aggregate point gives equal weight to the four workflows. [Evaluation overview and methodology](https://evals.typesafe.ai/)

The current aggregate workflow results are:

| Model | Agreement called “accuracy” | Cost per case | Time per case |
|---|---:|---:|---:|
| **Jev** | **67.8%** | **$0.0004** | **0.4 s** |
| Haiku 4.5 | 53.6% | $0.0195 | 12.5 s |
| Opus 5 | 73.1% | $0.1761 | 37.8 s |
| Sonnet 5 | 67.8% | $0.1174 | 78.1 s |
| DeepSeek v4 Flash | 64.4% | $0.0059 | 51.9 s |
| DeepSeek v4 Pro | 65.5% | $0.0413 | 86.5 s |
| OpenAI Luna | 66.8% | $0.0033 | 12.9 s |
| OpenAI Sol | 74.1% | $0.0836 | 23.3 s |
| OpenAI Terra | 67.9% | $0.0304 | 10.1 s |

All figures are TypeSafe's own and are read directly from the current [evaluation overview](https://evals.typesafe.ai/). On this table Jev roughly matches Sonnet 5 and Terra in agreement, but not Sol or Opus 5. Rounded aggregate ratios put Jev about 76 times cheaper and 25 times faster than Terra, not at the launch headline's maximum 444.6 times cheaper and 193.6 times faster. TypeSafe says those larger figures are the high end of expected real gains, not a universal multiplier. [Launch caveat](https://typesafe.ai/blog/introducing-system-one-models-and-jev)

The homepage places the 193.6x/444.6x headline near a displayed example of $0.000081 vs $0.013880 and 0.114 s vs 8.566 s. Those displayed pairs imply about 171x cheaper and 75x faster, respectively, so the nearby example is not the calculation behind the headline. [TypeSafe homepage](https://typesafe.ai/)

Jev's own per-workflow results show substantial task variation:

| Workflow | Cases | Jev agreement | Cost | Time | Highest listed workflow result |
|---|---:|---:|---:|---:|---|
| [Security incidents](https://evals.typesafe.ai/security_incidents.html) | 240 | 61.7% | $0.0001 | 0.3 s | Opus 5, 66.2% |
| [Agent trace observability](https://evals.typesafe.ai/agent_trace_observability.html) | 117 | 71.6% | $0.0003 | 0.5 s | Sol, 76.6% |
| [Invoice processing](https://evals.typesafe.ai/invoice_processing.html) | 150 | 61.8% | $0.0011 | 0.5 s | Sol, 79.1% |
| [Customer service](https://evals.typesafe.ai/customer_service.html) | 204 | 76.0% | $0.0001 | 0.4 s | Sol, 78.3% |

Invoice processing is the clearest weakness in this suite. Jev trails Sol by 17.3 percentage points there. Its customer-service result is much closer to the leader.

### What the benchmark does and does not establish

The benchmark establishes that TypeSafe can run its own four structured workflows much faster and more cheaply with Jev than with the listed generative APIs, while often getting similar model-consensus answers. It does not establish human-level correctness, factual truth, broad task generalization, or calibration.

Important limitations, several acknowledged by TypeSafe, are:

- The labels come from two other models. Agreement with their average is not ground truth. It also favors behavior similar to those model families.
- TypeSafe's model-capabilities team created the workflows. The launch post says they were not deliberately constructed to flatter Jev and are not from its training distribution, but acknowledges possible bias.
- The pages state 711 cases in total, but publish only five selected case payloads for each workflow. They do not publish confidence intervals, error bars, or statistical tests.
- No full downloadable 711-case corpus, raw predictions, labels, evaluation runner, or immutable benchmark revision is linked. The four public JavaScript payloads expose selected examples, not the full benchmark.
- Claude Fable 5.1 refused or failed on 88 security documents. TypeSafe substituted Opus 5 for those reference values, making the reference procedure non-uniform in that workflow.
- Cost and end-to-end latency depend on provider pricing, network location, default reasoning settings, load, and retries. TypeSafe says the LLM measurements used laptops on the US West Coast because the service was based there.
- There is no standard public benchmark table. TypeSafe deliberately rejects persistent public leaderboards in favor of dated, one-off evaluations that it says it will retire rather than optimize against. This may reduce benchmark hill climbing, but it also reduces comparison with the broader literature. [Anti-benchmaxxing](https://typesafe.ai/blog/antibenchmaxxing)

### Repeatability is not correctness

TypeSafe has published two small repeated-call studies:

- In the Choice study, one borderline moderation post was evaluated with eight questions over 15 repeats on `jev-1.13.0`. The raw winning-label repeatability was 90.8%. Adding an application rule that treated top probability below 0.60 as uncertain increased policy agreement to 99.2%, while sending 25.8% of decisions to the uncertain path. Claude Haiku 4.5 at temperature zero had 100% repeatability in the same example. The study correctly says it measures repeatability, not accuracy. [Choice consistency](https://docs.typesafe.ai/cookbooks/consistency_choice_cookbook)
- In the Noul study, one insurance claim was evaluated with 14 propositions over 15 repeats. The broad coverage proposition varied from roughly 0.43 to 0.53 and crossed the 0.5 decision boundary. The exclusion proposition varied from roughly 0.53 to 0.62. Review bands can stop those changes from becoming opposite automatic actions, but do not make the underlying predictions deterministic or correct. [Noul consistency](https://docs.typesafe.ai/cookbooks/consistency_noul_cookbook)

The Choice study injects a random identifier on every call. Consequently, it cannot cleanly separate sampler nondeterminism from sensitivity to an irrelevant state change. Both studies use one underlying item, so they are implementation examples rather than general stability estimates.

### Parallel-question performance

An official cookbook asks 13 questions about the GDPR Wikipedia article. Five runs on `jev-1.12` averaged $0.000497 and 0.27 seconds when batched, versus $0.006090 and 2.71 seconds for 13 sequential calls. TypeSafe reports the same answers and describes this as 12.2 times cheaper and 10.0 times faster. The latency comparison is against sequential calls; a concurrent single-call baseline would narrow the wall-clock gap. [Parallel questions](https://docs.typesafe.ai/cookbooks/parallel_questions)

Another official cookbook reranks 30 BM25 passages for each of 40 CLERC legal queries. It reports top-1 accuracy rising from 5% to 18% and top-10 accuracy from 38% to 62%. This is a small application result, not a general Jev benchmark, and there is no reported comparison against modern neural rerankers. [Re-ranking cookbook](https://docs.typesafe.ai/cookbooks/rerank_typesafe)

## Independent evidence available so far

These four repositories are independent of TypeSafe, public, and substantially more auditable than a screenshot. They include code, protocol details, aggregate artifacts, and in some cases saved outputs. They were all created within days of Jev's launch, are author-run rather than independently replicated, and are not peer reviewed. Public datasets may also overlap unknown training data. Treat the results as strong early tests, not settled model characterization.

| Study | Design | Main result | Calibration, sensitivity, and scope caveats |
|---|---|---|---|
| [AbdelStark/jev-benchmarks](https://github.com/AbdelStark/jev-benchmarks) | 300 held-out examples, 100 each from AG News, Banking77/BTZSC, and DAIR Emotion; Jev 1.13 vs local GLiNER2.5; pinned revisions and paired bootstrap intervals | Jev accuracy was **0.910 vs 0.700** on AG News and **0.870 vs 0.610** on Banking77. Emotion was **0.480 vs 0.440**, with the difference unresolved. | On Emotion, Jev was worse calibrated: Brier **0.846 vs 0.668**, NLL **5.588 vs 1.381**, and zero probability on the true label for 16% of examples. Jev p50 was about 236 to 256 ms; GLiNER was about 44 ms on the low-cardinality tasks and 296 ms on the 72-label task. The sample is small, and 5% error-budget thresholds were selected and evaluated on the same slice. Hosted-network and local-CPU latency are not normalized. |
| [anessbelbati/jev-rerank-bench](https://github.com/anessbelbati/jev-rerank-bench) | Eight English retrieval datasets, 1,617 scored questions; same 30 BM25 candidates; Jev variants vs Cohere Rerank 4, ZeroEntropy, DeepSeek, and an open Qwen approximation | Equal-dataset nDCG@10 was **0.692 for Jev's rubric vs 0.691 for Cohere Pro**. The interval on the gap was inconclusive. Jev averaged **422 ms and $0.45 per 1,000 queries** vs Cohere's **844 ms and $2.51**. Jev scored 71% on NevIR negation pairs vs Cohere's 67%. | Equal-query weighting reverses the headline order: Cohere **0.756 vs Jev 0.738**. Reversing the 30 candidates changed Jev's top passage on **24.7%** of questions. The study calls this order sensitivity, not repeatability. nDCG is ranking quality, not classification accuracy; exploratory intervals are uncorrected for multiple comparisons. |
| [Gaurav-Gosain/jev-sec-bench](https://github.com/Gaurav-Gosain/jev-sec-bench) | Prompt-injection detection on 662 messages plus 200 matched vulnerable/secure code pairs | With deployment context, injection detection reached **96.5% accuracy, 0.9927 AUROC, 0.0588 ECE**, and 325 ms p50. Adding the application's purpose to state raised recall from **74.9% to 95.1%**. Jev ranked the vulnerable sample above its secure twin in **178/200, or 89.0%**, of code pairs. | Absolute vulnerable-code accuracy was **71.5%** and ECE **0.1868**. The repository demonstrates label errors in the synthetic code corpus, so pairwise ranking is more trustworthy than absolute accuracy. Injection data is specific to one news-assistant policy, much of it German; the study used a single run and small per-class cells. |
| [anisselbd/jev-phishing-bench](https://github.com/anisselbd/jev-phishing-bench) | 2,000 synthetic emails; direct Jev verdict vs Claude Haiku 4.5; repeated passes; five decomposed signal questions and held-out regressions | Direct Jev was **62.6% accurate, 0.689 AUROC, 0.154 ECE** vs Haiku's **81.3%, 0.837, 0.097**. Jev p50 was 239 ms vs 687 ms and about 12 times cheaper. A held-out logistic regression over five Jev signals reached **95.0%**, statistically tied in accuracy with **93.2%** from the equivalent Haiku signals. | A two-feature regex already reached **91.8%**, showing the dataset is separable by construction. Jev's best single signal scored 89.4%, below both the regex and Haiku's equivalent 94.2% signal. The email bodies are synthetic, ground truth comes from URL reputation feeds, and direct verdicts are prompt-sensitive. Decomposition plus downstream fitting tests a system, not standalone Jev accuracy. |

Together, these studies suggest a more precise profile than the launch headline:

- Jev can be a strong zero-shot classifier and reranker in some domains, especially when many labels or judgments share one state.
- It is not uniformly faster than a compact local classifier, nor uniformly more accurate than a low-cost LLM.
- Calibration varies sharply by task. The very poor Emotion NLL and the direct phishing ECE are direct counterexamples to reading "calibrated decisions" as a universal property.
- Decomposing a decision into useful signals can work much better than asking for the final verdict, but hand-designed features, application context, and downstream code deserve part of the credit.
- Option order, question wording, and context framing can materially change behavior. Any deployment should test these axes explicitly.

## Documented failures and jagged edges

TypeSafe maintains an unusually direct limitations page for `jev-1.13`, last reviewed 17 September 2026. The following are published limitations, not inferred ones. [Jev 1.13 jaggedness](https://docs.typesafe.ai/model-jaggedness/jev-1.13)

1. **Literal interpretation.** Jev can follow the surface form of a question, including negation, while missing implied intent. Small wording changes can reverse the requested proposition. Write the precise proposition you want evaluated.
2. **Arithmetic and exact numbers.** It is not a calculator. Counting, exact arithmetic, and numeric precision are unreliable. Semantic names work better than raw numeric encodings, and high-level program descriptions work better than assembly or binary. A Score should not be used to reconstruct an exact quantity.
3. **Dates and times.** Comparisons are unreliable, especially with mixed formats, relative dates, and domain boundaries. Extract parts with bounded choices and perform calendar arithmetic in code.
4. **Indirection and multiple reasoning hops.** Double negation, indirect phrasing, and chained reasoning reduce accuracy. Decompose the task into direct, atomic questions.
5. **Long irrelevant state.** Accuracy declines as irrelevant material grows. The 64k context limit is a maximum, not a claim of constant quality throughout the window.
6. **Prompt injection and adversarial state.** Instructions embedded in the state can steer the model. TypeSafe says Jev is not designed to treat state as hostile by default. This is a material safety limitation for email, webpages, tickets, logs, retrieved documents, and any other untrusted input.
7. **Contradictory instructions or criteria.** Overlapping or conflicting definitions confuse the model. The options need to be mutually clear in the application designer's wording.
8. **No structural probability invariants across separate questions.** Semantically equivalent formulations can produce very different values. TypeSafe gives an example in which equivalent Noul and Choice questions produce 0.22 versus 0.01 for yes. A proposition and its negation can also fail to sum to one, with a published example of 0.72 and 0.47. Independent questions are not a coherent joint probability model.
9. **No text generation.** Constructing strings through chains of choices is slow and poor. Use an LLM when the job requires new prose, code, or an open-ended sequence.

Additional documented boundaries are:

- English is stronger than other languages. [Models](https://docs.typesafe.ai/models)
- Repeated requests can vary even without a user-visible sampling control. [Consistency cookbooks](https://docs.typesafe.ai/cookbooks/consistency_choice_cookbook)
- Confidence thresholds must be selected for the domain and risk. TypeSafe tells users to test on their own data. [Confidence](https://docs.typesafe.ai/confidence)
- The public workflow evaluation shows task jaggedness, with Jev ranging from 61.7% to 76.0% agreement across four domains. [Evaluation overview](https://evals.typesafe.ai/)

## “Zero hallucinations” needs a narrow definition

TypeSafe's site says Jev has "zero hallucinations." The defensible part of that claim is **output-shape safety**: a Choice can only select one supplied option, a Score can only use supplied levels, and a Noul can only return a yes probability. Jev cannot invent a prose citation, unsupported JSON key, or malformed free-form answer because free-form generation is not in the interface. [TypeSafe homepage](https://typesafe.ai/) · [System One](https://docs.typesafe.ai/concepts/system-one)

It can still make a false semantic decision, assign the wrong option high probability, contradict another question, follow an injected instruction in the state, or misread a date or number. The founder's statement that it can be confidently wrong and the company's jaggedness examples make this explicit. [Founder comment](https://news.ycombinator.com/item?id=49718780) · [Jaggedness](https://docs.typesafe.ai/model-jaggedness/jev-1.13)

Therefore:

- **Schema hallucination:** prevented by the closed output contract.
- **Semantic error:** not prevented.
- **Confident semantic error:** explicitly possible.
- **Cross-question logical inconsistency:** explicitly documented.

This distinction is essential for any high-stakes deployment.

## Runtime, pricing, and deployment

The current hosted product facts are:

| Property | Jev 1.13 |
|---|---|
| Version | `jev-1.13.0`; both `jev-latest` and `jev-preview` currently resolve to it |
| Input | Text, text-bearing JSON object, or array of text values |
| Context | 64k total request; 32k for state plus longest question |
| Price | $42 per billion input tokens, or $0.042 per million; output free |
| Published limits | 250k tokens/second and 1,200 requests/minute, dynamically adjusted |
| Typical claimed latency | 70 to 500 ms end to end in the launch material |
| Public interface | Hosted HTTPS API plus Python and TypeScript/JavaScript SDKs |
| Weights | Not published |
| Self-hosting | Not documented or offered publicly |
| Customer tuning | No per-account fine-tune or LoRA; same weights for all accounts |
| Customer data | TypeSafe says requests and responses are not used for training; enterprise zero-data-retention is available |

Sources: [Models](https://docs.typesafe.ai/models), [launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev), [SDK documentation](https://docs.typesafe.ai/sdk), and [legal/data-handling index](https://docs.typesafe.ai/legal).

The API can report overload as HTTP 529 and rate limiting as 429. Official SDKs retry eligible failures with backoff and honor `retry-after`; direct HTTP clients need to implement this. These retries affect tail latency and must be logged in a serious benchmark. [API reference](https://docs.typesafe.ai/api)

## What the public code contains

The official GitHub organization exposes ten repositories on the research date. The Jev-adjacent ones are:

- [`typesafe-sdk-python`](https://github.com/typesafe-ai/typesafe-sdk-python): generated/handwritten client types and HTTP plumbing for the hosted API.
- [`typesafe-sdk-js`](https://github.com/typesafe-ai/typesafe-sdk-js): the equivalent TypeScript/JavaScript SDK.
- [`system-one-adapter-python`](https://github.com/typesafe-ai/system-one-adapter-python): a drop-in `TypeSafeClient` replacement that asks third-party LLMs to emulate the System One contract. It supports structured-output and prompted-JSON modes for comparing cost, speed, and behavior. It is not Jev.
- [`skills`](https://github.com/typesafe-ai/skills): agent instructions and integration examples for the API.

None contains model weights, a tokenizer, architecture configuration, training data, RLCD code, a reward function, or a Jev inference server. The adapter is useful for reproducing TypeSafe's interface with an LLM baseline, but it cannot reproduce Jev's internal model.

The adapter's [public confidence utility](https://github.com/typesafe-ai/system-one-adapter-python/blob/main/src/system_one_adapter/_utils/confidence_metrics.py) computes Choice confidence from the top probability above the uniform baseline and Score confidence from concentration around the modal level. These formulas describe the LLM adapter. TypeSafe does not say that Jev's production service uses the same Score formula.

## Papers and model cards

I found **no paper or technical report about Jev**, and no paper that defines RLCD as an algorithm. Exact searches of arXiv for the model name, company name, and the expanded RLCD phrase produced no relevant result on 21 September 2026. The founder's launch-thread response says only that a paper has been discussed and the architecture remains private. [arXiv search: Jev and TypeSafe](https://arxiv.org/search/?query=Jev+TypeSafe+AI&searchtype=all) · [arXiv search: Reinforcement Learning for Calibrated Decisions](https://arxiv.org/search/?query=%22Reinforcement+Learning+for+Calibrated+Decisions%22&searchtype=all) · [Founder comment](https://news.ycombinator.com/item?id=49718824)

There is an unrelated 2023 paper titled [RLCD: Reinforcement Learning from Contrastive Distillation](https://arxiv.org/abs/2307.12950). It was published before Jev and expands the acronym differently. It is not TypeSafe's Reinforcement Learning for Calibrated Decisions and should not be used as its method paper.

Diogo Almeida is a coauthor of the 2022 InstructGPT paper, [Training language models to follow instructions with human feedback](https://arxiv.org/abs/2203.02155). That paper explains an earlier RLHF system. TypeSafe cites Almeida's role to establish the team's training background and contrasts RLCD with RLHF. It does **not** describe Jev, System One architecture, calibrated-decision training, or the Jev dataset. It cannot serve as a reproduction guide for this model.

No official Jev model card or checkpoint is linked from TypeSafe's site, documentation, GitHub organization, or launch material. The `Models` documentation is a service card with limits, pricing, aliases, language notes, and data-handling claims, not a research model card with lineage, parameter count, dataset composition, training compute, evaluations, and weights. [Models](https://docs.typesafe.ai/models)

Hugging Face API queries for models, datasets, and Spaces under the canonical `typesafe-ai` organization name returned no Jev artifacts on the research date. This is an artifact check, not a claim that every similarly named third-party account belongs to TypeSafe. [Models API](https://huggingface.co/api/models?author=typesafe-ai&limit=100) · [Datasets API](https://huggingface.co/api/datasets?author=typesafe-ai&limit=100) · [Spaces API](https://huggingface.co/api/spaces?author=typesafe-ai&limit=100)

The official pre-launch talk [What's next after RLHF?](https://ai.engineer/talks/cJ0EOzey--o-whats-next-after-rlhf) also does not fill the gap. When asked about a classifier-style head and whether reward arrives only at the end or throughout a process, Almeida declined to disclose the architecture or training mechanics.

## What can be reproduced today

### Hosted behavior

A careful user can reproduce an evaluation of the hosted model, subject to service updates and nondeterminism:

1. Pin `jev-1.13.0`, not an alias.
2. Freeze the complete request JSON, including exact state order, whitespace where relevant, instructions, criteria, and question order.
3. Save the full response, resolved model ID, token usage, wall-clock latency, HTTP status, retry count, date, client region, and SDK version.
4. Repeat each item enough times to measure variation. Do not treat one call as deterministic.
5. Use independently labeled data from the intended deployment domain. Do not use an LLM consensus as the only truth source for a high-stakes test.
6. Report task accuracy and per-class errors. For the probabilities, also report log loss or Brier score, reliability curves, calibration error by bucket, and selective accuracy/coverage at the exact thresholds the application will use.
7. Test paraphrases, negation, long irrelevant context, conflicting criteria, non-English cases, numbers and dates, and prompt injection. These target the failure modes TypeSafe already documents.
8. Re-run the suite before adopting a new version. Aliases can move and tuned thresholds may no longer be valid.

Steps 1, 2, 3, and 8 follow TypeSafe's versioning and API guidance. The metric and adversarial-test details are recommended independent evaluation practice, not a claim that TypeSafe used them internally.

### The interface with an alternative model

The official [System One LLM adapter](https://github.com/typesafe-ai/system-one-adapter-python) lets a user run the same typed request contract through supported generative-model APIs. That makes an interface-level baseline reproducible. It does not test the same architecture, sampling process, or training objective.

### Jev itself

Reproducing Jev would require at minimum:

- an architecture specification and exact configuration;
- a tokenizer and input serialization spec;
- the pretraining checkpoint or full pretraining recipe;
- training data or a sufficiently detailed construction pipeline;
- the RLCD mathematical objective, reward construction, and optimization algorithm;
- calibration and checkpoint-selection methods;
- training hyperparameters and compute budget;
- model weights or deterministic training code;
- an inference/sampling specification and validation vectors.

None is public. Any model built now from a text encoder plus classifier or candidate-scoring head would be a plausible System One-like baseline, but it would be an independent design, not a reproduction of Jev.

## An open approximation, not a Jev reproduction

The following is a concrete design that an independent team could build. It reproduces the useful **interface properties**: one shared state, dynamic natural-language candidates, parallel typed outputs, explicit probabilities, and no string generation. There is no evidence that Jev uses this architecture, data recipe, loss, or calibration method. It should be named as a new baseline, not Jev, System One, or RLCD.

### 1. Model shape

Use an open pretrained text encoder with two cooperating towers:

```text
state tokens S  -> shared state encoder -> Hs

for every question q and candidate c_i in parallel:
    [question instructions q ; candidate description c_i]
        -> candidate encoder -> h_qi
    late interaction / cross-attention(Hs, h_qi)
        -> scalar logit z_qi
```

The state tower runs once per request. The question/candidate tower runs on a padded batch containing every candidate from every question. A small cross-attention or late-interaction scorer combines each dynamic candidate with the reusable state representation. This is more general than a fixed classification head because the option names and descriptions arrive at inference time.

For a stronger but more expensive version, let question/candidate tokens cross-attend to cached state keys and values. For a faster version, pool the state and candidate embeddings and score their concatenation, elementwise product, and difference through a multilayer perceptron. Train and benchmark both. The faster version may lose fine-grained evidence alignment.

The typed heads are simple:

```text
Choice:
    p_i = softmax(z_i over the candidates)
    answer = argmax_i p_i

Noul:
    p_yes = sigmoid(z_yes)
    # or score explicit yes/no descriptions and use a two-class softmax

Score with K ordered levels:
    p_k = softmax(z_k over the level descriptions)
    score = sum(k * p_k for k in 0..K-1)
```

Return the full distributions. If a single confidence scalar is desired, define it publicly. For example, normalized entropy `1 - H(p) / log(K)` is transparent and uses the whole distribution. Do not call it the probability of correctness. For Noul, return `p_yes` without adding a second number unless that second number has an independently validated meaning.

### 2. Input serialization and invariants

Define a canonical, versioned serialization rather than relying on incidental JSON order:

```text
<STATE>{canonical state}</STATE>
<QUESTION>{instructions}</QUESTION>
<CANDIDATE>{candidate description}</CANDIDATE>
```

Do not place the application key or numeric Score index in the learned input. Randomize candidate order during training. Preserve a stable mapping back to caller keys in ordinary code. Enforce the following as explicit tests:

- permuting Choice options should permute the probability vector but not change its mapped meaning;
- batching questions should match evaluating them separately within a declared numerical tolerance;
- adding an irrelevant question should not change existing answers;
- a proposition and its explicit negation should be approximately complementary when both are evaluated under the same context;
- equivalent Noul and two-option Choice formulations should agree after calibration;
- malformed or out-of-schema outputs should be impossible because the server constructs JSON from tensors.

Jev itself does not satisfy all of these probability invariants, according to TypeSafe's jaggedness page. They are design goals for the approximation, not claims about the product.

### 3. Data construction

Build a versioned mixture with four parts:

1. **Licensed labeled datasets.** Convert classification, entailment, ranking, moderation, routing, ordinal scoring, and retrieval datasets into the common state/question/candidate form. Keep the source license and immutable revision for every row.
2. **Programmatically generated tasks.** Create controlled examples for dates, arithmetic, negation, option order, missing evidence, contradictory criteria, long distractors, and logically related propositions. These provide exact labels and targeted counterfactual pairs.
3. **Synthetic teacher ensemble.** Ask several strong teacher models independently for distributions over the same bounded candidates. Randomize prompt templates and candidate order. Preserve teacher disagreement as a soft target rather than forcing a majority label. Remove examples with parsing failures, suspiciously identical rationales, or source leakage.
4. **Human-reviewed ambiguous and adversarial cases.** Sample high-disagreement, high-loss, and high-impact items for adjudication. Include untrusted text containing prompt injections and text that argues for its own classification.

Split by source, template family, and semantic cluster before generating final examples. Otherwise paraphrases or synthetic siblings can leak across train and test. Keep a hidden calibration split and a separate frozen evaluation split. Publish dataset cards, generation prompts, teacher versions, filters, hashes, and exclusion counts.

### 4. Training objectives

Start with supervised probabilistic training. For a hard label `y`, use cross-entropy and multiclass Brier loss:

```text
L_hard = CE(p, y) + lambda_brier * sum_i (p_i - 1[i=y])^2
```

For a teacher distribution `t`, distill its uncertainty with KL divergence or soft cross-entropy:

```text
L_soft = -sum_i t_i * log(p_i)
```

Add narrowly scoped consistency losses:

```text
L_order       = divergence(p(original), unpermute(p(permuted)))
L_paraphrase  = divergence(p(q), p(paraphrase(q)))
L_complement  = abs(p(A) + p(not A) - 1)
L_batch       = divergence(p(batched), p(separate))
```

The complete objective can be a weighted sum of `L_hard`, `L_soft`, and the relevant invariance terms. Sample weights should prevent large synthetic families from overwhelming real labeled data. Report ablations for every loss and data component.

This recipe does not need reinforcement learning. If reinforcement learning is added, specify what action is sampled, what reward is observed, the proper scoring rule used as reward, the reference policy, KL constraint, and the policy optimizer. Calling ordinary supervised distillation "RLCD" would be misleading. TypeSafe has not published enough detail to know whether any part of this resembles its procedure.

### 5. Calibration

After model selection, freeze the weights and fit temperature scaling on the held-out calibration split separately for Choice/Noul and Score if justified. Do not tune the temperature, confidence thresholds, and reportable accuracy on the same items. Compare global scaling with domain-conditioned scaling, but reject the latter if the domain labels will be unavailable or unstable at deployment.

Publish:

- negative log likelihood and Brier score;
- reliability plots and expected/max calibration error with the exact bins;
- calibration by task, language, class count, state length, and confidence range;
- selective risk and coverage at predeclared error budgets;
- the share of exact-zero probabilities on true labels;
- pre-calibration and post-calibration results, including accuracy to catch unintended changes.

Calibration must be rechecked after quantization, pruning, distillation, or a tokenizer change.

### 6. Inference implementation

Group candidates by question, batch the full request, encode the state once, and run all candidate interactions as one padded tensor operation. Construct the output JSON outside the model. An implementation should expose:

- model and tokenizer revision;
- maximum state tokens, questions, and candidates;
- deterministic mode and seed where supported;
- precision and quantization mode;
- per-stage time for tokenization, state encoding, candidate encoding/scoring, calibration, and serialization;
- cache-hit behavior and batch/concurrency limits.

Benchmark latency against sequential and concurrent baselines. Report local model time separately from network time. Test scaling over state length, number of questions, candidates per question, and concurrent requests.

### 7. Failure-focused evaluation

The test plan should include ordinary task quality plus the exact areas where official and independent evidence finds problems:

- direct wording, paraphrases, scope changes, negation, and double negation;
- complementary propositions and equivalent Choice/Noul formulations;
- candidate-order reversal and irrelevant-candidate insertion;
- exact arithmetic, counting, numeric encodings, dates, time zones, and relative time;
- long relevant context plus increasing distractor ratios;
- prompt injection and self-referential text in the state;
- overlapping and contradictory criteria;
- English and separately reported non-English languages;
- repeated identical calls and batched-vs-separate evaluation;
- direct verdicts versus decomposed signals combined by fixed code or a held-out downstream model;
- distribution shift, low-base-rate classes, and abstention/selective-coverage curves.

Use human or programmatic ground truth where possible, not only another model's consensus. Freeze protocols before inference, publish raw probability vectors and failures, use paired confidence intervals, and distinguish exploratory from confirmatory comparisons. This would produce a reproducible open typed-decision model. It would still say nothing about Jev's undisclosed architecture unless TypeSafe later publishes enough material for a direct comparison.

## Questions TypeSafe would need to answer for a technical audit

1. What is the exact Jev 1.13 computational graph and parameter count?
2. What pretrained checkpoint or pretraining procedure precedes RLCD?
3. How are dynamic criteria represented and scored, and how does the parallel sampler work?
4. What is the RLCD objective in equations? Which parts are reinforcement learning rather than supervised probabilistic prediction or post-hoc calibration?
5. Who or what creates decision labels and probabilities? If synthetic generators are used, which models, prompts, filters, and quality controls are involved?
6. What are the data volumes, domain and language mixtures, licensing basis, deduplication rules, and contamination checks?
7. Which proper scoring rule or calibration loss is optimized? Is calibration post-hoc, end-to-end, or both?
8. What are reliability metrics by task, language, class balance, probability bucket, and distribution shift?
9. What causes repeated-call variation, and can customers request deterministic inference or a seed?
10. How does latency scale with state length, number of questions, number of criteria, concurrent load, and batch size?
11. Which hardware and inference implementation produce the quoted 70 to 500 ms range?
12. What safety training or isolation addresses prompt injection in untrusted state?
13. Will TypeSafe publish weights, a model card, evaluation cases, training code, or the discussed paper?

Until those questions have public answers, detailed claims about Jev's internal operation are speculation.

## Source map

### Product and training claims

- [Introducing System One models and Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev): launch, product thesis, latency and price claims, evaluation caveats, demos, RLCD FAQ, and high-level data statements.
- [TypeSafe homepage](https://typesafe.ai/): current positioning and headline comparison.
- [AI primer](https://docs.typesafe.ai/introduction/machine-learning-primer): RLHF/RLVR/RLCD framing and the definition of calibration.
- [The bitterest lesson](https://typesafe.ai/blog/bitterest-lesson): TypeSafe's data-first training philosophy.
- [Anti-benchmaxxing](https://typesafe.ai/blog/antibenchmaxxing): why TypeSafe does not plan a conventional persistent benchmark table.
- [Manifesto](https://typesafe.ai/manifesto): product philosophy about neural judgment inside deterministic software. It contains no model architecture or training recipe.
- [What's next after RLHF?](https://ai.engineer/talks/cJ0EOzey--o-whats-next-after-rlhf): pre-launch founder talk and audience questions; no architecture or reward mechanics are disclosed.
- [AI too good to be true, too bad to be useful](https://typesafe.ai/blog/ai-too-good-to-be-true-too-bad-to-be-useful-typesafe-ai): TypeSafe-hosted version of the broader product argument.

### Model and API documentation

- [Complete documentation index](https://docs.typesafe.ai/llms.txt)
- [Introduction](https://docs.typesafe.ai/introduction)
- [System One concept](https://docs.typesafe.ai/concepts/system-one)
- [Models](https://docs.typesafe.ai/models)
- [API reference](https://docs.typesafe.ai/api) and [OpenAPI JSON](https://api.typesafe.ai/openapi.json)
- [Choice](https://docs.typesafe.ai/primitives/choice), [Score](https://docs.typesafe.ai/primitives/score), and [Noul](https://docs.typesafe.ai/primitives/noul)
- [Confidence](https://docs.typesafe.ai/confidence)
- [Jev 1.13 jaggedness](https://docs.typesafe.ai/model-jaggedness/jev-1.13)
- [Legal and data handling](https://docs.typesafe.ai/legal)

### Evaluations and examples

- [Workflow evaluation overview](https://evals.typesafe.ai/)
- [Security incidents](https://evals.typesafe.ai/security_incidents.html)
- [Agent trace observability](https://evals.typesafe.ai/agent_trace_observability.html)
- [Invoice processing](https://evals.typesafe.ai/invoice_processing.html)
- [Customer service](https://evals.typesafe.ai/customer_service.html)
- Selected example payloads: [security](https://evals.typesafe.ai/security_incidents-cases.js?v=6c96b19f), [agent traces](https://evals.typesafe.ai/agent_trace_observability-cases.js?v=4a3821a2), [invoices](https://evals.typesafe.ai/invoice_processing-cases.js?v=8c2f8869), and [customer service](https://evals.typesafe.ai/customer_service-cases.js?v=066f789b)
- [Choice consistency](https://docs.typesafe.ai/cookbooks/consistency_choice_cookbook)
- [Noul consistency](https://docs.typesafe.ai/cookbooks/consistency_noul_cookbook)
- [Parallel questions](https://docs.typesafe.ai/cookbooks/parallel_questions)
- [Re-ranking](https://docs.typesafe.ai/cookbooks/rerank_typesafe)

### Code and founder clarifications

- [Official TypeSafe GitHub organization](https://github.com/typesafe-ai)
- [Python SDK](https://github.com/typesafe-ai/typesafe-sdk-python)
- [JavaScript SDK](https://github.com/typesafe-ai/typesafe-sdk-js)
- [System One LLM adapter](https://github.com/typesafe-ai/system-one-adapter-python)
- [Agent skills](https://github.com/typesafe-ai/skills)
- [Architecture and paper comment](https://news.ycombinator.com/item?id=49718824)
- [Zero-shot classifier exchange](https://news.ycombinator.com/item?id=49718490)
- [Parallel-output constraint](https://news.ycombinator.com/item?id=49719122)
- [General model and no task-specific training](https://news.ycombinator.com/item?id=49719245)
- [Confidently wrong clarification](https://news.ycombinator.com/item?id=49718780)

### Independent and attributed secondary evidence

- [jev-benchmarks](https://github.com/AbdelStark/jev-benchmarks)
- [jev-rerank-bench](https://github.com/anessbelbati/jev-rerank-bench)
- [jev-sec-bench](https://github.com/Gaurav-Gosain/jev-sec-bench)
- [jev-phishing-bench](https://github.com/anisselbd/jev-phishing-bench)
- [TechCrunch interview](https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/), used only for the attributed statement that the data is exclusively synthetic and the architecture remains private

## Final assessment

Jev is technically interesting because it changes the service contract rather than asking a chat model to imitate a classifier. The bounded outputs, shared-state fan-out, low reported latency, and low price are concrete and testable. The company also publishes more failure guidance than many model vendors.

The public evidence does not support a detailed mechanistic account of how Jev works or how it was trained. "New architecture," "parallel sampler," and "RLCD" are names for undisclosed internals. The published benchmark supports a claim of an unusually strong speed/cost tradeoff on four TypeSafe-designed workflows, not a claim that Jev is generally more accurate than frontier models. Its closed schema eliminates malformed or invented output types, not wrong judgments.

For deployment, treat Jev as a fast, probabilistic, hosted classifier/ranker whose probabilities require domain-specific validation and whose input state is vulnerable to adversarial steering. For research or reproduction, wait for the promised/discussed paper, weights, training details, and auditable calibration results. They are not public as of the date of this report.
