103 of 561 catalog entries

Benchmarks & research — built on Jev

Evaluation material — measured latency, task accuracy, and reproductions of published claims. Treat these as third-party work with their own methodology, and read the setup notes before quoting any number.

This directory does not publish its own benchmark claims. Figures quoted elsewhere on this site come from TypeSafe public materials and are labelled as such.

Project page111

JevK5

Apache-2.0 open-weight Jev alternative: noul / choice / score with probabilities, TypeSafe-style /v1/systemone, weights on Hugging Face.

On-site pageView
GitHub1.8k

SemIf

Semantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.

github.comOpen
GitHub959

jevlike

Train a small model that chooses among a changing list of text options, one probability per option in a single pass. Includes Doom, chess, and Wikispeedia demos.

github.comOpen
GitHub137

jev-visual

An educational Jev-like visual inference experiment on Apple Silicon: shared context, direct candidate scoring, and local visual demos.

github.comOpen
GitHub105

openjev

Open, Jev-compatible System One decision server on DiffusionGemma.

github.comOpen
GitHub89

jev-eval-agent

Personal-assistant agent built on Vercel's eve with 100 mocked tools, measuring how many steps it takes when Jev picks the tool versus the LLM.

github.comOpen
GitHub86

decider

One-pass typed decisions with calibrated probabilities (System One style model), fine-tuned from Qwen3.5-2B.

github.comOpen
GitHub74

reflex

A small open decision model: state + typed questions -> calibrated probabilities. A Jev / System One re-creation on Qwen3.5.

github.comOpen
Project page69

jevfire

JEV-inspired parallel decisions for CUDA LLMs. One context, many decisions. vLLM API, game-agent examples, and reproducible benchmarks.

On-site pageView
Project page68

jevmlx

Jev-style parallel constrained decisions for any MLX model on Apple Silicon. Typed, schema-valid JSON in one forward pass.

On-site pageView
Project page45

LitJev

A reproduction of Jev that turns any Qwen model into a fast decision model, serving the same /v1/systemone schema (Choice, Score, Noul) with no training and no generated answer text.

On-site pageView
GitHub44

open-jev (daseinlabs)

One-pass option scoring with a local Gemma 3 4B on Apple silicon via MLX, inspired by jevlike, with a Doom demo.

github.comOpen
GitHub32

typesafe-ai-benchmark

This is a LLM Gateway that mimics typesafe ai structured output. Like an imposter Jev.

github.comOpen
Project page30

jevgpt

A chatbot built on a model that cannot generate text (TypeSafe AI's Jev, driven autoregressively).

On-site pageView
GitHub23

mini-jev

Mini-Jev: what a Jev-style typed-decision interface looks like on a frozen Qwen3-4B — read the option letter's logits instead of generating JSON. Preregistered experiment, results, teaching bench.

github.comOpen
GitHub21

Verdict-open-jev

Non-autoregressive decision engine on ModernBERT (151M) with calibrated uncertainty (RLCD), TypeSafe AI Jev benchmark audit, and in-browser WebGPU playground.

github.comOpen
GitHub17

von

The open-source System One decision model. Sub-15ms, non-autoregressive, local drop-in alternative to TypeSafe Jev.

github.comOpen
GitHub16

jev-capability-atlas

Independent, evidence-based map of when TypeSafe's Jev actually holds up vs. breaks down — real API-call receipts, not a leaderboard. 中文為主的雙語 repo。

github.comOpen
GitHub15

jev-column-race

Jev vs Gemini 3.8 Flash: labelling 1,000 app reviews, 4.1× faster and 7× cheaper.

github.comOpen
GitHub15

jev-on-a-laptop

Unofficial study: Jev-style parallel typed decisions on stock 1.5B-8B models on an Apple Silicon laptop. Benchmarks, research notes, and a Hugging Face Space demo.

github.comOpen
GitHub15

open-jev (JoshuaSP)

Typed JSON inference with DiffusionGemma, with Every and Jev benchmark results.

github.comOpen
GitHub14

jev_local

Replicating Jev with a local LLM.

github.comOpen
GitHub12

jevbetter

A stronger one-pass scorer over a variable list of text options: hashed n-gram encoder, rival-aware attention, gated head, temperature scaling, benchmarked against jevlike.

github.comOpen
GitHub12

openjev (zhihz)

Local bilingual probability decisions from context, questions, and candidate answers. Independent research preview inspired by TypeSafe Jev.

github.comOpen
GitHub11

jev-benchmarks

Probability-aware evaluation for typed decision models: calibration, selective risk, latency, and reproducible benchmarks.

github.comOpen
GitHub11

system-one

Batched single-token choice inference for open language models, compatible with TypeSafe.

github.comOpen
GitHub10

Jev_apps

看看 Jev 能做什么:用中英文讲清热门应用、工作原理和各自优缺点。Explore Jev apps with plain-language examples, explanations, and comparisons.

github.comOpen
GitHub10

open-alternative-jev

Open alternative to Jev: typed, calibrated decisions from any open-weights LLM in one forward pass (HF + vLLM), with benchmarks.

github.comOpen
GitHub9

openvons

Openvons (open-Jev): 有限選択肢に確率で答える判断層 — テキスト / 画像 / 日本語音声コマンド.

github.comOpen
GitHub9

TypeAR

Type-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev.

github.comOpen
GitHub7

jev_stock

An experimental JEV-powered framework for forecasting short-term stock price direction from structured market data.

github.comOpen
GitHub7

system-one-open

Open replica of TypeSafe's Jev: typed calibrated decisions in one forward pass, on Gemma 4 E2B / Gemma 3 270M (Modal).

github.comOpen
GitHub6

jevcal

Stop guessing confidence thresholds: calibrate, threshold, and drift-check typed decision models (TypeSafe Jev) against an LLM teacher.

github.comOpen
GitHub6

typesafe-local

Inspired by TypeSafe Ai, Ask a local LLM typed questions, get calibrated probabilities instead of text. Structured output without generation or parsing. MLX / Apple Silicon.

github.comOpen
GitHub5

jev-korean-benchmark

Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence.

github.comOpen
GitHub5

jev-lm

A word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval.

github.comOpen
GitHub5

jev.nu

Nushell module for the TypeSafe System One API: typed decisions with calibrated probabilities.

github.comOpen
GitHub5

jevify

An agent skill to discover TypeSafe Jev opportunities, design typed questions, and learn from recent community experiments.

github.comOpen
GitHub5

LegalForecastBench

LegalForecast-MTD benchmark alpha and official evaluation workflows.

github.comOpen
GitHub4

daf-jev

Daf-jev: composable Python toolkit for TypeSafe's Jev (System One) decision API — question builders, confidence gates, evaluator, calibration, CLI, MCP server, agent skill.

github.comOpen
GitHub4

jev-align

Calibrated alignment verifier for LLM responses and agent plans — powered by Jev.

github.comOpen
GitHub4

jev-search-rerank-eval

Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.

github.comOpen
GitHub3

jev-behavior-study

Independent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification.

github.comOpen
GitHub3

jev-chat

A chatbot from typed Jev decisions: hierarchical speculative decoding over System One probabilities.

github.comOpen
GitHub3

jev-curate

High-throughput synthetic & pretraining dataset sifter powered by TypeSafe AI Jev (api.typesafe.ai). Stream, filter, and score Parquet & JSONL datasets at 1,500+ rows/sec using System One typed decisions (Choice, Score, Noul).

github.comOpen
GitHub3

jev-harness

Decision harness for TypeSafe Jev — confidence gates, shadow mode, recipes, and evals. Claude CLI 48.9s → Jev 1.3s on the same row-filter job.

github.comOpen
GitHub3

jev-little-airways

A show-and-tell capability study for Jev, TypeSafe's System One decision model.

github.comOpen
GitHub2

calibre

Calibration and confidence-based routing measured on Banking77: 80.2% accuracy at $0.103 per 500 decisions.

github.comOpen
GitHub2

claude-jev

Claude Code plugin that scores review findings, debug hypotheses and design options with TypeSafe's Jev — calibrated probabilities instead of one more opinion.

github.comOpen
GitHub2

jev-benchmark

Benchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs.

github.comOpen
GitHub2

jev-for-engineers

Eight minimal working examples of TypeSafe's Jev (a System One model) applied to mechanical and electrical engineering: CAD/CAE/CAM routing, FEM result triage, DFM screening, BOM alignment, hallucination-proof extraction. Zero dependencies.

github.comOpen
GitHub2

jev-gate

Not every coding task needs your best model. Experimental Jev-powered model routing for Claude Code — V3 prototype runs today, V4 routes at the task boundary.

github.comOpen
GitHub2

jev-mode

I kept watching coding agents burn context on decisions that aren't hard - triage 400 tickets, tag 600 files, route to one of six teams. jev-mode moves those verdicts to a typed-judgment model. I A/B'd it: 78% fewer tokens, 16x less work-attributable input, accuracy 96.1% vs 93.7%. Python, no deps, MIT.

github.comOpen
GitHub2

jev-research-eval

Reproducible Jev Ultrafast research-browser eval harness + field note (QC’d cases, suite runner, report generator). Not investment advice.

github.comOpen
GitHub2

jev-sec-bench

Blind security benchmarks for Jev, TypeSafe's System One model: prompt injection and vulnerable code detection, built on jev-go.

github.comOpen
GitHub2

mcts-agent

Discriminative Monte Carlo Tree Search using TypeSafe Jev System One Primitives and Gemini.

github.comOpen
GitHub2

new-api-plugin-typesafe

TypeSafe AI System One (Jev) task plugin for QuantumNous/new-api — native /v1/systemone, synchronous evaluation, token billing.

github.comOpen
GitHub2

research_desk

TypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers.

github.comOpen
GitHub2

siftr

Fast, cheap judgment for AI coding agents: semantic search, focused reads and list picking in ~2s. CLI + MCP server on TypeSafe Jev. Benchmarked on SWE-bench.

github.comOpen
GitHub2

system-one-gemma

Open-source Jev-style System One decision model. Gemma 3 270M with a scoring head — fast, calibrated decisions in a single forward pass. No text generation. Inspired by TypeSafe.ai's Jev.

github.comOpen
GitHub2

tenbin

MCP server and agent skill for the TypeSafe AI System One API (Jev): decompose a judgment into Choice / Score / Noul questions, lint them, measure on labelled data, and put calibrated thresholds in code.

github.comOpen
GitHub2

trade-jev

Backtest Jev (TypeSafe) as a BUY/SELL/HOLD trader on NQ L10 order-book data.

github.comOpen
GitHub2

zerosweep

Autonomous System-One Triage Engine & Benchmark powered by TypeSafe AI (Jev). 75ms inference, $0 output tokens, and RLCD epistemic safety gates.

github.comOpen
GitHub1

decisionbridge

A Jev-inspired decision interface for existing LLMs. Explicit choices, scores, calibration, and review thresholds.

github.comOpen
GitHub1

jev-agent-failure-benchmark

Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).

github.comOpen
GitHub1

jev-as-a-judge

Using Jev as an evaluator.

github.comOpen
GitHub1

jev-carryforward

What your last session knew, scored against what this one is doing. MCP server: a per-project ledger written as things happen, recalled per task with TypeSafe's Jev evaluation model via Vercel AI Gateway.

github.comOpen
GitHub1

jev-chess

Chess moves, evaluations, persona opponents, and game classification with TypeSafe AI System One.

github.comOpen
GitHub1

jev-exploration

Jev (TypeSafe) exploratory thread: claim audit, live demos, and runnable code.

github.comOpen
GitHub1

jev-freeform

An observable raw-character chat experiment powered entirely by TypeSafe Jev Choice.

github.comOpen
GitHub1

jev-phishing-bench

Jev (TypeSafe) vs Claude Haiku 4.5 on 2 000 phishing emails: accuracy, calibration, latency, cost. Reproducible benchmark.

github.comOpen
GitHub1

jev-rerank-bench

Can a decision model beat dedicated rerankers? TypeSafe Jev vs Cohere Rerank 4 vs ZeroEntropy zerank-2 vs a chat-model baseline: 14 datasets, every raw API response, bootstrap ranges on every gap.

github.comOpen
GitHub1

jev-synergy-screening

Jev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels.

github.comOpen
GitHub1

kyotsu-ai-bench

AI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard).

github.comOpen
GitHub1

modelsystem

Curated catalog of System One / Decision Models — contributions for modelsystem.one.

github.comOpen
GitHub1

openjev-experiments

Experiments with openjev, an open Jev-style option-logit runner, on local models.

github.comOpen
GitHub1

padflow-jev-evals

Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.

github.comOpen
GitHub1

qwen-rlcd

Jev-style calibrated decision model (Choice/Score/Noul) on Qwen3.5-0.8B.

github.comOpen
GitHub1

RISC-jeV

I tortured Jev into being a RISC-V CPU.

github.comOpen
GitHub1

typesafe-vs-deepseek

TypeSafe (Jev) vs DeepSeek-flash: side-by-side speed/token/cost/accuracy comparison across invoice extraction, email classification, and reranking.

github.comOpen
GitHub

FinancialPredictionJev

Using Jev to test how well it predicts financial markets(just like most llms as of september 2026, it doesnt do that good).

github.comOpen
GitHub

jev-alpha-bench

Does Jev predict stock returns from news? It reads the news well; there is no tradeable alpha. Three arms separate reading from recall.

github.comOpen
GitHub

jev-anotacao-sentencas

Jev (TypeSafe) vs. Gemini 3.8 Flash vs. GPT-5.6 Luna na anotação estruturada de sentenças do TJSP: qualidade, tempo e custo.

github.comOpen
GitHub

jev-deferred-crispification

Position paper: the Hidden-Markov and fuzzy primitives missing from TypeSafe AI's Jev and System-One decision models. Two lemmas, one principle (Deferred Crispification), one architecture (BSF-S1).

github.comOpen
GitHub

jev-finance-benchmark

Typesafe.ai model jev finance benchmark.

github.comOpen
GitHub

jev-headline-bench

Can Jev pick the winner of a real headline A/B test? 64.5% across 10,984 Upworthy randomized experiments, 74.7% when the difference was decisive.

github.comOpen
GitHub

jev-jp-address

Jev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証.

github.comOpen
GitHub

jev-lab

TypeScript experiments, evaluations, and latency benchmarks for TypeSafe's Jev model.

github.comOpen
GitHub

jev-pick-and-place-study

A small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.

github.comOpen
GitHub

jev-playground (hegargarcia)

Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.

github.comOpen
GitHub

jev-report

发明 RLHF 的人,这次做了个不会说话的模型:Jev 独立研究报告。52 页 PDF + 50 条中文实测复现包 + 143 条可回溯数据表.

github.comOpen
GitHub

jev-routing-experiment

Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena.

github.comOpen
GitHub

jev-secret-detection

Measures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets.

github.comOpen
GitHub

jev-shadcn-lint-eval

A small second eval for shadcn-ui/lint that uses TypeSafe's Jev to judge the linter's own output.

github.comOpen
GitHub

jev-spam-eval

Zero-shot spam filtering with TypeSafe Jev Noul questions, compared with TF-IDF baselines.

github.comOpen
GitHub

jev-trace-classifier

Application of TypeSafe Jev (noul judgment primitive) on the collusion.wiki corpus: agent vs human page authorship, head-to-head vs local Qwen3.8-Flash-Next.

github.comOpen
GitHub

misereru-slide-jev

Ongoing Japanese research deck on Jev and System One models, maintained as Markdown slides.

github.comOpen
External

Parallel Constrained Decoding (Qwen2.5-1B-RLCD)

Hugging Face Space exploring open-source parallel constrained decoding as an alternative to Jev.

huggingface.coOpen
GitHub

PocketJev

On-device iPhone visual decision tool using MLX and Qwen3-VL direct option logits.

github.comOpen
GitHub

shade-arena-jev-monitor

Evaluating TypeSafe's Jev as a fast monitor and action gate for agent sabotage in SHADE-Arena, compared with Gemini 2.5 Flash/Pro.

github.comOpen
GitHub

system-one-adapter-rust

Rust port of TypeSafe system-one-adapter (LLM-backed system_one evaluations).

github.comOpen
GitHub

thaiexam-jev-charts

Charts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models.

github.comOpen
GitHub

Typesafe_chess_eval

An evaluation of typesafe AI chess. As it turns out, the AI isn't doing really well even though chess is not a particularly open-ended game. Still, it's only a prototype and this probably wasn't optimzied for games.

github.comOpen
GitHub

typesafe-oracles

Evaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call.

github.comOpen

Jev patterns that pair with benchmarks and research

Browse other categories