103 of 561 catalog entries
Benchmarks & research — built on Jev
Evaluation material — measured latency, task accuracy, and reproductions of published claims. Treat these as third-party work with their own methodology, and read the setup notes before quoting any number.
This directory does not publish its own benchmark claims. Figures quoted elsewhere on this site come from TypeSafe public materials and are labelled as such.
JevK5
Apache-2.0 open-weight Jev alternative: noul / choice / score with probabilities, TypeSafe-style /v1/systemone, weights on Hugging Face.
SemIf
Semantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.
jevlike
Train a small model that chooses among a changing list of text options, one probability per option in a single pass. Includes Doom, chess, and Wikispeedia demos.
jev-visual
An educational Jev-like visual inference experiment on Apple Silicon: shared context, direct candidate scoring, and local visual demos.
openjev
Open, Jev-compatible System One decision server on DiffusionGemma.
jev-eval-agent
Personal-assistant agent built on Vercel's eve with 100 mocked tools, measuring how many steps it takes when Jev picks the tool versus the LLM.
decider
One-pass typed decisions with calibrated probabilities (System One style model), fine-tuned from Qwen3.5-2B.
reflex
A small open decision model: state + typed questions -> calibrated probabilities. A Jev / System One re-creation on Qwen3.5.
jevfire
JEV-inspired parallel decisions for CUDA LLMs. One context, many decisions. vLLM API, game-agent examples, and reproducible benchmarks.
jevmlx
Jev-style parallel constrained decisions for any MLX model on Apple Silicon. Typed, schema-valid JSON in one forward pass.
LitJev
A reproduction of Jev that turns any Qwen model into a fast decision model, serving the same /v1/systemone schema (Choice, Score, Noul) with no training and no generated answer text.
open-jev (daseinlabs)
One-pass option scoring with a local Gemma 3 4B on Apple silicon via MLX, inspired by jevlike, with a Doom demo.
typesafe-ai-benchmark
This is a LLM Gateway that mimics typesafe ai structured output. Like an imposter Jev.
jevgpt
A chatbot built on a model that cannot generate text (TypeSafe AI's Jev, driven autoregressively).
mini-jev
Mini-Jev: what a Jev-style typed-decision interface looks like on a frozen Qwen3-4B — read the option letter's logits instead of generating JSON. Preregistered experiment, results, teaching bench.
Verdict-open-jev
Non-autoregressive decision engine on ModernBERT (151M) with calibrated uncertainty (RLCD), TypeSafe AI Jev benchmark audit, and in-browser WebGPU playground.
von
The open-source System One decision model. Sub-15ms, non-autoregressive, local drop-in alternative to TypeSafe Jev.
jev-capability-atlas
Independent, evidence-based map of when TypeSafe's Jev actually holds up vs. breaks down — real API-call receipts, not a leaderboard. 中文為主的雙語 repo。
jev-column-race
Jev vs Gemini 3.8 Flash: labelling 1,000 app reviews, 4.1× faster and 7× cheaper.
jev-on-a-laptop
Unofficial study: Jev-style parallel typed decisions on stock 1.5B-8B models on an Apple Silicon laptop. Benchmarks, research notes, and a Hugging Face Space demo.
open-jev (JoshuaSP)
Typed JSON inference with DiffusionGemma, with Every and Jev benchmark results.
jev_local
Replicating Jev with a local LLM.
jevbetter
A stronger one-pass scorer over a variable list of text options: hashed n-gram encoder, rival-aware attention, gated head, temperature scaling, benchmarked against jevlike.
openjev (zhihz)
Local bilingual probability decisions from context, questions, and candidate answers. Independent research preview inspired by TypeSafe Jev.
jev-benchmarks
Probability-aware evaluation for typed decision models: calibration, selective risk, latency, and reproducible benchmarks.
system-one
Batched single-token choice inference for open language models, compatible with TypeSafe.
Jev_apps
看看 Jev 能做什么:用中英文讲清热门应用、工作原理和各自优缺点。Explore Jev apps with plain-language examples, explanations, and comparisons.
open-alternative-jev
Open alternative to Jev: typed, calibrated decisions from any open-weights LLM in one forward pass (HF + vLLM), with benchmarks.
openvons
Openvons (open-Jev): 有限選択肢に確率で答える判断層 — テキスト / 画像 / 日本語音声コマンド.
TypeAR
Type-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev.
jev_stock
An experimental JEV-powered framework for forecasting short-term stock price direction from structured market data.
system-one-open
Open replica of TypeSafe's Jev: typed calibrated decisions in one forward pass, on Gemma 4 E2B / Gemma 3 270M (Modal).
jevcal
Stop guessing confidence thresholds: calibrate, threshold, and drift-check typed decision models (TypeSafe Jev) against an LLM teacher.
typesafe-local
Inspired by TypeSafe Ai, Ask a local LLM typed questions, get calibrated probabilities instead of text. Structured output without generation or parsing. MLX / Apple Silicon.
jev-korean-benchmark
Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence.
jev-lm
A word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval.
jev.nu
Nushell module for the TypeSafe System One API: typed decisions with calibrated probabilities.
jevify
An agent skill to discover TypeSafe Jev opportunities, design typed questions, and learn from recent community experiments.
LegalForecastBench
LegalForecast-MTD benchmark alpha and official evaluation workflows.
daf-jev
Daf-jev: composable Python toolkit for TypeSafe's Jev (System One) decision API — question builders, confidence gates, evaluator, calibration, CLI, MCP server, agent skill.
jev-align
Calibrated alignment verifier for LLM responses and agent plans — powered by Jev.
jev-search-rerank-eval
Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.
jev-behavior-study
Independent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification.
jev-chat
A chatbot from typed Jev decisions: hierarchical speculative decoding over System One probabilities.
jev-curate
High-throughput synthetic & pretraining dataset sifter powered by TypeSafe AI Jev (api.typesafe.ai). Stream, filter, and score Parquet & JSONL datasets at 1,500+ rows/sec using System One typed decisions (Choice, Score, Noul).
jev-harness
Decision harness for TypeSafe Jev — confidence gates, shadow mode, recipes, and evals. Claude CLI 48.9s → Jev 1.3s on the same row-filter job.
jev-little-airways
A show-and-tell capability study for Jev, TypeSafe's System One decision model.
calibre
Calibration and confidence-based routing measured on Banking77: 80.2% accuracy at $0.103 per 500 decisions.
claude-jev
Claude Code plugin that scores review findings, debug hypotheses and design options with TypeSafe's Jev — calibrated probabilities instead of one more opinion.
jev-benchmark
Benchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs.
jev-for-engineers
Eight minimal working examples of TypeSafe's Jev (a System One model) applied to mechanical and electrical engineering: CAD/CAE/CAM routing, FEM result triage, DFM screening, BOM alignment, hallucination-proof extraction. Zero dependencies.
jev-gate
Not every coding task needs your best model. Experimental Jev-powered model routing for Claude Code — V3 prototype runs today, V4 routes at the task boundary.
jev-mode
I kept watching coding agents burn context on decisions that aren't hard - triage 400 tickets, tag 600 files, route to one of six teams. jev-mode moves those verdicts to a typed-judgment model. I A/B'd it: 78% fewer tokens, 16x less work-attributable input, accuracy 96.1% vs 93.7%. Python, no deps, MIT.
jev-research-eval
Reproducible Jev Ultrafast research-browser eval harness + field note (QC’d cases, suite runner, report generator). Not investment advice.
jev-sec-bench
Blind security benchmarks for Jev, TypeSafe's System One model: prompt injection and vulnerable code detection, built on jev-go.
mcts-agent
Discriminative Monte Carlo Tree Search using TypeSafe Jev System One Primitives and Gemini.
new-api-plugin-typesafe
TypeSafe AI System One (Jev) task plugin for QuantumNous/new-api — native /v1/systemone, synchronous evaluation, token billing.
research_desk
TypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers.
siftr
Fast, cheap judgment for AI coding agents: semantic search, focused reads and list picking in ~2s. CLI + MCP server on TypeSafe Jev. Benchmarked on SWE-bench.
system-one-gemma
Open-source Jev-style System One decision model. Gemma 3 270M with a scoring head — fast, calibrated decisions in a single forward pass. No text generation. Inspired by TypeSafe.ai's Jev.
tenbin
MCP server and agent skill for the TypeSafe AI System One API (Jev): decompose a judgment into Choice / Score / Noul questions, lint them, measure on labelled data, and put calibrated thresholds in code.
trade-jev
Backtest Jev (TypeSafe) as a BUY/SELL/HOLD trader on NQ L10 order-book data.
zerosweep
Autonomous System-One Triage Engine & Benchmark powered by TypeSafe AI (Jev). 75ms inference, $0 output tokens, and RLCD epistemic safety gates.
decisionbridge
A Jev-inspired decision interface for existing LLMs. Explicit choices, scores, calibration, and review thresholds.
jev-agent-failure-benchmark
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
jev-as-a-judge
Using Jev as an evaluator.
jev-carryforward
What your last session knew, scored against what this one is doing. MCP server: a per-project ledger written as things happen, recalled per task with TypeSafe's Jev evaluation model via Vercel AI Gateway.
jev-chess
Chess moves, evaluations, persona opponents, and game classification with TypeSafe AI System One.
jev-exploration
Jev (TypeSafe) exploratory thread: claim audit, live demos, and runnable code.
jev-freeform
An observable raw-character chat experiment powered entirely by TypeSafe Jev Choice.
jev-phishing-bench
Jev (TypeSafe) vs Claude Haiku 4.5 on 2 000 phishing emails: accuracy, calibration, latency, cost. Reproducible benchmark.
jev-rerank-bench
Can a decision model beat dedicated rerankers? TypeSafe Jev vs Cohere Rerank 4 vs ZeroEntropy zerank-2 vs a chat-model baseline: 14 datasets, every raw API response, bootstrap ranges on every gap.
jev-synergy-screening
Jev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels.
kyotsu-ai-bench
AI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard).
modelsystem
Curated catalog of System One / Decision Models — contributions for modelsystem.one.
openjev-experiments
Experiments with openjev, an open Jev-style option-logit runner, on local models.
padflow-jev-evals
Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.
qwen-rlcd
Jev-style calibrated decision model (Choice/Score/Noul) on Qwen3.5-0.8B.
RISC-jeV
I tortured Jev into being a RISC-V CPU.
typesafe-vs-deepseek
TypeSafe (Jev) vs DeepSeek-flash: side-by-side speed/token/cost/accuracy comparison across invoice extraction, email classification, and reranking.
FinancialPredictionJev
Using Jev to test how well it predicts financial markets(just like most llms as of september 2026, it doesnt do that good).
jev-alpha-bench
Does Jev predict stock returns from news? It reads the news well; there is no tradeable alpha. Three arms separate reading from recall.
jev-anotacao-sentencas
Jev (TypeSafe) vs. Gemini 3.8 Flash vs. GPT-5.6 Luna na anotação estruturada de sentenças do TJSP: qualidade, tempo e custo.
jev-deferred-crispification
Position paper: the Hidden-Markov and fuzzy primitives missing from TypeSafe AI's Jev and System-One decision models. Two lemmas, one principle (Deferred Crispification), one architecture (BSF-S1).
jev-finance-benchmark
Typesafe.ai model jev finance benchmark.
jev-headline-bench
Can Jev pick the winner of a real headline A/B test? 64.5% across 10,984 Upworthy randomized experiments, 74.7% when the difference was decisive.
jev-jp-address
Jev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証.
jev-lab
TypeScript experiments, evaluations, and latency benchmarks for TypeSafe's Jev model.
jev-pick-and-place-study
A small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.
jev-playground (hegargarcia)
Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.
jev-report
发明 RLHF 的人,这次做了个不会说话的模型:Jev 独立研究报告。52 页 PDF + 50 条中文实测复现包 + 143 条可回溯数据表.
jev-routing-experiment
Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena.
jev-secret-detection
Measures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets.
jev-shadcn-lint-eval
A small second eval for shadcn-ui/lint that uses TypeSafe's Jev to judge the linter's own output.
jev-spam-eval
Zero-shot spam filtering with TypeSafe Jev Noul questions, compared with TF-IDF baselines.
jev-trace-classifier
Application of TypeSafe Jev (noul judgment primitive) on the collusion.wiki corpus: agent vs human page authorship, head-to-head vs local Qwen3.8-Flash-Next.
misereru-slide-jev
Ongoing Japanese research deck on Jev and System One models, maintained as Markdown slides.
Parallel Constrained Decoding (Qwen2.5-1B-RLCD)
Hugging Face Space exploring open-source parallel constrained decoding as an alternative to Jev.
PocketJev
On-device iPhone visual decision tool using MLX and Qwen3-VL direct option logits.
shade-arena-jev-monitor
Evaluating TypeSafe's Jev as a fast monitor and action gate for agent sabotage in SHADE-Arena, compared with Gemini 2.5 Flash/Pro.
system-one-adapter-rust
Rust port of TypeSafe system-one-adapter (LLM-backed system_one evaluations).
thaiexam-jev-charts
Charts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models.
Typesafe_chess_eval
An evaluation of typesafe AI chess. As it turns out, the AI isn't doing really well even though chess is not a particularly open-ended game. Still, it's only a prototype and this probably wasn't optimzied for games.
typesafe-oracles
Evaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call.