Factuality & Hallucination Scorer
Score + Noul pattern for checking LLM answers against retrieved source documents.
Architecture overview
Compares generated answers to RAG chunks before user delivery. Emits a consistency Score and a contradiction Noul so application code can re-retrieve or escalate.
Textual LLM verifiers add latency and subjective prose that is hard to threshold.
Direct numeric Score / Noul outputs fit automated gates; cite official latency/pricing rather than inventing site benchmarks.
Official API schema (field names from docs.typesafe.ai)
Source: TypeSafe HTTP API reference. Question IDs below this table are directory example compositions — you choose them.
| Field | Type | Range | Notes |
|---|---|---|---|
| state | string | object | array | Plain text or structured JSON | Content to evaluate. Shared across all questions in one request. |
| model | string | e.g. "jev-latest" | Optional; defaults to TypeSafe flagship alias when omitted in SDKs. |
| questions.<id>.type | "choice" | "score" | "noul" | Exactly one of three primitives | Question ID is chosen by you; answers return under the same keys. |
| questions.<id>.instructions | string | Natural-language judgment | The actual question sent for inference (IDs are not sent to the model). |
| questions.<id>.criteria (choice) | Record<option, string | null> | 1–255 options | Map of option key → rubric description. |
| questions.<id>.criteria (score) | string[] | ≥ 2 ordered levels | Ordered level descriptions; score is a weighted position along them. |
| questions.<id>.criteria (noul) | { true?: string; false?: string } | Optional | Optional clarification of yes/no meanings. Answer field is noul ∈ [0, 1]. |
| answers.<id> (choice) | { type, choice, probabilities, confidence } | choice ∈ criteria keys; confidence ∈ [0, 1] | Probabilities sum to 1 across options. |
| answers.<id> (score) | { type, score, legend, probabilities, confidence } | score may fall between levels | legend maps level index → description. |
| answers.<id> (noul) | { type, noul } | noul ∈ [0, 1] | Probability that the answer is yes. No separate confidence field. |
| usage | { input_tokens, output_tokens } | Non-negative integers | Token accounting for the request. |
Example composition for this Jev AI tools workflow
Sample state and question keys are illustrative compositions for this directory page — not a separate official product API. Primitives remain Choice / Score / Noul.
Source: "Return policy allows refunds within 30 days. Shipping paid by customer." Answer: "You have 60 days with free return shipping."
Runnable call examples
Endpoint: https://api.typesafe.ai/v1/systemone (official). Requires your own TYPESAFE_API_KEY.
curl
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"state":{"source":"SOURCE","answer":"ANSWER"},"model":"jev-latest","questions":{"fact_score":{"type":"score","instructions":"Consistency with sources","criteria":["Severe hallucination","Partial","Mostly consistent","Strictly factual"]},"has_conflict":{"type":"noul","instructions":"Answer contradicts source?"}}}'TypeScript
import { noul, score, TypeSafeClient } from "@typesafe-ai/sdk";
const client = new TypeSafeClient();
const fact = await client.systemOne({
state: { source: docs, answer: llmOut },
model: "jev-latest",
questions: {
fact_score: score("Consistency with ground-truth documents", [
"Severe hallucination",
"Partial match",
"Mostly consistent",
"Strictly factual",
]),
has_conflict: noul("Does `answer` contradict `source`?"),
},
});Python
from typesafe_sdk import Noul, Score, TypeSafeClient
with TypeSafeClient() as client:
fact = client.system_one(
state={"source": docs, "answer": llm_out},
model="jev-latest",
questions={
"fact_score": Score(
instructions="Consistency with ground-truth documents",
criteria=[
"Severe hallucination",
"Partial match",
"Mostly consistent",
"Strictly factual",
],
),
"has_conflict": Noul(
instructions="Does `answer` contradict `source`?",
),
},
)Latency & cost (source attribution)
Official end-to-end latency range: ~70–500ms; many calls land near ~100ms from US West Coast (source: typesafe.ai / TypeSafe public materials, 2026-09). Official list price: $0.042 / 1M input tokens; output tokens free (source: typesafe.ai, as of 2026-09). Card latency figures are illustrative compositions within that published range — not independent lab measurements by this directory.
- Card illustration on this page: ~88ms (illustrative, within official range — not a lab run by jevaitools.com).
- Official list price: $0.042 / 1M input tokens; output tokens free (source: typesafe.ai, as of 2026-09).
This directory has not published an independent measurement script for this page. To measure yourself: call the official endpoint with your key, record wall-clock p50/p95 andusage.input_tokens, and keep the date of the run.
Comparison with generative LLMs on the same decision task
TypeSafe publishes workflow evaluations where Jev is compared with frontier LLMs on accuracy, cost, and latency (company materials, 2026). Those multipliers are vendor-reported ceilings, not results measured by this directory. Run the same questions through an LLM structured-output adapter on your labeled set before choosing a stack.
Suitable for
- Post-generation factuality gates in RAG
- Sampling audits on production answers
- Triggering re-retrieval when contradiction Noul is high
Not suitable for
- Producing corrected answers
- Open-web fact checking without retrieved sources in state
- Legal certification of truthfulness
Common failure modes
- Sources omitted from state → false confidence
- Score levels that mix “style” with “facts”
- Expecting citation strings from a System One model
Other Jev AI Tools
This site is an independent third-party directory and is not affiliated with, endorsed by, or operated by TypeSafe AI.