Model Overview · Internal Reference

Jev 1.13 — Classification and decision model

Jev reads language as a language model does, but does not generate text. It returns typed answers with calibrated probabilities and confidence values that software acts on directly.

01 — Overview

Reads, classifies, does not write

02 — Decision primitives

Noul, choice, score

Each question supplies type and instructions, plus its own criteria shape.

Noul

Boolean · probability 0–1

Returns a calibrated probability that a statement is true. For gates and guardrails.

{
  "type": "noul",
  "noul": 0.29
}

Choice

Single selection · up to 255 options

Returns the selected key, a probability per option, and a confidence. For routing and triage.

{
  "type": "choice",
  "choice": "billing",
  "confidence": 1.0
}

Score

Ordinal · 2–10 ordered levels

Returns a fractional score, a legend, probabilities per level, and a confidence. For severity and lead scoring.

{
  "type": "score",
  "score": 3.0,
  "confidence": 1.0
}
03 — Comparison

Against a standard language model

AttributeStandard LLMJev 1.13
Primary functionText generationDecision resolution
Output formUnconstrained prose or codeFixed JSON schema on every call
Output costBilled per generated tokenFree
Malformed outputPossible; needs defensive parsingImpossible by construction
ConfidenceNot suppliedCalibrated value plus distribution
LatencySeconds; grows with outputMilliseconds; flat
Per extra questionLatency risesFree; runs in parallel
Suited toDrafting, explanation, coding, dialogueRouting, triage, scoring, gating, moderation
04 — Use cases

Candidate applications

Worked examples, presented for assessment. None has been validated in production.

Audience sentiment at scale

score + choice · high-volume comment streams

A broadcaster wants the overall audience reaction to a live stream — supportive, critical, concerned — from 20,000 comments. Handing them to a language model would exhaust the token budget. Jev reads them in batches and returns a sentiment distribution per batch, which code tallies into one picture of the audience.

  • Comments are scored in batches, with a score question per batch and the results summed in code
  • The full distribution per comment is retained, so a rare strong reaction is not averaged away
  • Cost tracks input tokens alone, making full-corpus analysis feasible where per-comment generation is not
  • A low-confidence batch is flagged rather than folded silently into the total

Filtering what a search returns

noul × 4 · pre-generation filtering

A search assistant pulls matching documents and hands them to a language model to write an answer from. A document can match the keywords yet be outdated, contradict the question, or contain text aimed at the model. Each one is checked first.

  • Four yes/no questions per document: is it on topic, does it contain usable facts, does it contradict the question, is it trying to instruct the model
  • Thresholds apply in fixed order — security, conflict, relevance — and the first match decides
  • Documents that contradict go into a separate section of the prompt, so the answer can flag the disagreement
  • Re-tuning is a number change against stored answers, with no further calls
  • A filter, not a security boundary — the writing prompt still treats every document as untrusted

Sorting articles and comments into categories

choice · broad label when unsure

A publisher wants incoming articles and comments sorted into a fixed set of categories. Unclear cases look identical to clear ones, which is where sorting stalls. The confidence on the answer tells the two apart.

  • One question carrying every category, reliable to roughly 240 options
  • The decision reads confidence, not the top category's probability — a narrow lead and a scattered split are different situations at the same number
  • Below the threshold, the broader category is reported, worked out in code with no second request
  • Below that, the item goes to human review
  • In TypeSafe's published example, this fallback raised correct labels from 39 to 48 across 60 filings

Checking a name against earlier coverage

noul × N · advisory only

An editor asks whether an individual named in current coverage appears in earlier reporting on a similar offence. A string match cannot separate a genuine prior record from a namesake.

  • The state carries the two records being compared, not a full archive
  • Each identity assertion is a separate noul: surname, locality, vehicle or role, event-type consistency
  • All must clear threshold before a link is recorded; one failure routes to review
  • Probabilities are retained so a reviewer sees near-misses, not a bare rejection
  • Advisory only — it annotates a draft and never publishes. Low confidence does not mean no prior record exists
  • Thresholds are fitted per archive and are not transferable from the case above