Concept · updated 2026-10-07

Decision Models

Jev and the OpenAI Decisions API: AI that picks answers instead of writing them, and where it fits in the business. Every claim tagged.

Status: parked

V verified — safe P primary read — safest S secondary — hedge it ? unsourced — never on air

How to Use This

This is the plain-words layer. The proof layer is Decision Models. Every claim there has a tag and a source. Web version: https://concepts.rycolston.com/decision-models/

Part 1 explains Jev and the OpenAI Decisions API like you are ten. Part 2 is how they could help your business and your code.

Table of Contents

Part 1 — What these things are

  1. Two New Robots That Don't Talk

  2. The Three Kinds of Questions

  3. "How Sure Are You?"

  4. Jev vs. OpenAI: The Scorecard

  5. Where They Mess Up

Part 2 — Your business

  1. What You Already Have

  2. Where They Fit in Your Code

  3. Where They Do Not Fit

  4. The Rules That Still Apply

  5. The Safe First Step


Part 1 — What these things are

Chapter 1: Two New Robots That Don't Talk

Rests on: evidence §1, §2.

Most AI you know is a talker. You ask Claude a question. It writes you an answer, word by word. That is great for essays. It is slow and messy when a computer program only needs "yes" or "no."

Think of a school test. One kind of test says "explain in a paragraph." Another kind says "circle A, B, C, or D." A talker AI is built for the paragraph. Your code only wants the circle.

So two companies built AIs that only circle answers.

  • Jev comes from a small company called TypeSafe AI. It came out on September 15, 2026. The boss, Diogo Almeida, used to work at OpenAI. He helped invent the trick that made ChatGPT good at chatting.
  • The Decisions API comes from OpenAI. It was shown on September 29, 2026. It opened to everyone as a test version ("beta") on October 6, 2026.

They are rivals. Jev is not part of OpenAI. OpenAI built its own version two weeks later.

Jev's name comes from William Stanley Jevons. He noticed that when coal got cheaper to use, people used more coal, not less. TypeSafe thinks the same will happen with AI. Make it cheap enough, and people will use it everywhere.

Chapter 2: The Three Kinds of Questions

Rests on: evidence §2.

Both robots answer only three kinds of questions. You write the question and the allowed answers ahead of time.

  1. Yes or no. "Does this text say stop?" The robot gives a number from 0 to 1. 0.95 means "almost surely yes." Jev calls this a Noul. OpenAI calls it a predicate.

  2. Pick one. "Which of these 10 rules fits this reply?" The robot picks one and tells you its odds for every option. Both call this a choice.

  3. Rate it. "How upset is this person? Calm, upset, or very upset?" The robot gives a spot on your scale. It can land between steps, like 1.4. Both call this a score.

You can ask many questions at once. The robot looks at the same text for each question, side by side. Asking 13 questions at once was about 12 times cheaper than asking them one by one in TypeSafe's own test.

The robot cannot write. It can't draft an email. It can't explain why. It can't quote the sentence that proves its answer. It only picks.

Chapter 3: "How Sure Are You?"

Rests on: evidence §2, §4.

Each answer comes with a "how sure" number called confidence. 1 means "I'm certain." 0 means "I'm guessing."

That number is the whole point. Your code can say:

  • Very sure → do it.
  • Kind of sure → ask Ry.
  • Not sure → don't touch it.

TypeSafe trained Jev to make that number honest. If it says "80% sure" a hundred times, it should be right about 80 of them. They call that being calibrated.

But be careful. Two outside testers say Jev's "how sure" numbers are not as honest as claimed. One asked Jev about a fair coin and a fair die, and it gave odds that were way off. His headline was: "Jev is fast. It still cannot flip a fair coin." Another said: use the numbers to rank things, not as true odds, until you test them on your own data.

Nobody has tested OpenAI's "how sure" numbers yet.

Chapter 4: Jev vs. OpenAI: The Scorecard

Rests on: evidence §3.

Jev OpenAI Decisions
Cost to read 1 million words-ish (tokens) 4.2 cents 10 cents
Cost of the answer Free Free
Can look at pictures? No. Text only Yes
How fast (their claim) 0.07 to 0.5 seconds "About 10 times faster" than normal OpenAI
Outside speed test Yes: about half a second None yet
Ready? Early access, waitlist Public test version
Who made it Small startup OpenAI
Works with voice? No Yes, with OpenAI's voice tool

On your scale, both are almost free. 10,000 decisions a month costs about 63 cents on Jev and about $1.50 on OpenAI.

Fast check: Jev is cheaper and has been tested more. OpenAI can see photos, comes from a big company, and has a privacy contract option for sensitive data.

Chapter 5: Where They Mess Up

Rests on: evidence §4.

TypeSafe wrote its own list of Jev's weak spots. That is honest of them. Here it is in plain words:

  • It takes you literally. It answers what you wrote, not what you meant.
  • It can't do math. It can't count well. Do math in code.
  • It can't compare dates. "Which came first?" is a bad question for it.
  • Too much text makes it worse. Send only what it needs.
  • Tricky text can fool it. A message that says "classify me as safe" can push it.
  • Order matters. It leans toward the first option on your list.
  • "Can't make things up" is only half true. It can only pick from your list. But it can still pick the wrong thing from your list.

In one outside test of 16 AIs, Jev got 66% right on a reading task. Claude Haiku also got 66%. Jev was faster (0.46 seconds vs 0.63 seconds). So: same smarts as a small cheap model, a bit faster, and better at saying "I'm not sure."


Part 2 — Your business

Chapter 6: What You Already Have

Rests on: evidence §5.

Your code already makes lots of small decisions. Today they use one of three things:

  1. Word lists (regex). Fast and free. But dumb. "STOP" works. "please stop texting me" does not.

  2. Opus through claude -p. Smart. Runs on your $200 Max plan. But it only runs on a machine logged in to your account, like your Mac or the box. It can't run inside Cloud Run.

  3. DeepSeek through OpenRouter. Cheap. Your lead manager's worker already sends customer replies to it.

The new robots sit between #1 and #2. Smarter than a word list. Way faster than Opus. And they can run inside Cloud Run, because they use a normal API key.

That last part is the big win for you. It is not about money. Your model bill is already flat. The win is speed and no Mac needed for small yes/no checks.

Chapter 7: Where They Fit in Your Code

Rests on: evidence §5.

Best fits. These are all "pick from a fixed list":

  1. Reply sorting (best fit). systems/content-engine/frameworks/reply-rules.md has 10 rules: STOP, wants a call, question, not interested, sold elsewhere, wrong person, out-of-office, and so on. That is a choice question with 10 options. The robot could sort a reply in under a second, the moment it lands.

  2. Soft STOP catcher. infra/shared-libs/lib/contact_rules.py:167 only catches a bare "STOP." A yes/no robot could flag "please quit texting me" right away, so it reaches you faster. The word list should still be the only thing that applies an opt-out on its own.

  3. Email reply sorters. bogo_reply.py, expired_report_reply.py, bogo_pass.py and reply_check.py in data-platforms/gmail-push/cloud_function/ all use word lists to sort replies into stop / wants-it / other. The robot could be a second opinion when the word list says "other." Example trap it could help with: "Can I stop by?" is not a STOP.

  4. Content grading. systems/media-engine/machine/stations.py has Opus grade every hook and title against a checklist each morning. Each check is a yes/no. No customer data is involved. This is the safest place to test.

  5. Topic tags. ramble_ledger_topics.py picks 1 to 4 tags from a closed list. That is a choice question.

  6. Email inbox triage. data-platforms/email-triage/ sorts email with rules. The robot could handle the ones no rule matches.

New ideas, only with OpenAI (it sees pictures):

  1. Listing photo check. Before a listing goes live, ask: Is this photo dark? Blurry? Is a person in it? Is the toilet seat up? Is it a duplicate? Each one is a yes/no. Nothing does this in your code today.

  2. Voice actions. OpenAI pairs this with its voice tool. A caller says something, and the robot picks the action from your list. This is a future idea, not a first step.

Chapter 8: Where They Do Not Fit

Rests on: evidence §4, §5.

Do not use them where the answer must quote proof. Your rules need a quote for every customer statement. That rules out:

  • Full FUB person reads (lib/person_read.py)
  • Deal status (/deal-status)
  • Payoff sweep (/payoff-sweep)
  • Lost-lead reviews
  • Writing any customer email or text

Also do not use them for math, money numbers, or dates. Do those in code.

Chapter 9: The Rules That Still Apply

Rests on: evidence §5.

Your own rules do not go away:

  • The 99% test. docs/BUILD-PATTERNS.md §1f says a cheaper model must agree with Opus 99% of the time, on a few hundred real records, before it takes a job. That applies here too.
  • A new paid vendor. Your "no paid API" rule is only about Anthropic. But these would be new paid accounts. A vendor swap needs a kill rule and a trial window first.
  • Customer words leaving the house. No rule covers this yet. OpenAI keeps abuse logs for 30 days unless you get a special "zero retention" setup. TypeSafe offers zero retention only to big customers. Your DeepSeek worker already sends customer text out, so this is not new, but it should be a written rule.
  • The robot never sends anything. Any decision that leads to a send still goes through send_or_log and your approval gates.

Chapter 10: The Safe First Step

Verdict (Ry, 2026-10-07): parked. None of these move the needle for the business. The slow parts are approvals, full reads and lead flow, not one-second yes/no checks, and at this volume faster and cheaper saves almost nothing. Revisit only if reply volume grows a lot or a real-time use shows up (for example, a voice line that must pick an action while the caller waits). The steps below are kept for that day.

Rests on: evidence §3, §5, §6.

Start where nothing can hurt a customer:

  1. Test on content grading first. No customer data. Run Jev and OpenAI side by side with Opus on the same hooks and titles. Cost: pennies.

  2. Run it in shadow. The robot answers, but nothing acts on its answer. Log both answers. Count how often they agree.

  3. Kill rule: under 99% agreement on 300 items → stop. Trial window: two weeks.

  4. Only then try reply sorting, also in shadow first.

What is still unknown: how fast OpenAI really is from your machines, whether OpenAI's "how sure" numbers are honest, and how long the Jev waitlist is.

Evidence file

Every claim, tagged

The prose above is only as good as the tag on each sentence here. Read the tag before you say the sentence out loud.

Narrative companion: Decision Models - Curriculum — the read-through written to be taught. Web: https://concepts.rycolston.com/decision-models/

Decision Models (Jev and the OpenAI Decisions API) — Evidence File

Every claim below carries a confidence tag. Read the tag before you say the sentence out loud.

Tag Means Safe to teach?
V Verified. Two independent primaries agree Yes
P Primary. I pulled the original document this session and read the sentence Yes — strongest tier
S Secondary. A helper or a reporter quoted it. I did not open the original Only with hedge
? Unsourced. Believed true, never checked No. Do not say on air.

Name check. Ry asked about "Jev and OpenAI's Decision API". The OpenAI product is officially the Decisions API (with an s). Jev is not an OpenAI product. It is a rival model from the startup TypeSafe AI. The two get named together because OpenAI launched its version two weeks after Jev. V

Coverage (2026-10-07). Both products are under one month old. Every speed number from OpenAI is its own claim. Jev has a few outside tests; OpenAI's endpoint has none yet. Expect this file to age fast.


1. What they are

Jev P

TypeSafe launch post, 2026-09-15, by founder Diogo Almeida: "TypeSafe AI is releasing our first System One Model: a new class of frontier models built to make fast, structured decisions that software can use directly." Jev is "available today in early access."

OpenAI Decisions API P

OpenAI guide: "The Decisions API evaluates text, images, or both and returns typed answers about 10x faster than the Responses API. Get the probability that a condition is true, a choice from a fixed set, or a score against a rubric."

  • Status line on the same page: "The Decisions API is in public beta, and we expect to GA in the coming weeks." Model: gpt-6-luna only. Endpoint: POST /v1/decisions.
  • Source: developers.openai.com/api/docs/guides/decisions.md, fetched 2026-10-07.
  • OpenAI changelog, Oct 6, 2026: "Released the Decisions API in beta with gpt-6-luna." P (changelog.md)
  • Announced at DevDay on 2026-09-29 in "limited preview". S — NYU Shanghai RITS write-up (link). OpenAI's own DevDay recap page returned HTTP 403 to curl.

How the two relate P

NYU Shanghai RITS, 2026-09-30: "It arrives two weeks after TypeSafe AI launched Jev, a model built around the same idea." No source says Jev runs inside the OpenAI product or that the firms are partners.

Who built Jev P / S

  • TypeSafe docs: RLHF "was used to train InstructGPT and ChatGPT and was co-invented by Diogo Almeida, cofounder of TypeSafe." P (AI primer)
  • Launch post, Almeida: "At OpenAI, I helped build the methods that made language models useful at following instructions." P
  • $40M seed at a $200M valuation; co-founders Erik Gafni and Sasha Sheng. S — Dealroom (403 to curl).
  • Talks on a $1B+ round near a $10B valuation. S — Spanish blog citing The Information. Do not air without the original.
  • Named after William Stanley Jevons: "Every order of magnitude drop in the cost of intelligence unlocks orders of magnitude more use cases." P (launch post FAQ)

2. How they work (the mechanism)

The big idea: answers, not essays P

TypeSafe intro: "Large language models (LLMs) are designed to produce text for humans to read. When you need a model to make a judgment that your code will consume, that creates a mismatch." Jev "returns typed decisions and probabilities rather than generated text." (introduction)

Three question types — the same in both products V

Jev name OpenAI name Asks Returns
Noul predicate Is this true? one probability, 0 to 1
Choice choice Which one of these? the pick + a probability per option + confidence
Score score Where on this ordered scale? a weighted average level + probabilities + confidence

Both docs fetched this session. OpenAI: "score: the probability-weighted average of the level indices." TypeSafe: "The probability-weighted answer across the levels; can land between levels."

Many questions, one call P

  • TypeSafe: "Every question is evaluated in parallel and in isolation against the same state in one go. Adding questions barely changes the response time."
  • OpenAI: "Put independent questions in the same questions array to evaluate shared input." And: "For decisions that depend on an earlier answer, send separate requests."
  • TypeSafe cookbook: batching a 13-question briefing into one call was "12.2x cheaper and 10.0x faster with no change in answers." P (vendor's own test)

Parallel sampling, not word-by-word P

Launch post table: LLMs sample "Sequential. Generates one token at a time"; Jev is "Parallel. Generates all outputs in a single query."

Training method: RLCD P

"Reinforcement learning for calibrated decisions trains TypeSafe to return decisions and calibrated probabilities instead of generated text." Calibrated means: "Outcomes assigned a probability of 0.8 should occur about 80% of the time." The page adds: "These rates describe groups of predictions, not a guarantee about any single answer."

Confidence P

  • TypeSafe: confidence "is 1 when all the probability is on one outcome and 0 when the probability is spread evenly." Suggested use: high = act, medium = confirm or review, low = route to a human. "Thresholds scale with risk."
  • OpenAI: "Use labeled examples from your application to set thresholds for routing, filtering, or review. Choose thresholds based on the cost of false positives and false negatives."

"Can't hallucinate" — true only in a narrow sense P + V

  • Launch post: Jev "can't hallucinate"; "The model never makes type errors."
  • Pushback, Hacker News (jacobgold): "Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value." P (fetched)
  • Armin Ronacher via The Next Web: "it delegates the hallucination problem a little bit to the user." P (fetched)
  • Teach it as: it can only pick from your list. It can still pick the wrong item.

3. Specs side by side

Jev 1.13 (TypeSafe) OpenAI Decisions (gpt-6-luna)
Endpoint POST https://api.typesafe.ai/v1/systemone P POST https://api.openai.com/v1/decisions P
Input price $0.042 per 1M tokens P $0.10 per 1M tokens P
Output price Free P None charged: "You pay only for input tokens" P
Input types Text only. "Images, audio, and video are not supported (yet)." P Text and images. Images must be inline base64; "At most 128 images are allowed across the request." P
Context "64k tokens per request; 32k tokens for state plus the longest question" P Luna model page: "Maximum input tokens: 922,000" P (not stated for /v1/decisions itself)
Max options "a maximum of 255 options per Choice"; Score "up to 10" levels P Not published P (absence checked in the reference page)
Rate limits "100K tokens per second / 80 requests per second", and "can change without notice" P No Decisions-specific limit published. Luna Build tier: 5,000 RPM, 2,000,000 TPM P
Speed claim "70ms-500ms" end to end P (vendor) "about 10x faster than the Responses API" P (vendor); 150 ms vs 1.6 s on the DevDay slide S
Status Early access; waitlist P Public beta; GA "in the coming weeks" P
Refusals No refusal type in the API reference P Per question: "Other questions in the same request can still receive answers." P
Batch API Not offered ? Not offered — /v1/decisions is absent from the Batch endpoint list S
Data "Jev is not trained on customer requests or responses." ZDR for enterprise only. P Not used for training; abuse logs kept "30 days"; ZDR "Yes, see below for limitations"; HIPAA eligible under a BAA P
Language "English is the primary training language" P Not stated
Voice None Pairs with GPT-Live: "Combine the Live API and Decisions API to select actions and report their results through voice." P

Sources: Jev models page, Jev API reference, OpenAI guide, OpenAI reference, OpenAI your-data, gpt-6-luna. All fetched 2026-10-07.

Correction to outside articles. Reviews written before 2026-10-06 (OrcaRouter, eesel) say OpenAI "has published no rate for the Decisions API." That is now out of date: the live guide prices it at $0.10 per 1M input tokens. P

Cost math on Ry's scale P prices, my arithmetic

One reply plus the rules list is about 1,500 tokens (my estimate ?). 10,000 decisions a month = 15M tokens.

  • Jev: 15 × $0.042 = $0.63 a month.
  • OpenAI: 15 × $0.10 = $1.50 a month. At this volume, the dollar cost is close to zero on both. The money question is not the bill. It is the new vendor, and where customer words go.

4. Where they are weak

TypeSafe's own list of failure modes P

From Jev 1.13 jaggedness, last reviewed 2026-10-02:

  • "jev-1.13 answers the question you wrote, not the one you meant."
  • "Jev is not a calculator." "jev-1.13 does not count reliably."
  • "jev-1.13 reads dates as text, not as ordered quantities."
  • "Accuracy falls as the state grows with content unrelated to the decision."
  • Prompt injection: "Content written to adversarially steer the model … can move the answer."
  • Option order: "jev-1.13 leans toward the option that comes first."
  • No writing: "jev-1.13 is not trained to generate text."

Calibration is disputed P

  • Alex Molas, 2026-09-23: "The same model can be calibrated on one dataset but not on another." Advice: "treat Jev's outputs as good scores (they rank examples well) rather than good probabilities." (link)
  • Maximum Effort, 2026-10-02: "If Jev were to completely punt on the answer, and spread the probability evenly over all possible choices, it would score a mean TV of 0.546. But Jev scores 0.518." Headline: "Jev is fast. It still cannot flip a fair coin." (link)

Accuracy is middling; speed is real P

wotai.co independent test, 2026-09-18, 16 models on 150 passages: Jev and Claude Haiku 4.5 both scored 66.0%; Jev median 455 ms vs Haiku 631 ms; Jev flagged "uncertainty on 34.7% of rows". The tester: "Several models beat it on accuracy and one beat it on calibration." (link)

  • TypeSafe's own benchmark: "Jev averages 67.8% agreement with the reference answers." S (DataCamp, 403 to curl)
  • TypeSafe says its benchmark reference answers lean toward OpenAI and Anthropic models. P (launch post)

OpenAI's endpoint has no outside test yet S

OrcaRouter, about 2026-10-01: "there is no independent measurement of it in the two days since launch." Nothing newer found.

Adoption signals P / S

  • The Next Web: "Within 24 hours of arriving on Vercel's AI Gateway, Jev reached nearly 13% of the company's paid teams." P (fetched)
  • Jev on OpenRouter: P50 0.21 s, P95 0.34 s as of 2026-10-01. S (Firecrawl citing OpenRouter)

5. Ry's codebase — where decisions are made today

Mapped by a read-only Opus helper on ~/rylobasic; spot-checked lines marked P.

# Where Decision Today Fit
1 infra/shared-libs/lib/contact_rules.py:167 _STOP_RE Is this text a STOP? Regex; whole message must be the keyword. "please stop texting me" goes to Ry. P yes/no, as a second look only
2 data-platforms/gmail-push/cloud_function/bogo_reply.py BOGO reply: stop / bogo / other Regex S choice
3 .../expired_report_reply.py:107 classify Report-email reply: stop / report / other; auto-reply? Regex + headers P choice + yes/no
4 .../bogo_pass.py BOGO pass text: stop / confirm / other Word sets S choice
5 .../reply_check.py is_plain_yes Plain "yes" to the morning text ("Are you against me sending one?" → "No" means yes) P Word lists yes/no
6 systems/content-engine/frameworks/reply-rules.md + paperclip/leadmgr Which of 10 reply rules applies DeepSeek worker reads, Opus judges P choice (10) — cleanest fit
7 systems/media-engine/machine/stations.py:421 Grade hooks/titles against checklists Opus claude -p, daily 04:50 CT S yes/no per check
8 systems/media-engine/machine/ramble_ledger_topics.py 1–4 topic tags from a closed list Opus S choice
9 data-platforms/lifelog/judge_loops.py Open loop: "open" or "noise" Opus hourly S yes/no (private texts)
10 data-platforms/email-triage/cloud_function/rules.py Email category Rules table S choice, as fallback when no rule matches
11 data-platforms/fub/hygiene/referee.py FUB stage or ESCALATE Opus claude -p daily S choice — but it gates a write switch
— Person reads, deal-status, payoff sweep, lost-lead review Need a quoted piece of evidence Opus Not a fit. These products cannot quote.

Ry's rules that apply P

  • docs/RAILS.md:258: "Every Claude call runs on Ry's Max plan through Claude Code… No paid Anthropic API key anywhere." It also says "Cloud Run cannot hold Max."
  • docs/BUILD-PATTERNS.md:183 (§1f): a cheaper model must reach the "same conclusions 99% of the time" before it takes a job.
  • Global rules: "Vendor swaps name kill criteria and a trial window up front."
  • Gap found: no written rule covers sending customer text to a new outside model vendor. Customer text already goes to OpenRouter (DeepSeek worker). S (helper report)
  • OpenRouter account check, 2026-10-07: Zero Data Retention is required for all other models, training toggles are off, and the day's DeepSeek calls were served by DigitalOcean, Together and inference.net. P (Privacy and Logs pages, viewed in Chrome)

6. Verdict

Ry, 2026-10-07: "i dont think any of these things move the needle for me or the busines". Parked. No pilot, no vendor account.

7. What is still unknown

  • Real Decisions API latency on Ry's network. Only OpenAI's "10x" claim exists. ?
  • Whether OpenAI's confidence is calibrated at all. No test found. ?
  • Jev access for a new account today (waitlist vs instant). ?
  • Whether TypeSafe's price is subsidized. TypeSafe itself: "We can't prove it isn't subsidized." P
  • OpenAI per-call limits (max questions, max choices). Not published. P absence