How to Use This
This is the plain-words layer. The proof layer is Decision Models. Every claim there has a tag and a source. Web version: https://concepts.rycolston.com/decision-models/
Part 1 explains Jev and the OpenAI Decisions API like you are ten. Part 2 is how they could help your business and your code.
Table of Contents
Part 1 — What these things are
-
Two New Robots That Don't Talk
-
The Three Kinds of Questions
-
"How Sure Are You?"
-
Jev vs. OpenAI: The Scorecard
-
Where They Mess Up
Part 2 — Your business
-
What You Already Have
-
Where They Fit in Your Code
-
Where They Do Not Fit
-
The Rules That Still Apply
-
The Safe First Step
Part 1 — What these things are
Chapter 1: Two New Robots That Don't Talk
Rests on: evidence §1, §2.
Most AI you know is a talker. You ask Claude a question. It writes you an answer, word by word. That is great for essays. It is slow and messy when a computer program only needs "yes" or "no."
Think of a school test. One kind of test says "explain in a paragraph." Another kind says "circle A, B, C, or D." A talker AI is built for the paragraph. Your code only wants the circle.
So two companies built AIs that only circle answers.
- Jev comes from a small company called TypeSafe AI. It came out on September 15, 2026. The boss, Diogo Almeida, used to work at OpenAI. He helped invent the trick that made ChatGPT good at chatting.
- The Decisions API comes from OpenAI. It was shown on September 29, 2026. It opened to everyone as a test version ("beta") on October 6, 2026.
They are rivals. Jev is not part of OpenAI. OpenAI built its own version two weeks later.
Jev's name comes from William Stanley Jevons. He noticed that when coal got cheaper to use, people used more coal, not less. TypeSafe thinks the same will happen with AI. Make it cheap enough, and people will use it everywhere.
Chapter 2: The Three Kinds of Questions
Rests on: evidence §2.
Both robots answer only three kinds of questions. You write the question and the allowed answers ahead of time.
-
Yes or no. "Does this text say stop?" The robot gives a number from 0 to 1. 0.95 means "almost surely yes." Jev calls this a Noul. OpenAI calls it a predicate.
-
Pick one. "Which of these 10 rules fits this reply?" The robot picks one and tells you its odds for every option. Both call this a choice.
-
Rate it. "How upset is this person? Calm, upset, or very upset?" The robot gives a spot on your scale. It can land between steps, like 1.4. Both call this a score.
You can ask many questions at once. The robot looks at the same text for each question, side by side. Asking 13 questions at once was about 12 times cheaper than asking them one by one in TypeSafe's own test.
The robot cannot write. It can't draft an email. It can't explain why. It can't quote the sentence that proves its answer. It only picks.
Chapter 3: "How Sure Are You?"
Rests on: evidence §2, §4.
Each answer comes with a "how sure" number called confidence. 1 means "I'm certain." 0 means "I'm guessing."
That number is the whole point. Your code can say:
- Very sure → do it.
- Kind of sure → ask Ry.
- Not sure → don't touch it.
TypeSafe trained Jev to make that number honest. If it says "80% sure" a hundred times, it should be right about 80 of them. They call that being calibrated.
But be careful. Two outside testers say Jev's "how sure" numbers are not as honest as claimed. One asked Jev about a fair coin and a fair die, and it gave odds that were way off. His headline was: "Jev is fast. It still cannot flip a fair coin." Another said: use the numbers to rank things, not as true odds, until you test them on your own data.
Nobody has tested OpenAI's "how sure" numbers yet.
Chapter 4: Jev vs. OpenAI: The Scorecard
Rests on: evidence §3.
| Jev | OpenAI Decisions | |
|---|---|---|
| Cost to read 1 million words-ish (tokens) | 4.2 cents | 10 cents |
| Cost of the answer | Free | Free |
| Can look at pictures? | No. Text only | Yes |
| How fast (their claim) | 0.07 to 0.5 seconds | "About 10 times faster" than normal OpenAI |
| Outside speed test | Yes: about half a second | None yet |
| Ready? | Early access, waitlist | Public test version |
| Who made it | Small startup | OpenAI |
| Works with voice? | No | Yes, with OpenAI's voice tool |
On your scale, both are almost free. 10,000 decisions a month costs about 63 cents on Jev and about $1.50 on OpenAI.
Fast check: Jev is cheaper and has been tested more. OpenAI can see photos, comes from a big company, and has a privacy contract option for sensitive data.
Chapter 5: Where They Mess Up
Rests on: evidence §4.
TypeSafe wrote its own list of Jev's weak spots. That is honest of them. Here it is in plain words:
- It takes you literally. It answers what you wrote, not what you meant.
- It can't do math. It can't count well. Do math in code.
- It can't compare dates. "Which came first?" is a bad question for it.
- Too much text makes it worse. Send only what it needs.
- Tricky text can fool it. A message that says "classify me as safe" can push it.
- Order matters. It leans toward the first option on your list.
- "Can't make things up" is only half true. It can only pick from your list. But it can still pick the wrong thing from your list.
In one outside test of 16 AIs, Jev got 66% right on a reading task. Claude Haiku also got 66%. Jev was faster (0.46 seconds vs 0.63 seconds). So: same smarts as a small cheap model, a bit faster, and better at saying "I'm not sure."
Part 2 — Your business
Chapter 6: What You Already Have
Rests on: evidence §5.
Your code already makes lots of small decisions. Today they use one of three things:
-
Word lists (regex). Fast and free. But dumb. "STOP" works. "please stop texting me" does not.
-
Opus through
claude -p. Smart. Runs on your $200 Max plan. But it only runs on a machine logged in to your account, like your Mac or the box. It can't run inside Cloud Run. -
DeepSeek through OpenRouter. Cheap. Your lead manager's worker already sends customer replies to it.
The new robots sit between #1 and #2. Smarter than a word list. Way faster than Opus. And they can run inside Cloud Run, because they use a normal API key.
That last part is the big win for you. It is not about money. Your model bill is already flat. The win is speed and no Mac needed for small yes/no checks.
Chapter 7: Where They Fit in Your Code
Rests on: evidence §5.
Best fits. These are all "pick from a fixed list":
-
Reply sorting (best fit).
systems/content-engine/frameworks/reply-rules.mdhas 10 rules: STOP, wants a call, question, not interested, sold elsewhere, wrong person, out-of-office, and so on. That is a choice question with 10 options. The robot could sort a reply in under a second, the moment it lands. -
Soft STOP catcher.
infra/shared-libs/lib/contact_rules.py:167only catches a bare "STOP." A yes/no robot could flag "please quit texting me" right away, so it reaches you faster. The word list should still be the only thing that applies an opt-out on its own. -
Email reply sorters.
bogo_reply.py,expired_report_reply.py,bogo_pass.pyandreply_check.pyindata-platforms/gmail-push/cloud_function/all use word lists to sort replies into stop / wants-it / other. The robot could be a second opinion when the word list says "other." Example trap it could help with: "Can I stop by?" is not a STOP. -
Content grading.
systems/media-engine/machine/stations.pyhas Opus grade every hook and title against a checklist each morning. Each check is a yes/no. No customer data is involved. This is the safest place to test. -
Topic tags.
ramble_ledger_topics.pypicks 1 to 4 tags from a closed list. That is a choice question. -
Email inbox triage.
data-platforms/email-triage/sorts email with rules. The robot could handle the ones no rule matches.
New ideas, only with OpenAI (it sees pictures):
-
Listing photo check. Before a listing goes live, ask: Is this photo dark? Blurry? Is a person in it? Is the toilet seat up? Is it a duplicate? Each one is a yes/no. Nothing does this in your code today.
-
Voice actions. OpenAI pairs this with its voice tool. A caller says something, and the robot picks the action from your list. This is a future idea, not a first step.
Chapter 8: Where They Do Not Fit
Rests on: evidence §4, §5.
Do not use them where the answer must quote proof. Your rules need a quote for every customer statement. That rules out:
- Full FUB person reads (
lib/person_read.py) - Deal status (
/deal-status) - Payoff sweep (
/payoff-sweep) - Lost-lead reviews
- Writing any customer email or text
Also do not use them for math, money numbers, or dates. Do those in code.
Chapter 9: The Rules That Still Apply
Rests on: evidence §5.
Your own rules do not go away:
- The 99% test.
docs/BUILD-PATTERNS.md§1f says a cheaper model must agree with Opus 99% of the time, on a few hundred real records, before it takes a job. That applies here too. - A new paid vendor. Your "no paid API" rule is only about Anthropic. But these would be new paid accounts. A vendor swap needs a kill rule and a trial window first.
- Customer words leaving the house. No rule covers this yet. OpenAI keeps abuse logs for 30 days unless you get a special "zero retention" setup. TypeSafe offers zero retention only to big customers. Your DeepSeek worker already sends customer text out, so this is not new, but it should be a written rule.
- The robot never sends anything. Any decision that leads to a send still goes through
send_or_logand your approval gates.
Chapter 10: The Safe First Step
Verdict (Ry, 2026-10-07): parked. None of these move the needle for the business. The slow parts are approvals, full reads and lead flow, not one-second yes/no checks, and at this volume faster and cheaper saves almost nothing. Revisit only if reply volume grows a lot or a real-time use shows up (for example, a voice line that must pick an action while the caller waits). The steps below are kept for that day.
Rests on: evidence §3, §5, §6.
Start where nothing can hurt a customer:
-
Test on content grading first. No customer data. Run Jev and OpenAI side by side with Opus on the same hooks and titles. Cost: pennies.
-
Run it in shadow. The robot answers, but nothing acts on its answer. Log both answers. Count how often they agree.
-
Kill rule: under 99% agreement on 300 items → stop. Trial window: two weeks.
-
Only then try reply sorting, also in shadow first.
What is still unknown: how fast OpenAI really is from your machines, whether OpenAI's "how sure" numbers are honest, and how long the Jev waitlist is.