Introduction
At DevDay 2026, OpenAI introduced its new OpenAI Decisions API. Just 14 days earlier, TypeSafe AI launched Jev, a fast model made for the exact same purpose.
This simple guide explains what a Decisions API is, how it works, and provides a clear comparison of OpenAI Decisions API vs TypeSafe Jev to help you choose the right tool for your software.
What Is a Decisions API and Why Do You Need It?
In traditional AI agent workflows, software wastes time and money asking big models simple yes-or-no questions, such as:
- Does this action need human approval?
- Which code path should we run?
- Is this shell command safe?
Old Chat Approach vs New Decisions API Approach
- Old Chat Approach:
PromptTextOutputTextParserSchemaCheckAction - This approach generates unnecessary words, costs more money, and risks parsing errors.
- New Decisions API Approach:
State+BoundedQuestionTypedDecisionPolicyCheckAction - This approach skips text generation completely. The model directly picks from options you set in code.
How Does a Decisions API Work in 4 Steps?
Every Decisions API follows four simple steps:
- Capture the State: Collect only the needed data (like a ticket text, a screenshot, or a tool output).
- Ask a Bounded Question: Ask a simple question with clear choices, such as “Which team gets this ticket?”.
- Pick from Pre-Defined Options: The model picks one answer from your custom list.
- Run Code Safety Checks: Your app receives the selection and checks user permissions before taking action.
Read Also: Launching Jev Model: How does it Change the Game for AI Engineers
OpenAI Decisions API vs TypeSafe Jev: Benchmark & Comparison
While both models solve the same problem, they are built differently. The OpenAI Decisions API uses a constrained version of GPT-6 Luna, whereas TypeSafe Jev is a purpose-built model that cannot generate text at all.
Key Comparison Table
Feature / Metric | OpenAI Decisions API | TypeSafe Jev |
Status | Limited Preview | Generally Available (GA) |
Engine | Constrained GPT-6 Luna | Purpose-built System One Model |
Question Types | Bounded choice selection | Choice, Score, Noul (mixable) |
Image Input | Supported (Text + Images) | Not Supported (Text & JSON only) |
Latency Speed | 150 ms (OpenAI DevDay claim) | 0.21 s (P50) / 0.34 s (P95) (Live measured) |
Token Price (Per 1M) | $0.10 input / $0.50 output | $0.042 input / $0 output (Free) |
Probabilities | Returns overall confidence | Returns per-option probabilities + score |
Speed and Cost Breakdown
Latency Performance
- OpenAI Decisions API: OpenAI reported an average speed of 150 ms at DevDay 2026, which is 10x faster than asking GPT-6 Luna through standard chat.
- TypeSafe Jev: Live OpenRouter telemetry shows a stable 210 ms (P50) latency across real user traffic.
Cost Efficiency
TypeSafe Jev is significantly cheaper because it generates zero output tokens, costing just $0.042 per million input tokens. The OpenAI Decisions API runs on standard Luna rates ($0.10 input / $0.50 output per million tokens).
Intelligence Index vs. Cost Efficiency
Data from Artificial Analysis places Gemini 4 Argon in a favorable position regarding intelligence versus operational cost.
Simple Code Example
Below is a complete Python example using Pydantic and the `responses.parse()` method to enforce a strict decision schema. You can learn more about how schema enforcement works in the official OpenAI Structured Outputs documentation
python
from openai import OpenAI
from pydantic import BaseModel
from typing import Literal
client = OpenAI()
# 1. Define the strictly bounded decision schema
class RoutingDecision(BaseModel):
selected_team: Literal[“billing_team”, “tech_support”, “fraud_team”]
confidence_score: float
# 2. Call the API with strict schema enforcement
response = client.responses.parse(
model=”gpt-6-astra”,
input=[
{“role”: “system”, “content”: “Analyze the ticket and route to the correct team.”},
{“role”: “user”, “content”: “Customer was charged twice on invoice #9948.”}
],
text_format=RoutingDecision,
)
# 3. Access the typed decision output safely
decision = response.output_parsed
print(decision.selected_team) # Guaranteed to be one of the Literal values
Handling Edge Cases & Refusals
When working with user-generated text, always verify that the model did not issue a safety refusal before processing the result:
Python
for output in response.output:
if output.type == “message”:
for item in output.content:
if item.type == “refusal”:
print(f”Request refused: {item.refusal}”)
Where Does a Decisions API Fit in Your Agent Architecture?
A reliable AI agent stack uses 5 clear layers:
- Orchestrator: Keeps track of the task, state, and retries.
- Generative Model: Understands goals, plans steps, and writes text for humans.
- Decision Model (Decisions API): Answers fast, bounded questions (routing, risk, next action).
- Policy Engine: Enforces safety rules, rate limits, and account permissions.
- Action Layer: Runs approved tool calls and logs results.
Important Rule: Never let the Decisions API grant permissions. It can suggest that an action looks safe, but your code policy engine must always verify user permissions first.
Final Thoughts: Which One Should You Choose?
- Choose OpenAI Decisions API if: You already use the OpenAI ecosystem or need to process visual context (screenshots and image inputs) for computer-use agents.
- Choose TypeSafe Jev if: You want a fully public, production-ready solution today (GA) with the lowest possible token cost and detailed per-option probabilities.

