Person in a blue sweater speaking on stage with a headset, standing against a wooden and black backdrop.

OpenAI’s Decisions API echoes a startup’s fast AI model as agent security becomes the next battleground

OpenAI’s new decision models API mirrors a startup’s fast classifier and could reshape AI agent security, routing and cost.

In short

OpenAI has introduced a limited-preview Decisions API that appears to compete with TypeSafe AI’s Jev, a fast low-cost decision model for automation. The new tool could help make AI agents cheaper to route and safer to monitor.

  • OpenAI launched a limited-preview Decisions API aimed at fast, constrained choices for models and agents.
  • The product closely resembles TypeSafe AI’s Jev, a startup model built for cheap, high-speed classification.
  • Decision models may become important for monitoring AI agents and reducing security risks.
  • A hackathon demo suggested Jev-style monitoring could be far cheaper than using a frontier LLM.
  • The category could become a core layer of AI infrastructure as agents become more common.

OpenAI used its latest Dev Day to quietly unveil a new Decisions API, a product that appears aimed at the same fast, low-cost “decision model” niche already being pursued by startup TypeSafe AI. The move matters because these lightweight models could become a critical layer for pricing, routing, and monitoring AI agents as companies try to make them faster, cheaper, and safer.

The announcement, made by CEO Sam Altman during Tuesday’s event, suggests OpenAI sees growing demand for a model that can choose from a small set of options at high speed instead of generating long, open-ended responses. That approach is especially relevant as AI systems increasingly act autonomously in software workflows, where every extra millisecond and every extra cent can matter.

OpenAI did not provide a full technical breakdown of the new API, and the company released it only as a limited preview. But the product’s framing, along with reactions across the developer community, makes one thing clear: the market for so-called “decision models” is moving from startup experiment to mainstream infrastructure.

What OpenAI announced at Dev Day

OpenAI’s new Decisions API is designed to narrow a model’s job to a predefined set of choices. Rather than asking a large language model to freely explain, summarize, or generate text, developers can present it with a bounded decision task — for example, selecting a category for an image, or choosing the next behavior for an AI agent.

At Dev Day, Altman described the system as a way to focus OpenAI’s Luna model on a fixed set of options, with the goal of making it dramatically faster while preserving core capabilities such as image understanding, broad language coverage, and safety controls.

Altman said concentrating the model on a specific choice can make it “extremely fast” without giving up image understanding, language support, or safety protections.

That positioning matters. In the current AI stack, not every task needs a heavyweight, general-purpose model. Many practical jobs in software are classification problems dressed up as AI: should this action proceed, which bucket does this item belong in, or is this output safe to execute? The Decisions API appears designed for exactly those jobs.

Why a “decision model” is different from a chatbot

A decision model is optimized for constrained outputs, not freeform conversation. Instead of producing paragraphs, it returns a probability-weighted choice from a shortlist, which can make it much cheaper to run and much easier to slot into production systems.

That distinction explains why the new product drew comparisons to Jev, a model TypeSafe AI launched earlier this month. Jev is built as a kind of supercharged classifier on top of a large language model, allowing developers to submit a list of possible answers and receive ranked probabilities quickly and inexpensively.

In practice, that means decision models can sit behind the scenes of an application, handling tasks like:

  • image categorization
  • routing requests to different agent behaviors
  • checking whether an AI action matches the user’s goal
  • flagging outputs for human review
  • filtering risky or irrelevant tool use

For developers trying to run AI products at scale, this is attractive because it replaces some expensive reasoning calls with a much smaller, faster inference step.

How does OpenAI’s new API compare with TypeSafe’s Jev?

OpenAI’s API appears to target the same use case as Jev, but the exact overlap remains unclear because the product is still in limited preview. TechCrunch said it had not yet observed developers stress-testing the API in the wild, so the real-world feature set and performance characteristics are still being measured.

Still, the resemblance is notable. TypeSafe AI’s Jev has been positioned as an answer to a recurring complaint in AI software: frontier models are often too slow and too expensive for the kind of repeated, tiny decisions that software makes constantly.

OpenAI’s move suggests it has reached the same conclusion. The company is now effectively offering a specialized layer for decision-making, rather than relying solely on broad generative models for every part of an agent’s workflow.

Product Company Primary use Pricing/efficiency goal Status
Decisions API OpenAI Constrained choices for images and agents High speed with lower inference cost Limited preview
Jev TypeSafe AI Fast probabilistic classification Cheap, rapid decision-making Recently launched
Frontier LLMs Multiple labs Open-ended reasoning and generation High capability, higher cost Widely deployed

The competition here is not just about product features. It is about where the AI value chain is heading. If the industry keeps breaking work into smaller and smaller model calls, then specialized decision layers could become as important as the headline-grabbing generative models that power chatbots and image tools.

Why are decision models gaining attention now?

Decision models are attracting interest because the economics of AI are pushing companies toward more efficient workflows. Large language models are powerful, but they are also comparatively slow and expensive when used for every step in a software pipeline.

That is especially true in agentic systems, where an AI tool may need to evaluate a situation, choose an action, verify the result, and then move on to the next step. If each of those stages requires a top-tier model, costs can escalate quickly.

Developers using Jev have reportedly found that the model can reduce both latency and spend, while still delivering useful signals for downstream systems. In other words, a smaller model can act as the control plane for a larger AI application.

TypeSafe CEO Diogo Almeida, who previously worked as an OpenAI engineer and co-invented reinforcement learning, has argued that the real challenge is not speed alone but intelligence at low cost.

Almeida has said that cheap and fast computation is easy to find, but that the hard part is improving the amount of intelligence delivered per dollar.

That philosophy helps explain why the market is moving in this direction. Software builders do not merely want a model that works. They want one that works often enough, quickly enough, and cheaply enough to be deployed everywhere in a product.

What is “System One” thinking?

“System One” is TypeSafe’s term for intuitive, rapid thinking, as opposed to “System 2,” which refers to slower, deliberate reasoning. The company’s argument is that many software decisions should behave more like instinct than analysis.

That framing is useful for understanding the current wave of AI tooling. Not every task needs deep reasoning. In a lot of production settings, what developers actually need is a fast, reliable classification or routing decision that keeps a larger workflow moving.

OpenAI’s product framing suggests it is comfortable adopting that same logic, even if it is not using the same terminology.

How could decision models improve AI agent security?

Decision models could improve AI agent security by making it affordable to inspect nearly every action an agent takes. That is one of the most promising use cases emerging around OpenAI’s new API and similar tools.

OpenAI has already had to address incidents in which agents behaved badly on the open internet, and one of the company’s security responses has involved using a separate model to watch for suspicious actions. The problem with that approach is cost: monitoring with a frontier model can be expensive enough to limit how often it is used.

Shapor Naghibzadeh, a cybersecurity veteran and founder of the startup QueryStory, believes a decision model could make continuous oversight much more practical. He built a hackathon demo last weekend that uses Jev to compare each agent action with the original task, then decide whether to block, flag, or allow it.

The workflow is straightforward:

  1. If an action is clearly out of bounds, stop it.
  2. If the model is uncertain, route the action to a human or a secondary review layer.
  3. If the action aligns with the task, let it proceed.

That kind of triage is valuable because it gives companies a way to add guardrails without turning every action into a costly AI reasoning call.

Naghibzadeh said a system like this could have prevented the Hugging Face incident, while costing far less than using a frontier model for the same monitoring job.

According to his example, monitoring with Jev cost just $2.94 compared with $372 using a frontier large language model. Even if the exact numbers vary by implementation, the order of magnitude is the real story: cheap oversight changes what is feasible.

What does the cost difference mean for AI agents?

The cost difference means that safety checks may become routine instead of selective. If monitoring every agentic step becomes inexpensive enough, developers can shift from sampling-based oversight to near-continuous review.

That has several implications:

  • agents can be watched more often without blowing up budgets
  • security policies can be applied consistently across workflows
  • developers can catch bad actions earlier in the chain
  • human reviewers can focus only on ambiguous cases

In effect, decision models could become the gatekeepers of agent ecosystems. Rather than replacing frontier models, they would help govern them.

This is a subtle but important change. The AI industry often talks about large models as the main event, but the next wave of software may depend just as much on the smaller systems that decide when those larger models should act, what they should do next, and whether they should be trusted at all.

Why OpenAI’s move matters for the broader market

OpenAI’s entry into this category validates a market that was already beginning to form. The company is not first to the space, and it likely will not be the last major player to launch a similar product. But when the leading frontier lab starts offering a tool in a niche, it usually signals that the niche has moved from curiosity to strategic importance.

For startups like TypeSafe, that creates both risk and opportunity. Big labs can copy popular product directions quickly, but they can also enlarge the market by proving demand exists. If OpenAI customers start building on a Decisions API, the category could gain legitimacy fast.

At the same time, OpenAI’s scale could compress margins for smaller entrants unless they can differentiate on technical quality, workflow integration, or data advantages.

That is where TypeSafe is trying to build defensibility. Almeida says the company’s edge comes from synthetic data that helps it generate statistically meaningful outputs. In a market where many models can be made fast and cheap, he argues, the deeper challenge is maintaining a strong intelligence-to-cost ratio.

How strong is the moat in this market?

The moat is likely to come from calibration, training data, and task-specific reliability rather than raw speed alone. If several vendors can offer a low-cost classifier, then the one that best matches its confidence levels to real-world outcomes will win developer trust.

That is important because a decision model’s usefulness depends heavily on calibration. A model that says “confident” when it should be uncertain can be more dangerous than no model at all, especially in security-sensitive or agentic systems.

In that sense, the battle is not just about whether a model can make a choice quickly. It is about whether it can make the right choice with the right confidence signal attached.

Where this fits in the future of AI infrastructure

The rise of decision models reflects a broader trend in AI infrastructure: decomposition. Instead of one giant model doing everything, developers are increasingly building stacks of specialized models, each responsible for a narrow task.

That architecture is appealing because it can reduce cost, improve latency, and make systems easier to debug. It also makes safety engineering more practical, since the model that monitors an agent does not need to be as expensive as the model doing the agent’s main work.

For enterprise buyers, this may be the most important lesson in OpenAI’s announcement. The value of AI will increasingly come not just from conversational fluency or raw reasoning, but from the invisible plumbing that routes, validates, and constrains model behavior.

In that world, a decision model is not a side project. It is a core control surface.

What happens next?

The next phase will be about adoption, calibration, and proof. Developers will want to know whether OpenAI’s Decisions API is actually comparable to Jev in speed, cost, and reliability. They will also want to know how it behaves under edge cases, how well it handles uncertainty, and whether it can be trusted in production agent systems.

OpenAI’s limited preview means the evidence is still early. But the market signal is already strong. A model class once viewed as a niche classification tool is now being positioned as a foundational layer for agent safety, routing, and automation.

If that trend continues, the companies that win the next phase of AI may not be the ones that build the biggest models. They may be the ones that build the best decision layers around them.

Milestone What happened Why it matters
Earlier this month TypeSafe AI launched Jev Introduced a fast, low-cost decision model for automation
Tuesday’s Dev Day OpenAI announced Decisions API Signaled that frontier labs are entering the same category
Last weekend Hackathon demo used Jev for monitoring Showed how cheap oversight could protect AI agents
Now Developers are debating calibration and safety The category is moving toward mainstream infrastructure

For OpenAI, the announcement is one more sign that the company is broadening beyond generative chat into the operational layers that make AI software usable at scale. For TypeSafe and other startups, it is evidence that the category they are chasing is real enough to attract a giant.

And for developers, it points to a future in which the most important AI call may not be the one that writes the answer, but the one that decides what happens next.

Frequently asked questions

What is OpenAI’s Decisions API?

OpenAI’s Decisions API is a new limited-preview tool that lets a model choose from a predefined set of options instead of generating open-ended text. It is designed to make tasks like classification, routing, and agent-behavior selection faster and cheaper.

How is it related to TypeSafe AI’s Jev?

It appears to target the same product category as Jev, TypeSafe AI’s fast decision model. Both systems are meant to return probabilistic choices from a shortlist, which makes them useful for low-cost automation and agent control.

Why are decision models important for AI agents?

Decision models are important because they can monitor or route agent actions at a much lower cost than frontier large language models. That makes it practical to check more steps, block risky behavior, and flag uncertain actions for review.

Could decision models improve AI safety?

Yes. Decision models could improve AI safety by screening each agent action against the original task and stopping bad behavior early. Because they are cheaper to run, developers can apply oversight more often without making the system prohibitively expensive.

What gives startups like TypeSafe a competitive edge?

Startups like TypeSafe may stand out through better calibration, synthetic data, and reliability rather than speed alone. As more companies offer similar tools, the moat will likely come from how accurately the model’s confidence matches real-world outcomes.

Share this 🚀