Estimated Reading Time: 4 minutes
Key Takeaways
- Not every AI task requires a large generative model; many enterprise workloads are fundamentally decision problems.
- Decision-oriented models such as Nimble and Jev are designed for structured classification, scoring, routing, and yes/no decisions rather than generating free-form text.
- In Agent Router, this approach helps determine whether a task should go to a script, lightweight agent, coding environment, or larger LLM.
- The business case is compelling: lower AI costs, greater privacy through local execution, and reduced dependency on expensive frontier models.
- OpenAI’s emerging Decisions API suggests that structured decision intelligence is becoming a distinct layer of the enterprise AI stack.
- The future is likely hybrid: deterministic software, decision models, local models, and frontier LLMs working together rather than sending every problem to the largest available model.
- The key architectural question is shifting from “Which AI model is smartest?” to “What is the smallest and most appropriate intelligence needed for this decision?”.
Introduction
Over the past few years, the default architecture for AI applications has been surprisingly simple: when software needs intelligence, send the problem to a large language model.
That works remarkably well when the task is genuinely generative: writing, coding, research, analysis, planning or complex reasoning.
But many enterprise AI workloads are not generative at all.
They are decisions.
Which system should handle this request? Is this transaction suspicious? Does this message require escalation? Which model should process this prompt? Is the task simple, complex or somewhere in between?
For these questions, asking a large language model to reason, generate a response, format it as JSON and then have software parse that response can be unnecessarily expensive and complex.
A new category of decision-oriented AI models is emerging to address exactly this problem.
Agent Router: a practical example
I encountered this problem while developing Agent Router.
Agent Router is a system I built in my spare time to help me managing the tokens usage through several AI models, after running out of credits one time too many.
It receives a task and determines how it should be executed. Depending on the request, it can route work to a deterministic script, a lightweight coding agent, a more capable coding environment, or a general-purpose LLM.
The business logic is straightforward:
A simple deterministic operation should not consume an expensive reasoning model. A small code change does not necessarily require the same resources as a repository-wide architectural refactoring.
Initially, this kind of classification naturally suggests another LLM call.
But that creates an interesting inefficiency: using an expensive general-purpose model simply to decide whether we need an expensive general-purpose model.
Even using a local tiny model (which eventually was my choice), still requires a general-purpose model.
Decision models offer a different approach.
From generating answers to making decisions
TypeSafe recently introduced Jev, which it describes as a “System One” model: a model optimized for making fast, structured decisions rather than producing prose. Instead of asking it to explain what it thinks, software defines the possible decisions in advance. Jev returns typed answers together with probabilities that can be incorporated directly into application logic.
Bespoke Labs' Nimble, now available through Ollama, follows a similar pattern and can run locally.
The application provides some context plus predefined questions such as:
Is this a coding task?
Does it require repository access?
Is the scope small, medium or large?
Which execution path is appropriate?
Nimble evaluates those questions and returns choices, yes/no decisions or scores. It does not need to generate paragraphs explaining its reasoning.
Ollama supports up to 64 such questions in a single request.
For my Agent Router app, this architecture is particularly attractive.
Instead of:
Prompt → LLM → generated classification → parser → routing
the architecture becomes:
Prompt → decision model → typed decisions → routing policy
The distinction sounds small. At scale, it is not.
Why business should care
The first benefit is cost.
Enterprise AI systems may eventually make millions of small decisions. Paying generative-model economics for every classification, routing or policy decision quickly becomes difficult to justify.
The second is privacy and local execution.
A model such as Nimble can run through Ollama inside infrastructure controlled by the organization. Prompts used for internal routing or classification therefore do not necessarily need to leave the environment.
The third is reduced dependence on frontier models.
Agent Router illustrates an increasingly important architecture: reserve sophisticated generative models for tasks that genuinely require sophisticated generation or reasoning.
Everything else can potentially be handled by smaller models, deterministic software or specialized decision engines.
This is not about replacing LLMs. It is about using them where they create the most value.
A category worth watching
Jev and Nimble are unlikely to be isolated examples.
OpenAI has also announced a Decisions API in limited preview, aimed at bounded questions with predefined answers for uses such as classification, request routing and selecting an agent's next action. Public technical details remain limited, so it is too early to compare its economics or behavior directly with Jev or Nimble. But its appearance reinforces the broader direction of travel.
The significance is larger than model routing.
Enterprise AI architectures are beginning to separate two fundamentally different forms of machine intelligence:
Generative intelligence, used when software needs to create, reason, explore or communicate.
And decision intelligence, used when software needs to choose, classify, score or branch.
For years, we have used large language models for both because they were the most accessible intelligent component available.
That may be changing.
The next generation of enterprise AI platforms may not be built around one increasingly powerful model. They may instead combine deterministic software, specialized decision models, local models and frontier LLMs—each used precisely where its economics and capabilities make sense.
In that architecture, the smartest AI system may not be the one using the largest model.
It may be the one that knows when not to use it.
Want to find out more about how decision models can help your business?
Book a call with me.
Watch a video about Jev.
Frequently Asked Questions (FAQ)
Q: What is a decision model?
A: A decision model is an AI model optimized to make structured choices rather than generate open-ended text. Typical outputs include classifications, scores, yes/no decisions, or selections from predefined options.
Q: How is a decision model different from a large language model?
A: Large language models are designed for broad generative tasks such as writing, reasoning, coding, and analysis. Decision models are narrower: they evaluate a defined question and return a structured answer. This can make them faster, cheaper, and easier to integrate into application logic.
Q: Why not simply use an LLM for every decision?
A: You can, but it is often inefficient. If the application only needs to decide between a few known options, using a large generative model may introduce unnecessary inference cost, latency, and complexity.
Q: What role does Nimble play in Agent Router?
A: In Agent Router, Nimble can act as the structured decision layer. It can evaluate factors such as task type, complexity, repository access, tool requirements, and scope before the routing policy decides which execution path should handle the task.
Q: What is Jev?
A: Jev is a decision-oriented model from TypeSafe designed around structured “System One” decisions. Rather than generating arbitrary responses, it answers predefined questions using typed outputs such as choices, scores, and probabilistic yes/no decisions.
Q: How does Nimble compare with Jev?
A: Both follow a similar decision-oriented approach. One important difference for local architectures is that Nimble can run through Ollama, making it attractive when privacy, local inference, and infrastructure control are priorities.
Q: Where does the OpenAI Decisions API fit?
A: The OpenAI Decisions API represents the same broader architectural direction: separating bounded decisions such as classification, routing, or action selection from general-purpose generative reasoning. It gives enterprises another option for building dedicated decision layers rather than relying exclusively on free-form LLM calls.
Q: Can decision models replace LLMs?
A: No. They solve a different class of problem. Decision models are well suited to bounded choices and structured evaluation, while LLMs remain better suited to open-ended reasoning, content generation, complex analysis, and tasks where the possible answer cannot be predefined.
Q: What are the main business benefits?
A: The strongest benefits are typically lower AI operating costs, the ability to keep some processing local, and reduced dependence on large frontier models. They can also make application behavior easier to control because the allowed outputs are defined in advance.
Q: What does the future enterprise AI architecture look like?
A: Increasingly, it is likely to be hybrid. Deterministic software will handle predictable logic, decision models will handle classification and routing, local models will address privacy-sensitive workloads, and frontier LLMs will be reserved for tasks that genuinely require advanced generative intelligence.













