Fast Decision Models: Why Not Every Choice Needs a Large Language Model

A new class of AI model can make fast, structured decisions in one pass. Used as the first step, it lets a system call a large language model only when it has to.
Quick Answer
Most software that “uses AI” sends every question to a large language model (LLM). That is slow, costly, and hard to check when the question is simple, such as “is this login suspicious?” or “which team should get this ticket?” A System One model is built for these choices. You give it a description of the situation and a list of questions with fixed answers. It returns the chosen answers, how likely each option is, and how confident it is, in one pass. A confidence gate then decides whether to act, or to escalate to a language model or a person. We tested five sample decisions on tev1:4b, an open-source System One-style model, running on Ollama on a workstation with an NVIDIA RTX 3060 GPU. Each returned a structured answer in 3–4 output tokens. The clearest lesson came from a login case: when two options were nearly tied, the system must not act on either.
1. Two Ways to Think
Psychologists, most notably Daniel Kahneman in Thinking, Fast and Slow, describe human thinking as two modes.
System 1 is fast and automatic. You recognise a friend’s face, or brake for a ball rolling into the road, without deciding anything. It works well on familiar situations with a known set of options, and it is where shortcuts and biases come from.
System 2 is slow and deliberate. You use it to work through a tax form or compare two mortgage offers. It handles new situations, but it costs time and attention.

Healthy minds route between the two: most everyday decisions stay fast, and deliberate thinking is pulled in when the familiar pattern fails.
Most AI systems today use only System 2. Every question, from “is this a duplicate ticket?” to “should we block this account?”, goes to a large language model that writes out a full answer. That works, but it is slow, billed per token, and hard to check when the answer is a paragraph. For routine choices with a known set of options, that is overkill. Software needs a component that makes a quick, structured, auditable decision, and knows when not to trust it.
2. What Is a System One Model?
A System One model is a new kind of model built for one job: answering structured questions about a situation. It is not a chatbot, and it does not write prose. The first commercial example is Jev, from TypeSafe AI, which entered early access in September 2026. Open-source models built in the same style are appearing too, such as tev1:4b on Ollama.
The difference from a language model is in what comes back.

On the left, the system must write text and then extract a decision from it, and extraction can fail. On the right, the answer is already a decision with a confidence attached.
Three things define the approach:
- Input is a situation plus questions. The situation is a set of facts. Each question has a fixed list of allowed answers. There are three question types: pick one option (choice), place the situation on a rubric (score), or judge whether a statement is true (noul, a probability from 0 to 1).
- Output is typed. The answer is always one of the options you listed, with a probability for each and a confidence score from 0 to 1.
- Calibration is the goal. The stated aim is that higher confidence should mean higher real accuracy. That lets confidence drive decisions. General LLMs are known to be overconfident when asked for probabilities in text, which is one reason these models are trained differently.

Two limits to keep in mind. A System One model returns a number, not a reason, so explanations for audits have to come from logs, rules, or a language model. And vendor claims about speed and accuracy need checking against your own cases. TypeSafe’s documentation also recommends asking one factor per question and combining the answers in your own code, which keeps each judgement easy to test.
3. What Goes In, What Comes Out
The five samples cover five different domains. Each one sends a situation and a list of questions. For each question, the model returns the option it picked, how likely each option is, and a confidence score. Percentages are rounded.
Test setup. Ollama 0.35.1, the latest release as of 5 October 2026. Model tev1:4b, an open-source System One-style model. Hardware: a single NVIDIA RTX 3060 GPU on a workstation.
Scenario 1: Release monitoring

Scenario 2: Customer support ticket

Scenario 3: AI agent choosing a tool

Scenario 4: Login security check

Scenario 5: Game character deciding how to fight

Read each output the same way. The top option is the answer. The other percentages show how close the call was. The confidence tells your system whether to act on that answer.
4. What Went Wrong, and Why the Gate Matters
Four of the five answers look sensible. The support ticket was rated P0 and sent to reliability, which is what a support lead would do. The bandit’s attack fits its aggressive policy. The release answer leans toward a rollback but is unsure.
The login case is the important one. The risk call was high, which is right for a privileged account with 47 failed logins. The action call was allow, which is wrong. “Allow” (38%) and “block” (37%) were almost tied, and the confidence on the action was 0.16. The model was effectively guessing.
If a system applied “allow” automatically, a probable account takeover would get through. Better wording of the question would not fix this. What fixes it is a rule that refuses to act on a near-tie or on low confidence, and sends the case to a stronger check or a person. In security, that rule should always be in place.
What these samples do not show matters too. They are five hand-picked cases, not a test against real outcomes. They cannot tell us how often the model is right or whether its confidence matches its accuracy. Before adopting the pattern, a team should label a few hundred real cases, check that confidence tracks accuracy, and time the calls on its own hardware.
5. Production Use Cases
Each sample stands in for a business workflow that already exists in most organisations. The question is which decisions are made often enough, and under enough pressure, that a fast, auditable first answer would change the outcome.
Software release incidents
When a release goes wrong, the first ten minutes decide how long customers feel it. Today, the on-call engineer usually reads several dashboards, works out whether the problem is the new release, and then makes a call under pressure. A fast decision layer can read the same signals (error rate, response time, resource use, time since deployment) and return a clear first recommendation with a confidence score.
The business gain is less about the model being smarter than the decision being quicker and more consistent. Two engineers looking at the same spike should not reach opposite conclusions because one is tired. A rollback is also disruptive: it can undo good changes and needs its own follow-up. So the rule is simple. A high-confidence rollback recommendation can start the process, and a low-confidence one goes to a person with the context already gathered.
Customer support intake
Enterprise customers do not all wait equally well. A production outage reported by a large account should reach the reliability team within minutes, while a billing question can wait in the normal queue. Support teams often lose time because the first person to read a ticket has to classify it, and misrouted tickets bounce between teams.
A System One model can make the priority and routing call the moment a ticket arrives, using what is already known about the customer and the current service health. The business effect is faster response where it matters most and fewer tickets passed between teams. The risk is an under-prioritised ticket from a high-value customer, so the rule for top accounts should be that a low-confidence priority always goes to a human.
AI agent cost and routing
Many companies now run AI agents that search the web, query internal databases, check billing data, or send email. Each step can call a large language model, and every call costs money and time. Many requests do not need the most expensive path. A question about a customer notice may need an email search first, not a full research loop.
A fast decision layer can pick the likely first tool and flag whether private data or calculation is needed before the expensive model starts. The saving comes from skipping the wrong first steps. The safety side is small here, since a wrong first tool wastes time rather than causing harm. That makes this one of the lowest-risk places to start.
Account and login risk
Security teams face a flood of login events, and most are harmless. The few that are not can cause serious damage, especially when they involve privileged accounts such as finance or cloud administration. Blocking too much frustrates staff. Blocking too little lets attackers in.
A fast model can score each login against the signals a security analyst would check: device, location, failed attempts, password age, and account privilege. Its answer feeds the response. Low-risk events pass, clear high-risk events are challenged or blocked, and anything ambiguous goes to the security team. Section 4 shows why the ambiguous band matters: the login case had a near-tie between “allow” and “block”, and the right behaviour there is to escalate, not guess.
Game and simulation AI
Non-player characters in games and simulations have to respond in real time, across many characters at once, without a large cost per character. Hand-written rules become predictable and repetitive. Calling a large language model for every character’s choice is too slow and too expensive at scale.
A System One model can give each character a structured choice (attack, defend, retreat) and a graded emotional state, such as fear, on each tick. The designer keeps control through the options and the policy, and the variety comes from the probabilities. The risk is low, since a poor choice in a game is a visible but harmless mistake, which makes this a good first place to try the pattern.
What ties these together
Across the five cases, three things recur. The decisions are frequent and time-sensitive, and the options are already known. Each answer feeds another system, such as a paging tool, a ticket queue, a model router, a fraud pipeline, or a game engine. And the gate’s strictness should match the cost of being wrong (Section 6).

Each use case needs its own labelled data, thresholds, and escalation rules before it runs in production. QSS Technosoft can help teams decide which decisions to automate first and integrate System One-style models into their existing support, operations, security, and product systems, with the gates and audit trails described here.
6. The Architecture: Fast First, Language Model When Needed

Three rules make the gate work:
- The gate is deterministic. It compares confidence and the gap between the top two options against thresholds you set and test. It does not ask the model whether it is sure.
- The margin matters as much as the confidence. A 38% option beside a 37% rival is not a decision, however confident the headline number looks.
- Stakes set the strictness. A support ticket can act at a moderate threshold, with a reviewer correcting mistakes. A security action needs a high threshold, and near-ties always escalate.
The language model is not removed. It handles the minority of cases the fast model is unsure about, where its ability to reason and explain is worth the cost.
7. When to Use It
Good fit:
- High-volume decisions with a known, finite set of options
- Real-time paths where seconds are too slow
- Decisions that must be logged, audited, and thresholded
- Workflows where a language model is already in use and you want to stop paying for the easy cases
Look elsewhere if:
- The answer must come with a written explanation for a customer or regulator. Pair the model with a language model or rule trace.
- The options cannot be listed in advance. Free-text generation is the right tool there.
- You have no labelled outcomes to measure against.
- The stakes are high and there is no escalation path.

9. Sources
- TypeSafe AI, Introduction: https://docs.typesafe.ai/introduction
- DataCamp, “Jev: TypeSafe’s System One Model Explained”: https://www.datacamp.com/blog/system-one-models-jev
- MarkTechPost, “TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text”: https://www.marktechpost.com/2026/09/19/typesafe-ai-releases-jev/
- The Register, “TypeSafe AI debuts model for machines that plays Doom”: https://www.theregister.com/ai-and-ml/2026/09/16/typesafe-ai-debuts-model-for-machines-that-plays-doom/5296711
At QSS Technosoft, we are happy to walk through where a System One layer would fit in your stack.
Interested in applying this pattern to your decision flows?
Book a Discovery Call →


