Book a Discovery Call

Fast Decision Models: Why Not Every Choice Needs a Large Language Model

Shreyak ChakrabortyBy Shreyak Chakraborty QSS Technosoft October 5, 2026 10 min read
Fast decision models: a System One model handles routine AI decisions in one pass and escalates uncertain cases to a large language model through a confidence gate

A new class of AI model can make fast, structured decisions in one pass. Used as the first step, it lets a system call a large language model only when it has to.

Quick Answer

Most software that “uses AI” sends every question to a large language model (LLM). That is slow, costly, and hard to check when the question is simple, such as “is this login suspicious?” or “which team should get this ticket?” A System One model is built for these choices. You give it a description of the situation and a list of questions with fixed answers. It returns the chosen answers, how likely each option is, and how confident it is, in one pass. A confidence gate then decides whether to act, or to escalate to a language model or a person. We tested five sample decisions on tev1:4b, an open-source System One-style model, running on Ollama on a workstation with an NVIDIA RTX 3060 GPU. Each returned a structured answer in 3–4 output tokens. The clearest lesson came from a login case: when two options were nearly tied, the system must not act on either.

1. Two Ways to Think

Psychologists, most notably Daniel Kahneman in Thinking, Fast and Slow, describe human thinking as two modes.

System 1 is fast and automatic. You recognise a friend’s face, or brake for a ball rolling into the road, without deciding anything. It works well on familiar situations with a known set of options, and it is where shortcuts and biases come from.

System 2 is slow and deliberate. You use it to work through a tax form or compare two mortgage offers. It handles new situations, but it costs time and attention.

System one vs system two routing

Healthy minds route between the two: most everyday decisions stay fast, and deliberate thinking is pulled in when the familiar pattern fails.

Most AI systems today use only System 2. Every question, from “is this a duplicate ticket?” to “should we block this account?”, goes to a large language model that writes out a full answer. That works, but it is slow, billed per token, and hard to check when the answer is a paragraph. For routine choices with a known set of options, that is overkill. Software needs a component that makes a quick, structured, auditable decision, and knows when not to trust it.

2. What Is a System One Model?

A System One model is a new kind of model built for one job: answering structured questions about a situation. It is not a chatbot, and it does not write prose. The first commercial example is Jev, from TypeSafe AI, which entered early access in September 2026. Open-source models built in the same style are appearing too, such as tev1:4b on Ollama.

The difference from a language model is in what comes back.

Llm vs system one pipeline

On the left, the system must write text and then extract a decision from it, and extraction can fail. On the right, the answer is already a decision with a confidence attached.

Three things define the approach:

  • Input is a situation plus questions. The situation is a set of facts. Each question has a fixed list of allowed answers. There are three question types: pick one option (choice), place the situation on a rubric (score), or judge whether a statement is true (noul, a probability from 0 to 1).
  • Output is typed. The answer is always one of the options you listed, with a probability for each and a confidence score from 0 to 1.
  • Calibration is the goal. The stated aim is that higher confidence should mean higher real accuracy. That lets confidence drive decisions. General LLMs are known to be overconfident when asked for probabilities in text, which is one reason these models are trained differently.
System One model question types: choice picks one option, score places the situation on a rubric, and noul returns a probability from 0 to 1 that a statement is true

Two limits to keep in mind. A System One model returns a number, not a reason, so explanations for audits have to come from logs, rules, or a language model. And vendor claims about speed and accuracy need checking against your own cases. TypeSafe’s documentation also recommends asking one factor per question and combining the answers in your own code, which keeps each judgement easy to test.

3. What Goes In, What Comes Out

The five samples cover five different domains. Each one sends a situation and a list of questions. For each question, the model returns the option it picked, how likely each option is, and a confidence score. Percentages are rounded.

Test setup. Ollama 0.35.1, the latest release as of 5 October 2026. Model tev1:4b, an open-source System One-style model. Hardware: a single NVIDIA RTX 3060 GPU on a workstation.

Scenario 1: Release monitoring

Scenario1 release monitoring

Scenario 2: Customer support ticket

Scenario2 support ticket

Scenario 3: AI agent choosing a tool

Scenario3 agent tool choice

Scenario 4: Login security check

Scenario4 login security

Scenario 5: Game character deciding how to fight

Scenario5 game character

Read each output the same way. The top option is the answer. The other percentages show how close the call was. The confidence tells your system whether to act on that answer.

4. What Went Wrong, and Why the Gate Matters

Four of the five answers look sensible. The support ticket was rated P0 and sent to reliability, which is what a support lead would do. The bandit’s attack fits its aggressive policy. The release answer leans toward a rollback but is unsure.

The login case is the important one. The risk call was high, which is right for a privileged account with 47 failed logins. The action call was allow, which is wrong. “Allow” (38%) and “block” (37%) were almost tied, and the confidence on the action was 0.16. The model was effectively guessing.

If a system applied “allow” automatically, a probable account takeover would get through. Better wording of the question would not fix this. What fixes it is a rule that refuses to act on a near-tie or on low confidence, and sends the case to a stronger check or a person. In security, that rule should always be in place.

What these samples do not show matters too. They are five hand-picked cases, not a test against real outcomes. They cannot tell us how often the model is right or whether its confidence matches its accuracy. Before adopting the pattern, a team should label a few hundred real cases, check that confidence tracks accuracy, and time the calls on its own hardware.

5. Production Use Cases

Each sample stands in for a business workflow that already exists in most organisations. The question is which decisions are made often enough, and under enough pressure, that a fast, auditable first answer would change the outcome.

Software release incidents

When a release goes wrong, the first ten minutes decide how long customers feel it. Today, the on-call engineer usually reads several dashboards, works out whether the problem is the new release, and then makes a call under pressure. A fast decision layer can read the same signals (error rate, response time, resource use, time since deployment) and return a clear first recommendation with a confidence score.

The business gain is less about the model being smarter than the decision being quicker and more consistent. Two engineers looking at the same spike should not reach opposite conclusions because one is tired. A rollback is also disruptive: it can undo good changes and needs its own follow-up. So the rule is simple. A high-confidence rollback recommendation can start the process, and a low-confidence one goes to a person with the context already gathered.

Customer support intake

Enterprise customers do not all wait equally well. A production outage reported by a large account should reach the reliability team within minutes, while a billing question can wait in the normal queue. Support teams often lose time because the first person to read a ticket has to classify it, and misrouted tickets bounce between teams.

A System One model can make the priority and routing call the moment a ticket arrives, using what is already known about the customer and the current service health. The business effect is faster response where it matters most and fewer tickets passed between teams. The risk is an under-prioritised ticket from a high-value customer, so the rule for top accounts should be that a low-confidence priority always goes to a human.

AI agent cost and routing

Many companies now run AI agents that search the web, query internal databases, check billing data, or send email. Each step can call a large language model, and every call costs money and time. Many requests do not need the most expensive path. A question about a customer notice may need an email search first, not a full research loop.

A fast decision layer can pick the likely first tool and flag whether private data or calculation is needed before the expensive model starts. The saving comes from skipping the wrong first steps. The safety side is small here, since a wrong first tool wastes time rather than causing harm. That makes this one of the lowest-risk places to start.

Account and login risk

Security teams face a flood of login events, and most are harmless. The few that are not can cause serious damage, especially when they involve privileged accounts such as finance or cloud administration. Blocking too much frustrates staff. Blocking too little lets attackers in.

A fast model can score each login against the signals a security analyst would check: device, location, failed attempts, password age, and account privilege. Its answer feeds the response. Low-risk events pass, clear high-risk events are challenged or blocked, and anything ambiguous goes to the security team. Section 4 shows why the ambiguous band matters: the login case had a near-tie between “allow” and “block”, and the right behaviour there is to escalate, not guess.

Game and simulation AI

Non-player characters in games and simulations have to respond in real time, across many characters at once, without a large cost per character. Hand-written rules become predictable and repetitive. Calling a large language model for every character’s choice is too slow and too expensive at scale.

A System One model can give each character a structured choice (attack, defend, retreat) and a graded emotional state, such as fear, on each tick. The designer keeps control through the options and the policy, and the variety comes from the probabilities. The risk is low, since a poor choice in a game is a visible but harmless mistake, which makes this a good first place to try the pattern.

What ties these together

Across the five cases, three things recur. The decisions are frequent and time-sensitive, and the options are already known. Each answer feeds another system, such as a paging tool, a ticket queue, a model router, a fraud pipeline, or a game engine. And the gate’s strictness should match the cost of being wrong (Section 6).

System One model production use cases ordered by cost of a wrong decision, from game AI and AI agent routing to support intake, release incidents, and login security, with confidence gate strictness rising alongside

Each use case needs its own labelled data, thresholds, and escalation rules before it runs in production. QSS Technosoft can help teams decide which decisions to automate first and integrate System One-style models into their existing support, operations, security, and product systems, with the gates and audit trails described here.

6. The Architecture: Fast First, Language Model When Needed

Confidence gate escalation

Three rules make the gate work:

  • The gate is deterministic. It compares confidence and the gap between the top two options against thresholds you set and test. It does not ask the model whether it is sure.
  • The margin matters as much as the confidence. A 38% option beside a 37% rival is not a decision, however confident the headline number looks.
  • Stakes set the strictness. A support ticket can act at a moderate threshold, with a reviewer correcting mistakes. A security action needs a high threshold, and near-ties always escalate.

The language model is not removed. It handles the minority of cases the fast model is unsure about, where its ability to reason and explain is worth the cost.

7. When to Use It

Good fit:

  • High-volume decisions with a known, finite set of options
  • Real-time paths where seconds are too slow
  • Decisions that must be logged, audited, and thresholded
  • Workflows where a language model is already in use and you want to stop paying for the easy cases

Look elsewhere if:

  • The answer must come with a written explanation for a customer or regulator. Pair the model with a language model or rule trace.
  • The options cannot be listed in advance. Free-text generation is the right tool there.
  • You have no labelled outcomes to measure against.
  • The stakes are high and there is no escalation path.
When to use a System One model: good fit for high-volume, real-time, auditable decisions with known options; look elsewhere when explanations are required, options are open-ended, or there is no labelled data or escalation path

9. Sources

At QSS Technosoft, we are happy to walk through where a System One layer would fit in your stack.

Interested in applying this pattern to your decision flows?

Book a Discovery Call →
Shreyak Chakraborty
About the Author — Shreyak Chakraborty

Shreyak Chakraborty is a Technical Lead at QSS Technosoft working on AI, agentic AI, machine learning, deep learning, knowledge graphs, data systems, IoT and robotics simulation. LinkedIn →

Frequently Asked
Questions

Close. A classifier usually returns one label. A System One model returns probabilities over options, a confidence score, and handles several questions in one request, designed to feed a gate.

No. The answer is guaranteed to be a valid option, but not necessarily the right one. The login action in Section 4 is a valid output and a poor choice. Correctness has to be measured against labelled cases.

No. Jev is TypeSafe AI’s model. tev1:4b is an open-source System One-style model on Ollama. Its training and behaviour are its own, so results from our samples describe tev1:4b only.

No. TypeSafe reports 40x–200x faster than frontier LLMs, $0.042 per million input tokens, and a 0% structured output error rate. We have not benchmarked these. Treat them as claims until your own measurements confirm them.

Label past decisions, see how accuracy changes at different confidence and margin levels, and set thresholds per decision type. Use stricter thresholds for actions that are hard to reverse.

No. It reduces how often the LLM is needed. The fast model handles routine cases, and the LLM handles the ambiguous, explanatory, or novel ones.

WhatsApp