Analysis · AI in production · October 2026
Your AI thinks too long: System 1 AI explained
For two years the industry taught language models to think slowly. Most questions a business asks them have a short list of allowed answers, and a new class of models picks one in milliseconds.
The short version
Daniel Kahneman split human thinking into a fast, automatic System 1 and a slow, effortful System 2. Since OpenAI's o1 in September 2024, AI vendors have poured effort into the System 2 side: models that "think" before they answer. That helps on maths and code, and it costs tokens and seconds on every call. In September 2026 TypeSafe AI, co-founded by former OpenAI researcher Diogo Almeida, launched Jev, a model that writes no text at all. It gets a situation and a list of allowed answers and returns one of them with a probability. Remigiusz Kinas has already built an open Polish counterpart, basal, on top of Bielik. Research from Meta shows where this works (judging answers with 4 tokens instead of 2,118) and where it collapses (school maths). In my own small test on 30 Polish customer-service decisions, a thinking model was slightly more accurate (29 against 27) but took 3.9 seconds per decision against basal's 118 milliseconds, and at the stricter of its two shipped confidence thresholds basal decided 23 cases alone without a single error. The article ends with a checklist for trying this on your own process.
A question for the first second
A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?
If "10 cents" arrived before you finished reading, you are in good company. Shane Frederick put this question into his Cognitive Reflection Test in 2005, and Kahneman reports that more than half of the students at Harvard, MIT and Princeton gave the intuitive answer. The right one is 5 cents: the bat costs $1.05, which is one dollar more.
Kahneman built Thinking, Fast and Slow (2011) around this kind of moment. System 1 answers at once and without effort: it recognises a face in a crowd, and it hands you "10 cents". System 2 is what you need for 17 × 24. It is slow and tiring, and it is the part that can check what System 1 proposed. The labels themselves come from the psychologists Keith Stanovich and Richard West (2000). Kahneman treats the two systems as a convenient way to talk about the mind; he does not place them anywhere in the brain.
One honest caveat before I build on the book. Not all of it survived psychology's replication crisis. The chapter on social priming did not hold up, and Kahneman said so himself in 2017: "I placed too much faith in underpowered studies." The two-systems model and the bat-and-ball result are on much firmer ground, and they are all this article relies on.
Two years of teaching AI to think slowly
The AI field borrowed Kahneman's vocabulary early. At NeurIPS in December 2019 Yoshua Bengio gave a keynote titled "From System 1 Deep Learning to System 2 Deep Learning": neural networks had become excellent at perception, and the next step was reasoning. In November 2023 Andrej Karpathy described the large language models of the time as pure System 1, sampling one word after another with no way to stop and check.
Then the industry went after System 2 in earnest. OpenAI's o1 (September 2024) spends extra computation writing a long internal chain of reasoning before it answers. Anthropic presented Claude 3.7 Sonnet (February 2025) as the first hybrid model that can answer at once or think at length. Qwen3 (April 2025) shipped one model with a thinking switch, and in July 2025 the Qwen team split it into separate fast and thinking models again, saying that was the way to get the best quality out of each. GPT-5 (August 2025) added a router that decides for you when a question deserves the slow path.
For a business the bill arrives in two currencies. Thinking is paid in tokens: on Claude, for example, thinking tokens are billed as output tokens, the most expensive kind. And it is paid in time. A thinking model can take tens of seconds before the first word of the answer. A chat widget can live with that, a phone line cannot: across ten languages, Stivers and colleagues (PNAS, 2009) found that people answer each other within 0 to 200 milliseconds of the end of a turn.
Jev: a model that does not write
On 15 September 2026 TypeSafe AI, a San Francisco company, came out of stealth with a $40M seed round led by DCVC. Its CEO Diogo Almeida spent about four years at OpenAI working on InstructGPT, ChatGPT and GPT-4, after Google Brain. The launch post says it plainly: "We were inspired by Daniel Kahneman, Thinking, Fast and Slow." The model is named after William Stanley Jevons, the economist behind the Jevons paradox, which says that making a resource cheaper tends to increase how much of it we use.
Jev does not generate text. You send it a situation and a typed question, and it returns a probability distribution over the answers you allowed. There are three kinds of question: choice (pick one of these), score (rate this) and a yes/no type TypeSafe calls noul. The company quotes 70 to 500 milliseconds per decision, $0.042 per million input tokens and nothing for output, because there is no output text to pay for.
question: Is this a complaint? (yes / no) answer: yes 0.990 112 ms"Moja paczka miała przyjść w poniedziałek, jest czwartek, a status w śledzeniu od trzech dni to „w drodze”. Gdzie jest moja przesyłka?" (My parcel was due Monday, it is Thursday, and tracking has said "in transit" for three days.)
question: Which team handles it? (complaints / billing / delivery / technical / sales) answer: delivery 0.997 117 ms
A generative model given the same job writes a sentence, and your code has to hope the sentence contains one of the words it expects. A decision model cannot answer outside the list. TypeSafe markets this as a model that "can't hallucinate", which is true only in a narrow sense: it can still pick the wrong option from the list, it just cannot invent a new one. And the headline numbers, 193.6 times faster and 444.6 times cheaper, come from TypeSafe's own tests on four in-house workflows, scored against the averaged answers of two large models rather than human-checked labels. TypeSafe itself calls them "on the higher end". There is no technical paper and no public weights, so for now those are claims, not results.
The first independent use I know of is academic. Researchers at the University of Texas at Dallas published Jev-Mem (arXiv, 21 September 2026), which uses Jev to run an AI agent's memory: deciding what kind of memory an observation is, which memories relate, and when to stop searching. A normal language model writes only the final answer. On the LoCoMo long-conversation benchmark, where another language model grades the answers, they report an overall score of 0.777 against 0.700 for their strongest baseline, with 158 seconds to build the memory against 1,044 seconds for the fastest competitor and 0.93 seconds per query. It is one benchmark, the strongest baseline is the authors' own earlier system, and the text disagrees with its own tables in places, so I read it as an early signal of how people will use these models.
Where a fast decision is enough
Look at the traffic that runs through a typical customer-service or back-office process, and most of the questions an AI is asked there have a closed answer. Which team gets this ticket. Is the customer angry. Is this a complaint under the returns policy. Does the form have every required field, as long as the rule is simple. Should a person take over now. Is this reply safe to send. None of them needs an essay, and every one of them runs thousands of times a day, so speed and cost add up.
There is also a strong research result behind the idea. In "Distilling System 2 into System 1" (Meta, July 2024) Ping Yu and colleagues let a model reason at length many times over, kept the answers it agreed with itself on, and trained it to give those answers directly. Used as a judge of answer quality, the fast version agreed with human raters 58.4% of the time using 4 tokens, against 49.1% for the slow reasoning method using 2,118 tokens, and it beat GPT-4 on that task. On a test of resisting leading, biased questions the fast version scored 81.3% against 76.0% for the slow method, with 56 tokens instead of 147.
The same paper is also the best warning label. On GSM8k, a set of school maths word problems, the fast version scored 7.1%, against 52.8% with step-by-step reasoning. Some questions cannot be answered without working them out, for models as for people. The practical rule is to give the fast model decisions with a closed answer, and keep the slow model, or a human, for anything that has to be worked out.
Speed also changes which products you can build at all. A voice agent that pauses for several seconds on every turn feels broken, whatever the quality of the answer. A decision in tens of milliseconds fits inside the gap people leave in normal conversation, so routing, escalation and safety checks can run on every turn without the caller noticing.
The Polish version already exists
At the end of September, two weeks after Jev's launch, Remigiusz Kinas, one of the people behind the Bielik models, released basal: an open, Apache-2.0 family of decision models for Polish and English, with the same kind of interface as Jev. TypeSafe never published how Jev is trained, so basal is Kinas's own recipe: Bielik-4.5B-v3.0 and Bielik-1.5B-v3.0 fine-tuned on 63,663 decision examples (deadlines from the Polish holiday calendar, amounts and VAT, questions about statutes, whether the evidence is sufficient), each either generated by code, grounded in statute text or checked by two independent verifier models, about a fifth of them in English. On top comes calibration, so that a 0.9 means the model is right about nine times in ten. One design choice matters for Polish in particular: the Bielik v3 tokenizer needs 17 to 51% fewer tokens for Polish text than the alternatives he measured, which makes every decision cheaper.
His results are strong on his home ground. On his own set of 7,081 Polish decisions, basal-4.5B is right 88.4% of the time against 78.0% for Jev. The number I like most is coverage: if you only let the model decide when it is confident enough to keep errors near 1%, basal-4.5B still decides 58.6% of the cases on its own, and Jev 18.1%. Everything else goes to a person or a slower model. On an H100 GPU one decision takes 12.5 milliseconds. Kinas is also candid about the limits: that test set is his own and unpublished, and on the public English JevBench basal-4.5B scores 0.740 against 0.861 for Jev. That is exactly why an independent test on someone else's data is worth running.
My own test on thirty Polish decisions
Kinas's numbers come from his own test set, so I wrote a different one: thirty short Polish messages of the kind an online shop, a telco or a bank receives every day. Ten ask whether a message is a complaint, ten route it to one of five teams, five ask whether a person should take over, and five check a complaint form against a four-part completeness rule. I ran basal-4.5B and basal-1.5B on my Mac (Apple M5 Max, plain PyTorch, both option orders averaged, as basal does) against Qwen3-14B, a general model about three times larger, with its thinking mode on and off.
The thinking model was the most accurate: 29 of 30 with thinking on, 28 with it off, and 27 for both basal models. With thirty cases one answer is worth 3.3 points, so I would not read much into that gap. The speed gap is large: basal-4.5B needed a median 118 milliseconds per decision and basal-1.5B 50 milliseconds, while Qwen3 with thinking needed 3.9 seconds, after writing a median 224 tokens of reasoning first. With thinking off, Qwen3 answered in 289 milliseconds.
Two of basal-4.5B's three errors came from the completeness rule, a check of four conditions: an order number, a description of the defect, what the customer wants, and a bank account number, but only when they want their money back. basal called 4 of the 5 forms complete, including a refund request with no bank account. Qwen3 with thinking went through the conditions one by one and got all five right. Checking a list of conditions one by one is the kind of task the Meta paper found fast models fail at.
For a manager, the more useful number is about confidence. basal ships with confidence thresholds fixed in advance by its author. At the stricter of its two (confidence of at least 0.913, set for roughly 1% errors), basal-4.5B decided 23 of the 30 cases on its own and got all 23 right. The remaining 7 would have gone to a person, and all three of its mistakes were among them. The thinking model is right more often, but in this setup it gives no comparable confidence number, so there is no clean way to know which of its answers to trust without a second check.
The limits are real and I want them on the record. I wrote the cases and the labels myself, with no second annotator, and one label (a question about a Wi-Fi router before buying it, which I filed under sales) is debatable, since three of the four systems, including the thinking model, sent it to technical support. Thirty cases make a demonstration. And a Mac is not an H100: my simple port skips basal's speed optimisations (Kinas reports 12.5 milliseconds on an H100), while Qwen3 ran on Apple's fast MLX path with 4-bit weights, so on speed the comparison, if anything, flatters the thinking model. The cases, labels, scripts and raw outputs are on GitHub at agentGreg/system-1-ai, so you can rerun them or add your own.
Which decisions to hand over
Two questions sort almost every AI decision in a process. How open is the answer: a fixed list of options, or free-form text? And what does a mistake cost? Closed answers with a cheap mistake belong to a fast model outright. Closed answers with an expensive mistake still suit a fast model, as long as you use its probability as a gate and hand everything below the threshold to a person. Free-form work, a reply to draft or a thread to summarise, belongs to a generative model, with thinking switched on where the task needs working out. And where the answer is open and the stakes are high, AI supports a person who decides.
If you want to try this on your own process, this is the order I would do it in.
- List the closed decisions. Walk one real process end to end and write down every point where a person or a model picks from a known set of answers. Most teams find more than they expect.
- Label a few hundred real cases. Take them from your own tickets and documents, labelled by the people who do the work today, because your inbox has its own language.
- Measure coverage at your error target as well as accuracy. Decide how many mistakes you can live with, then ask what share of cases the model handles alone at that level. That share is your saving.
- Route what it is unsure of. Below the threshold, send the case to a person or to a slower model. The fast model's probability is the switch.
- Time it end to end, including the network. A model that answers in 12 milliseconds on a GPU is slower across the Atlantic.
- Re-check regularly. Your customers change their language faster than your test set does.
What I'm looking at next
Two things. Whether basal's calibration holds on data that did not come from its author's pipeline, which is the gap Kinas himself leaves open, and I have Polish data from my own research on model confidence to test it with. And what a domain fine-tune on top of basal-4.5B gives on real customer-service decisions, measured against a large model acting as judge, on cost and on agreement with people. Stay tuned.
This is an analysis of public work plus a small experiment of my own. Jev's figures are the vendor's claims; basal's figures come from its technical report; the Meta figures from the paper below.
TypeSafe AI, introducing Jev → basal on GitHub → My test set and scripts → Yu et al., Distilling System 2 into System 1 → Jiang et al., Jev-Mem → Stivers et al., turn-taking →
Thanks to Remigiusz Kinas for publishing basal openly, with an unusually honest technical report. Related reading on this site: A confidence dial you can read, and turn, my paper on how well Polish models know what they don't know.