For much of the generative AI boom, one question has dominated business experimentation:
What can AI produce for us?
Write an email.
Summarise a document.
Create some marketing copy.
Draft a proposal.
Produce some code.
Turn these notes into a presentation.
These applications matter. They can save time, increase capacity and make useful capabilities available to more people.
But OpenAI’s o1 models suggest that we should start asking a different question.
What can AI help us work out?
OpenAI introduced o1-preview in September 2024 as a new series of models designed to spend more time reasoning before responding. The company reported particularly strong performance in areas such as mathematics, science and competitive programming.
The full o1 model followed in December, and OpenAI has since released o3-mini, another model explicitly designed around reasoning.
There are plenty of reasons to be cautious about translating benchmark performance into business outcomes.
But the direction is interesting.
The frontier isn’t only moving towards AI that produces better answers.
It is moving towards AI that can spend more computational effort working through the problem before producing them.
For businesses, that could eventually be much more important than generating better emails.
Generation was the obvious place to start#
It isn’t surprising that content generation became one of the first widely understood applications of generative AI.
The value is immediate.
Give a model some information and ask it to produce something.
The interaction is simple enough for almost anyone to understand.
It also maps neatly onto activities that happen constantly inside organisations.
People write emails.
They produce documents.
They summarise meetings.
They create presentations.
They analyse spreadsheets.
They write software.
Reducing the time required to perform those activities can create real value, particularly when they happen frequently.
But focusing too heavily on the output can obscure something important.
In many valuable activities, producing the answer isn’t the difficult part.
Working out what the answer should be is.
A salesperson can write a customer email reasonably quickly.
Deciding which customer deserves attention might be more consequential.
A finance team can produce a report.
Understanding what the numbers imply might matter more.
A manager can write a plan.
Identifying the right plan is the difficult bit.
A company can produce a proposal.
Working out how to structure the commercial offer may have much greater economic significance than writing the document itself.
This distinction becomes important as AI gets better at reasoning.
The opportunity moves from helping people produce work towards helping them think through work.
Reasoning changes the shape of the opportunity#
The term “reasoning model” risks making this sound more mysterious than it needs to be.
For a business leader, the technical details are less important than the behavioural difference.
Traditional language-model interactions can feel remarkably immediate: ask a question and receive an answer.
Models such as o1 are designed to allocate more computation to working through difficult problems before responding.
OpenAI says its approach improves with both additional reinforcement learning during training and additional time spent reasoning when answering.
That matters because many commercially important problems aren’t simple retrieval or generation tasks.
They contain constraints.
Trade-offs.
Incomplete information.
Multiple possible answers.
Dependencies.
Exceptions.
Uncertainty.
A useful answer may require several intermediate steps rather than a plausible response generated immediately.
That’s much closer to the texture of real business problems.
Should we increase the price of this product?
Why is this customer segment becoming less profitable?
Which sales opportunities deserve attention?
What is causing this operational bottleneck?
Which supplier creates the greatest risk?
What assumptions in this forecast appear weakest?
How should we allocate limited resources?
Which option best satisfies these competing requirements?
None is simply a writing task.
They are reasoning tasks.
And businesses contain an enormous number of them.
The next important unit of AI productivity may not be the document produced. It may be the problem understood better.
Better decisions can be worth more than faster outputs#
This connects to something I’ve written about before.
The cost of making a decision and the value of making a better decision can be radically different.
Imagine an employee spends 30 minutes researching something before making a commercial decision.
AI reduces the research to five minutes.
That’s useful.
Twenty-five minutes have potentially been released.
But suppose AI also helps the employee identify an important factor they would otherwise have missed.
Now the economics have changed.
The value no longer comes primarily from saving 25 minutes.
It comes from changing the decision.
Perhaps the difference is insignificant.
Perhaps it affects £100.
Perhaps it affects £100,000.
The time spent making the decision tells us very little about its economic importance.
This is why reasoning capability could expand the business case for AI beyond conventional productivity.
If AI can help people interrogate information, compare alternatives, identify inconsistencies, test assumptions and explore consequences, its economic value may increasingly appear in the quality of human judgement, not merely the amount of human labour removed.
And when those improvements occur repeatedly, small changes can create enormous value.
That’s a very different opportunity.
AI doesn’t need to become the decision-maker#
There’s an obvious leap from this argument that I don’t think businesses need to make.
If AI becomes better at reasoning, should we hand important decisions to it?
Not necessarily.
There is a large and economically interesting territory between:
AI writes things for humans
and
AI makes decisions instead of humans.
AI can contribute to a decision without owning it.
Imagine a manager considering whether to make a significant investment.
AI could help organise the available information.
It could identify missing data.
It could challenge assumptions in the business case.
It could construct alternative scenarios.
It could compare the proposal with previous decisions.
It could identify risks.
It could argue against the preferred option.
It could surface questions the management team hasn’t considered.
The human still decides.
But the decision process has changed.
This may be one of the most practical ways businesses should initially think about increasingly capable reasoning models.
Not:
Which decisions can we automate?
But:
Which decisions would benefit from better intelligence around the person making them?
That’s a much broader question.
Reasoning could make expertise more accessible#
There’s another potentially important consequence.
Many organisations contain people whose value comes partly from knowing how to approach difficult problems.
Experienced employees don’t merely possess information.
They know which questions to ask.
They recognise patterns.
They know where problems usually hide.
They understand which variables matter.
They can decompose a complicated situation into manageable pieces.
That expertise is expensive and scarce.
Reasoning models could potentially make some forms of structured problem-solving more accessible to people who don’t possess the same depth of experience.
That doesn’t mean expertise becomes unnecessary.
The opposite may be true.
An expert who understands a problem can often recognise when an AI answer is superficially convincing but wrong.
They understand context the model doesn’t.
They know when an exception matters.
They can distinguish a technically correct answer from a commercially useful one.
But AI may allow expertise to travel further.
A senior employee could potentially encode approaches, criteria and knowledge that help less experienced colleagues reason through problems more effectively.
A specialist could serve more internal demand.
A manager could interrogate unfamiliar information before escalating to an expert.
If that happens, AI isn’t simply automating expertise.
It is increasing the number of places where something resembling expert-level analytical support can economically be applied.
That could be enormously valuable.
But intelligence without context has limits#
There is a problem.
A frontier model may know an extraordinary amount about the world while knowing remarkably little about your business.
It doesn’t automatically understand your customers.
Your margins.
Your products.
Your contracts.
Your internal terminology.
Your previous decisions.
Your processes.
Your systems.
Your employees.
Your risk appetite.
Your strategy.
Your exceptions.
Or the strange historical reasons your organisation does something in a particular way.
That matters enormously.
Ask a reasoning model a mathematics problem and the necessary information may be contained within the problem itself.
Ask it whether your business should change its pricing strategy and the situation is very different.
The quality of the reasoning can only be as useful as the information and context available to reason with.
This creates an important distinction between general intelligence and organisation-specific usefulness.
The frontier models may continue becoming dramatically more capable.
But businesses will still have to work out how those capabilities connect to their own information, systems and operating context.
A model can be brilliant at reasoning and still reason badly about your business if it doesn’t know enough about your business.
That could become one of the defining implementation problems of enterprise AI.
The benchmark isn’t the business case#
OpenAI’s published results for o1 are impressive.
The original o1-preview ranked in the 89th percentile on Codeforces competitive-programming questions and performed strongly on difficult mathematics and science evaluations.
Those results tell us something useful about the direction of model capability.
They don’t tell us the return on investment for a manufacturer in Manchester.
This sounds obvious, but it’s an important discipline.
Frontier-model announcements naturally focus on benchmarks because model developers need ways to measure technical progress.
Businesses need different measurements.
Did sales conversion improve?
Did forecast accuracy improve?
Did fewer defects occur?
Did employees resolve customer problems faster?
Did utilisation increase?
Did risk decrease?
Did management make better decisions?
Did revenue increase?
Did margin improve?
Did something become possible that wasn’t economically viable before?
A model can move substantially up a benchmark without creating any additional value for a particular organisation.
Equally, a seemingly modest capability improvement could unlock an extremely valuable application inside another one.
Technical progress creates possibility. Business context determines value.
That distinction will become more important as frontier models improve.
More intelligence creates more opportunities to choose between#
There’s an interesting connection here to another problem.
Last month I argued that businesses won’t be able to pursue every AI opportunity they identify.
Reasoning models potentially make that problem larger.
If AI moves from generating outputs towards helping with analysis, problem-solving and decisions, the number of plausible applications expands again.
Look across an organisation and ask:
Where do people create documents?
You’ll find plenty of opportunities.
Now ask:
Where do people solve problems?
Almost everywhere.
Where do people compare alternatives?
Where do they investigate?
Where do they diagnose?
Where do they plan?
Where do they make judgements under uncertainty?
Where do they decide what happens next?
The potential surface area becomes enormous.
This doesn’t mean businesses should deploy reasoning models everywhere.
It means the opportunity-discovery problem becomes richer.
And the prioritisation problem becomes harder.
The cost of thinking matters too#
Reasoning also introduces another economic variable.
More computation isn’t free.
If a model spends greater computational effort working through a problem, businesses need to consider whether the value of the answer justifies the cost and latency involved in producing it.
This is already becoming visible in how reasoning models are developing.
OpenAI released o3-mini in January as a smaller reasoning model designed to offer strong performance, particularly in science, mathematics and coding, while reducing cost and latency compared with larger reasoning approaches.
That points towards an interesting future design question.
Not every problem deserves the most capable model.
A simple classification task doesn’t necessarily need expensive reasoning.
A high-consequence decision might.
Businesses may eventually need to become much more sophisticated about matching the cost of intelligence to the value of the problem.
That is an economic optimisation problem in its own right.
We shouldn’t use the most intelligence everywhere simply because we can.
We should use enough intelligence where the economics justify it.
From producing work to understanding problems#
Generative AI’s first wave made an extraordinary capability widely accessible.
People could describe what they wanted in ordinary language and software could produce something useful.
That alone has significant consequences.
But reasoning models suggest another stage is emerging.
AI systems may increasingly help people not merely execute a known task, but work through what should happen in the first place.
That shifts the business conversation.
From writing faster to understanding better.
From producing outputs to analysing situations.
From completing tasks to solving problems.
From answering questions to helping determine which questions matter.
And potentially, from saving the cost of work to improving the economic consequences of decisions.
There are still enormous limitations.
Reasoning models can be wrong.
They can lack crucial context.
Benchmark performance doesn’t guarantee business performance.
More reasoning can introduce additional cost and latency.
High-consequence decisions still require appropriate human judgement, controls and accountability.
But the direction deserves attention.
Because if AI becomes increasingly capable of helping businesses reason through difficult problems, the most interesting question may no longer be:
What work can AI produce for us?
It may become:
What problems are valuable enough to think about differently?
