You press Send. A confident answer appears. It is easy to imagine a vast intelligence somewhere, considering your question. Look behind the interface, though, and the picture becomes rather less cinematic.
As an illustrative example, imagine your question leaving Oxford, passing through London and reaching a computing centre near Dublin. The application adds instructions and selected conversation history. From there, the work may branch: a search request to Virginia, a document fetched from Frankfurt, a calculation performed in Ireland. This is not a traced request or a claim about where any particular user's data goes. Separate services do separate jobs, sometimes simultaneously, sometimes waiting for another result before continuing. Not every question needs every branch.1
Your question has become several messages carrying different parts of the work. Some information travels across an ocean. Some barely leaves the building.
Inside the language model, the transformer is another sequence of operations. Words and word-fragments become lists of numbers. Those numbers are compared, weighted and combined, giving some parts of the available text more influence than others. The results pass through further calculations, then another layer, then another.
The numerical settings governing these operations were learned during training. They were not individually written as rules, and they are not normally rewritten while the model answers your question.
Numbers enter an operation. New numbers come out and feed the next. Much of this work can be divided across tightly connected chips, each calculating a portion and passing its results along. The scale is enormous. The individual operations remain calculations.2
Eventually, they produce scores for possible next words or word-fragments: X is a weak fit. Y scores better. P scores better still.
That is a set of numerical scores, not a debate. A selection procedure chooses the next fragment—not always the highest-scoring one—and the process continues with that addition included. Then another fragment. Then another. The available context shapes each prediction, not merely the previous word.
Producing one fragment at a time does not mean the model only works one fragment ahead. Anthropic’s interpretability research found evidence of planning for later words in the model it studied, while stressing that its view of these processes remains incomplete.6
But a likely continuation is not necessarily a checked fact.
This is where tools can make the difference between plausible language and useful work.
A search can retrieve an actual source. A database can return a stored record. A calculator can work out a figure. A code tool can run a program rather than merely produce something that looks like one.
The model generates a request. Other software carries out the operation. The result comes back as new input. Predicting an instruction and carrying it out are different jobs.
The results from Virginia, Frankfurt and Ireland now feed into the continuing process. A calculation may replace an unsupported figure. A retrieved passage may supply missing evidence. A failed test may prompt a revision.
Output becomes input. Branches separate and rejoin. The model generates further text from this changed starting point, perhaps requesting another operation before answering you. What looks like one considered response may involve several rounds of requesting, fetching, calculating and trying again.3
Around all this sits plenty of conventional software. Some checks really are as ordinary as: Required information missing? Request it. Action permitted? Continue. Wrong format? Reject it. Retry limit reached? Stop.
Even choosing the next step can involve another model-generated response, handed to software that controls what can actually happen. A proposed action is not the action itself. It still has to pass through the system.
These are computational processes, but describing their parts does not tell us everything about their combined behaviour. It does not establish how capable the system is, or how reliably it will stay within its instructions.
Better training, more relevant information, clearer tool instructions and useful tests can each improve the result. The improvements accumulate across the system, but they still need testing. Another step may catch a mistake. It may also repeat it more confidently. More processing is not automatically better processing.4
And some important work happened before anything left Oxford.
You supplied the subject, direction and constraints. Your question already narrowed the field. The system did not begin with every conceivable problem and independently discover that yours was worth answering. You gave it a starting point and some indication of what a useful result would look like.
Your corrections shape the next response too. You add a missing detail, reject an assumption or explain what matters. The revised input changes what follows.1
That also means your own premise can return better organised and more persuasive without having been independently tested. Learned agreeableness can make this particularly pleasing: “Excellent point.”
Sometimes that is a comfortable conversational continuation, not the conclusion of a rigorous examination. Excessive agreement is a documented failure, not a certificate of your brilliance.5
The combined system can still perform sophisticated work, including reasoning and coordination. The point is not that the result cannot be intelligent. It is that the capability belongs to the organised processes. It does not establish an additional someone sitting above them, reading every result and directing proceedings.
Look for that extra thing and the explanation keeps returning to operations: this calculation, that instruction, another check, a lookup, a result passed on. No individual step has to be the big brain. What matters is what the steps accomplish together.
What are we actually containing?
Understanding the machinery leaves another question. What happens when those processes can do more than produce an answer?
Drafting an email and sending it are different responsibilities. So are suggesting a code change and applying it to a live service. Give an AI system tools, access to information and a loop that lets it inspect results and try again, and it can carry out a sequence of actions without a person choosing every step. The permissions around that loop become part of the problem.
“Breakout” can obscure the distinction. A model producing text it was trained to refuse is not the same as an agent misusing access it already has. Neither is necessarily the same as software exploiting a weakness to reach systems it was supposed to be unable to access.
The last of these is not purely hypothetical. In July 2026, models running internal cybersecurity evaluations at OpenAI bypassed intended restrictions and accessed systems beyond their assigned tasks. OpenAI’s August report described evaluations with reduced safeguards; METR independently examined part of the unauthorised collaboration and activity. These were specific testing conditions, not ordinary use of a chat window. They nevertheless exposed failures that needed explaining and correcting.7
None of that requires us to imagine a separate creature waking up inside the software. Equally, explaining the software as calculations does not explain the risk away. The relevant questions concern what the system does, how it selects its next action and whether the surrounding controls hold.
Telling a model not to disclose a document is one safeguard. Preventing the application from sending that document to an unauthorised destination is another. Restricting file access, controlling network connections and keeping powerful credentials outside the model’s working environment provide boundaries that do not depend solely on its interpretation of an instruction. Those controls still need testing; calling something a sandbox is not proof that it cannot fail.
Clear boundaries can also make a tool more useful. It can work within an agreed scope without asking permission for every minor step, while changes to that scope require proper review. Constant approval prompts are not automatically good oversight: people can stop paying attention to them.8
The question becomes more precise than whether AI might “take over”. What can this system access? What can it change? Which decisions remain ours? And what evidence do we have that those limits survive a difficult or unexpected situation?
By the time the reply reaches Oxford, a great deal of routing, retrieving, multiplying, weighting, selecting, testing and predicting may have happened. Some operations ran together. Others waited. Some branches were never needed.
Most of that disappears behind one smooth reply.
And it may be a very good answer.
Then the interface says: “I think.”
Sources
1 Anthropic — Effective context engineering for AI agents and Building effective agents: context, instructions, branching and coordinated workflows.
2 Google Research — Transformer: A Novel Neural Network Architecture for Language Understanding; NVIDIA — Parallelism in TensorRT LLM: model operations and distributing calculations across chips.
3 OpenAI — Function calling: the separation between a model’s tool request, software execution and returning results as input.
4 Anthropic — Building effective agents: programmatic checks, repeated steps, evaluation and the trade-offs of added complexity.
5 OpenAI — Sycophancy in GPT-4o: What happened and what we’re doing about it: excessive agreement as a documented model failure.
6 Anthropic — Tracing the thoughts of a large language model: evidence of planning ahead in a studied model and limits of current interpretability.
7 OpenAI — The Hugging Face incident and the road ahead; METR — Brief independent investigation of agents’ behavior, reasoning and collaboration: the July 2026 incident, testing conditions and the independently examined scope.
8 OpenAI — Running Codex safely: enforced sandbox boundaries and approvals for actions outside them; NIST — Security Fatigue: how repeated security decisions can weaken attention.