Let me tell you a story

My fellow AI engineers 👋 – or whatever we’re calling ourselves these days. We went from Data Miner, to Data Scientist, to AI Engineer so quickly that it’s becoming hard to describe what we actually are.

Let me tell you a story. The first time I saw an LLM, GPT-2, to be precise, I really wasn’t that impressed!
I could see it was a meaningful step forward, but I thought it was just another NLP trend. We had already gone from Seq2Seq models for machine translation to increasingly capable autocomplete systems, so this felt like the next incremental step.

Then I changed my mind 🤔. GPT-3 was the real turning point. Not because it was flawless, it wasn’t, but because it fundamentally changed what a language model could be. Instead of training a dedicated model for every task, you could simply describe what you wanted in natural language. GPT-4 pushed this even further with a dramatic leap in reasoning and reliability. Then came GPT-5 with its 400k context window, making it practical to work over entire codebases, books, or complex technical documentation in a single conversation. Since late 2025, I don’t think I’ve written a single line of code entirely by myself, and I don’t think I’m an exception anymore...

But here’s the question nobody really answers: how did we go from a stochastic token predictor 🦜 to a tool that Terence Tao (maths genius) uses to do research in maths?


First of all: What is an LLM, really?

Let’s go back to basics. Strip away all the hype, an LLM is just: f(text_x) → text_y

That’s it. A function that takes text as input and returns text as output. Under the hood, it’s predicting the next token, one at a time, autoregressively. GPT-2 did this. GPT-4o does this. Claude does this. The fundamental mechanism hasn’t changed.

"The cat sat on the" → couch: 42%, floor: 28%, car: 8%...
                               ↓
                            "couch"

One token is selected, appended to the context, and the process repeats. Token after token, this simple loop generates conversations, code, and everything else we ask it to produce.

💡
So why does GPT-5 feel dramatically more capable than GPT-2 if the underlying mechanism is still next-token prediction?

Underlying #0 – Scale, baby, scale

A certain orange-haired 🍊 president once said, “Drill, baby, drill.” In AI, the equivalent isn’t oil, it’s compute and data. Every new generation of LLMs has been trained with more GPUs, more parameters, and better optimization than the previous one.

GPT-2 had 1.5 billion parameters. GPT-3 jumped to 175 billion: a 100x leap in a single generation. Frontier proprietary models like GPT-4 or Claude no longer disclose their parameter count. But open models give us a sense of the scale: Qwen3 reached 235 billion parameters, and Kimi K2 reached 1 trillion, and those are the ones we can actually verify…

What surprised researchers wasn’t that bigger models performed better. It was how predictably they improved. In 2020, OpenAI showed that language models obey remarkably simple scaling laws: as model size, data, and compute increase together, the training loss decreases according to power laws. Even more surprisingly, entirely new capabilities emerge as a consequence of this scaling. Models don’t simply become better at predicting the next word. They become surprisingly capable at tasks once reserved for human experts: writing, translating, coding, mathematical reasoning, etc.

Why does this happen? The honest answer is that we still don’t fully know. 🤯 Scaling laws are one of the strongest empirical results in modern deep learning, but there is still no accepted theoretical explanation for why they hold so consistently. The important takeaway is simple: scaling works. And for all these years, every generation of frontier models has largely been the result of pushing the same recipe further than anyone else.

✔️
Of course, scale isn’t the whole story. If it were, GPT-5 would just be a much larger GPT-2. It clearly isn’t. Let’s now look at the five key abilities that turned a simple next-token predictor into today’s AI assistants.

Ability #1 – Remember

The context window

Before we talk about agents, tools, and shiny things, we need to talk about something more fundamental: how much text can the function actually see at once?

Early models had much smaller context windows. GPT-2 had 1,024 tokens, while GPT-3 had 2,048. It’s roughly 1-3 pages of text. That’s not much. You could ask them to rewrite a piece of code or summarize an article, but not much more. That limitation shaped everything about how we built systems around LLMs. You couldn’t feed them a full document. You couldn’t keep a long conversation. You had to chunk, truncate, summarize, and stitch everything together by hand.

# Before: the chunking nightmare
def process_document(doc, max_tokens=3000):
    chunks = split_into_chunks(doc, size=max_tokens)
    summaries = []
    for chunk in chunks:
        summary = llm.complete("Summarize: " + chunk)
        summaries.append(summary)
    # Then summarize the summaries... and hope you didn't lose anything important haha
    return llm.complete("Combine summaries: " + "\n".join(summaries))

This wasn’t just inconvenient, it was architecturally constraining 🚧. The entire RAG ecosystem (e.g. embeddings, vector databases, retrieval pipelines) was built largely as a workaround for small context windows. If you could fit everything in context, you wouldn’t need to retrieve.

Now, context windows have gone from 2k to 128k to 1M tokens in today’s frontier models. One million tokens is roughly 750,000 words (or 500-1500 full pages of Letter/A4): an entire codebase, a year of Slack messages, or a full legal contract corpus.

Why does this matter? Because an LLM can only reason over what it can see. Increasing the context window doesn’t make the model inherently smarter but it definitely gives it more information to think with 🧠. Instead of asking “What is the answer based on this paragraph?”, we can now ask “What is the answer based on the entire repository?” The quality of the response can improve simply because the model no longer has to guess what happened outside its context.

This also changes how we design AI systems. Rather than spending engineering effort deciding what to retrieve, summarize, or discard, we can increasingly provide the raw source material directly. In many workflows, context engineering is replacing prompt engineering.

✔️
The context window is the most underrated architectural change in the LLM ecosystem. Most of what we call agent capabilities is just what happens when you give the function enough room to think.

Conversation history

Even with a large context window, maintaining a conversation across turns used to be your problem. You serialized the history, you managed the size and you sent it all back on every call – again and again 🔁 – manually.

# Before: manual context management
history = []
def chat(user_message):
    history.append({"role": "user", "content": user_message})
    response = llm.complete(messages=history)  # resend everything, every time
    history.append(response.message)
    return response.text
# If history exceeded the context window: good luck my friend 🤞

Now, some providers handle this with a thread or response ID. You pass a reference and not the full history.

# Now: stateful by reference
r1 = llm.complete(input="Hi, my name is Yassine")
r2 = llm.complete(input="Do you remember my name?", previous_id=r1.id)
# That's it.

This is a significant improvement because you no longer need to manage conversation state yourself. Instead of repeatedly serializing, transmitting, and deserializing the entire history, you simply reference the previous response. This reduces implementation complexity and network overhead.

Long-term memory

Conversation history lets the assistant use what was said earlier in the same thread. Long-term memory goes a step further: it allows the system to keep useful information even after that conversation is over.

The implementation can actually be quite simple. After a few interactions, or at the end of a session, the system can ask the LLM: “Is there anything here worth remembering later?” If there is, those useful facts are extracted, cleaned up, and stored so they can be reused in future conversations.

# Process:
After few interactions / each end of a session: "What is worth remembering?" → written to LONG_TERM_MEMORY.md
In future conversations: retrieve relevant memories → inject them back into the context

# In the LONG_TERM_MEMORY.md:
## Project context 
- API framework: FastAPI 
- Authentication: Keycloak 
- Database: PostgreSQL 

## Past incidents 
- API latency was previously caused by pgSQL connection pool exhaustion

A few days later, the engineer comes back with: “We are seeing latency again on the API.” The system can pull back the previous incident, restore the useful details into the context, and let the model continue from there. The important part is that the LLM did not remember anything by itself. The application decided what was worth keeping, stored it somewhere persistent, then brought it back when it became useful again.

💡
Think of long-term memory as: deciding what is worth keeping, then knowing when to bring it back.

RAG (Retrieval-Augmented Generation)

But even a 1M token context window remains limited. The same challenge applies to large or private knowledge bases that may change over time, such as internal documentation, legal contracts, and product catalogs. To generate relevant answers, you need to retrieve the right documents and inject them into the LLM’s context before generation.

Before, that meant building the entire pipeline yourself: chunk documents, embed them, store vectors, retrieve on query and inject into context.

# Before: the full RAG stack, hand-rolled
chunks = split(document)
embeddings = embed_model.encode(chunks)
vector_db.store(embeddings)

# At query time
relevant = vector_db.search(embed_model.encode(query), top_k=5)
answer = llm.complete(context=relevant, question=query)

Now, providers expose this as a native tool.

# Now: RAG as a native capability
kb = llm.create_knowledge_base(files=["contracts.pdf", "policies.pdf"])
answer = llm.complete(
    input="What does the contract say about termination?",
    tools=[file_search(kb)]
)

💪
RAG isn’t dead! It’s still the right architecture for large, private and dynamic corpora. But before, you had to build it. Now it’s an architectural choice, not a forced one.

Ability #2 – Speak Precisely

We saw that an LLM can be viewed as a function that generates text. The challenge is that its output is inherently unstructured and probabilistic. That’s ideal for conversations, but much less so for software systems. If you want to extract structured data from a document, insert records into a database, or pass the output to another service, you need a format that can be parsed deterministically.

Before, extracting structured data meant praying:

# Before: asking the model to behave
response = llm.complete(
'Extract info as JSON: Alice, 32, admin. Respond ONLY with valid JSON PLEASE.' )
try:
    data = json.loads(response.text)
except json.JSONDecodeError:
    # Was it wrapped in ```json? Did it add a comment? A missing comma?
    # Start your retry logic here, gooood luck haha, etc..
    ...

Nowadays, you define a schema. And the provider enforces it at generation time.

# Now: schema-enforced output
class UserSchema(BaseModel):
    name: str
    age: int
    role: str  # "admin" | "user" | "guest"

response = llm.parse(input="Extract: Alice, 32, admin", schema=UserSchema)
# Guaranteed: correct types, no missing fields, no malformed JSON
user = response.parsed  # ready to use

What’s actually happening under the hood? The provider converts your schema into a grammar, then applies constrained decoding. At each generation step, instead of sampling freely from its 100,000 possible tokens, let’s say, the model can only pick tokens that are valid given the current state of the schema. The invalid ones are simply masked out before sampling.

If the model just generated {"age":, and the schema says age must be a number, "thirty-two" is simply not an option anymore. The model can only continue with tokens that can produce a valid number.

As I said, there is no post-hoc validation here: this is a hard constraint during generation!

❌
Constrained decoding guarantees the shape, not the meaning. {"age": 847} is perfectly valid output. Business logic is still your job.

Ability #3 – Act

So far, so good! Our LLM can now reason, remember, and respond precisely. But there’s one fundamental limitation left: the model is blind to everything outside its context window and is totally passive. No live data, no external APIs, no actions. It can’t check the weather, query a database, or trigger a deployment. If the data isn’t in the training data or the prompt, it simply doesn’t exist to it.

Tool calling broke both walls at once 👊.

Declaring a tool

Before, giving the model access to an external function meant a prompt hack and a lot of faith:

# Before: prompt-based function calling
prompt = """
If you need the weather, write: CALL_WEATHER(city="CITY_NAME")
Otherwise respond normally.
"""
# Then parse the text output with regex
match = re.search(r'CALL_WEATHER\(city="(.+?)"\)', response.text)
if match:
    result = get_weather(match.group(1))
    # re-inject and call again...

Now, you declare tools as schemas. The model returns a structured call so no parsing and no prayers 🙏

# Now: structured tool declaration
tools = [{
  "name": "get_weather",
  "description": "Returns current weather for a city", # Prompt engineering
  "parameters": {"city": {"type": "string"}}
}]

response = llm.complete(messages=msgs, tools=tools)

if response.tool_calls:
    call = response.tool_calls[0]
    # call.name == "get_weather"
    # call.arguments == {"city": "Paris"}
    # In the runtime, you are responsible for DOING the tool call.
else:
    print(response.text)  # Model chose to answer directly

Under the hood? Three mechanisms, none of them fully documented by providers:

  1. Context injection: The tool definition is serialized and injected into the system prompt as text. The model reads it like any other instruction. This is why description matters as much as the schema because it’s what the model uses to decide whether to call the tool.
  2. Fine-tuning: Models are trained on millions of tool calling examples with specific output formats, likely using special tokens for each provider (<|tool_call|>, <function_calls>, etc.).
  3. Constrained decoding: Once the model decides to call a tool, arguments are generated under schema constraints as we saw in Ability #2.

💡
Treat your tool description like a system prompt because vague descriptions get misused. Be really precise with your tool definitions!

Native provider tools

Some tools are so common (think of web search, code execution, or document retrieval) that providers have built them in directly. You activate them, and the provider does the magic.

# Before: rolling your own web search
results = bing_api.search(query)
context = format_results(results)
answer = llm.complete(f"Context:\n{context}\n\nQuestion: {query}")

# Now: one line
answer = llm.complete(input=query, tools=[web_search()])

❌
The trade-off is this: you lose observability. The provider decides what to search and you don’t see the raw results. If you need control or custom retrieval logic, go back to explicit tool calling.

MCP: The standard for tool discovery

Before MCP, every tool integration was bespoke🤵‍♂️. Each application had to define its own tool schemas, authentication, error handling, execution layer, etc. Thus, integrating the same tool across multiple projects often meant rewriting the same plumbing over and over again.

MCP (Model Context Protocol) standardizes this interface. Instead of hardcoding tool definitions into every application, an MCP server exposes them through a common protocol. Any MCP-compatible client can connect, discover the available tools at runtime, and invoke them without knowing in advance what the server provides.

Think of it as USB for LLM tools, not because it distributes tools, but because it defines a common interface between tool providers and tool consumers.

# Before MCP: hardcoded tool declarations per project
tools = [tool_a, tool_b, tool_c]  # written by hand, every time

# Now with MCP: discovered at runtime
tools = mcp_client.list_tools()  # whatever the server exposes today

What’s actually happening under the hood?

AT STARTUP 🚀

🔌 MCPClient connects → initialize()
    🤝 handshake + protocol / transport negotiation (HTTP or stdio)
    🔍 list_tools() → discover available tools at runtime
    🧰 passed to the LLM as tools=[...], just like any standard tool declaration

AT INFERENCE ⚡

🧠 LLM decides → tool_call("fetch_data", {"id": "123"})
🏃 Runner → MCPClient.call_tool("fetch_data", {"id": "123"})
         📡 JSON-RPC over the negotiated transport
         🛠️ MCP Server executes the tool and returns the result
🏃 Runner → 📥 injects the result into the context
🧠 LLM continues generation

💪
The runner doesn’t know what the server exposes until it asks. That’s precisely its value, because it decouples tool availability and usage from the LLM.

Skills

A skill packages procedural knowledge: how the agent should perform a recurring task. It is typically stored in a SKILL.md file containing metadata, instructions, and optionally references to scripts or other resources.

Under the hood, the agent does not need every skill fully loaded into context. It can first see lightweight metadata such as each skill’s name and description, then load the detailed SKILL.md only when the task requires it.

Tool → gives the model the ability to act on the external world
Skill → gives the model additional know how for how to perform a task

This pattern is now used across agent frameworks and providers such as Anthropic and OpenAI.

⚠️
Something really important! At the API level, you manipulate separate objects such as messages, instructions, tools, and skills. But they are ultimately serialized into a single model input: the text_x in our original f(text_x) → text_y. This input may use special control tokens and a specific formatting scheme for each provider, but fundamentally it is still just text.

Ability #4 – Reason & Orchestrate

We already know that older LLMs, or base LLMs, answer once and stop. They’re one shot functions: no iteration, no self correction and no multi step planning. The problem is that real world tasks rarely fit into a single question response. Let’s take: “Analyze our Q3 sales, cross reference them with market trends, and flag the three biggest risks.” That’s hard to do in one shot. It’s actually a workflow made of several complex tasks 🔁.

To turn the LLM from a responder into an actor, we need to give it two things: the ability to reason before acting, and the ability to chain that reasoning across multiple steps.

1️⃣ Reason – Making the LLM reason inside a single call

The simplest form of prompting for reasoning requires no special API feature: just ask the model to think out loud 🎤.

Chain of Thought (CoT): “Think step by step and give your final answer at the end” is the classic example for a CoT. The model writes its reasoning as part of its output. It’s just tokens, but structured tokens that force the model to work through the problem before committing to an answer. The reasoning IS the output. You see every step.

User → LLM("think out loud step by step, solve X")
     → "First, I need... Then I should... Therefore the answer is..."
     → User
# 1 call, reasoning is visible in the response.

Extended thinking (Hidden CoT): Some models (o3, Claude, DeepSeek R1) go further. They dedicate an internal token budget to reasoning before producing the final answer. The reasoning is no longer part of the output because it happens upstream, invisibly. Still 1 API call for 1 LLM call. The thinking is just longer, and sometimes exposed.

User → LLM → [internal reasoning tokens, hidden] → final answer → User

# 1 call, the model "thinks" before answering.

The transparency varies wildly by provider:

  • Claude: Provides thinking block response.thinking , on current models, these can contain summarized thinking.
  • DeepSeek R1: Same philosophy, the reasoning is exposed inside <think>...</think> before the answer.
  • OpenAI o3: The reasoning happens, you pay for it, but you only get a summary in a response variable called response.reasoning_summary.

There is also a notable variant in reasoning, the Tree of Thoughts 🌳, where the runner explores multiple reasoning branches in parallel instead of a single linear path, then evaluates and keeps only the most promising ones before continuing.

❌
With OpenAI, reasoning tokens are billed as output tokens, even if you see none of it. On a complex problem, that can mean thousands of invisible tokens on your bill.

2️⃣ ReAct – Orchestrating tools

CoT is powerful, but the model is still isolated; it can’t fetch live data mid-reasoning, can’t act on the world. ReAct (Yao et al., 2022), which stands for Reason 🤔 and Act ⚒️, fixes this by interleaving reasoning and acting in a loop:

Reason → Act → Observe → Reason → Act → Observe → ... → Answer

Reason  : LLM thinks about what to do next
Act     : LLM calls a tool (search, database, API, etc...)
Observe : Runner executes the tool & injects the result back
Repeat  : LLM reasons again with new information

Each iteration is a separate LLM call. The LLM doesn’t remember previous steps on its own, so the runner maintains state and re-injects history at every call.

And under the hood? Three mechanisms make the ReAct pattern work, although providers only document part of the implementation:

  • Prompting: Before the inference, the runner injects a system prompt that teaches the model how to solve the task in a CoT mode: reason step by step and not first-shot, use the available tools whenever they reduce uncertainty, wait for tool results before planning the action and only return a final answer once enough information has been gathered.
  • ReAct loop: Each iteration is like a brand new inference. The model reasons over the current context, either emits a tool call or a final answer, then stops. If a tool is requested, the runner executes it, injects the observation back into the conversation, and asks the model to reason again from this updated state.
  • Stateless: The model has no execution state between calls. The runner recreates it by replaying the system prompt, user request, previous assistant messages, tool calls and tool outputs before every inference. How providers compress or summarize long conversations to fit the context window remains largely undocumented.

Wait a minute! A ReAct loop does not always finish in one uninterrupted run. The runtime can persist the agent’s state between steps using checkpoints, allowing execution to pause, wait for human approval, recover from a failure, or resume later from where it stopped.

Reason → Act → Checkpoint → Wait → Resume → Act → Answer

💡
ReAct is not a library. It’s a pattern. LangChain, OpenAI SDK, LangGraph; they all implement variations of this same loop. Understanding ReAct means understanding what every agent framework does under the hood.

3️⃣ Multi-agents: 1 or more ReAct loops for the same goal

This is where things get much more open-ended. There is no single recipe anymore. You can use the ReAct loop we just saw, spawn 10 agents in parallel and keep the best answer, or create a specialized subagent whenever the main agent hits a particularly difficult problem.

That subagent can therefore investigate the issue with its own context, memory and tools, then return its findings to the parent agent. See more about this pattern in LLM Engineering: Context, Reasoning, and Agentic Systems.

✔️
Yes you got it! Agentic is less a specific architecture than a design space. The question is no longer “Which agent pattern are we using?” but “How much autonomy, delegation and orchestration do we want to give to each LLM in the system?”

In practice – before and after

Before agent SDKs, you implemented the ReAct loop yourself.

# Before: artisanal ReAct loop
messages = [{"role": "user", "content": user_input}]
while True:
    response = llm.complete(messages=messages)
    messages.append(response.message)
    if "CALL_WEATHER" in response.text:  # parsing text with regex 😬
        result = get_weather(extract_city(response.text))
        messages.append({"role": "user", "content": f"Observation: {result}"})
    else:
        return response.text  # hope it's actually done...
# Reason → Act → Observe → Repeat. All by hand.

Now, the SDK implements the ReAct loop for you.

# Now: declare intent, let the runner orchestrate the ReAct loop
agent = Agent(
    instructions="You analyze data and answer questions.",
    tools=[get_weather, query_db, search_web]
)
result = agent.run("Analyze last quarter's sales and flag any risks.")

🏖️ SANDBOXING – Once an LLM can actually act on external systems, permissions become part of the architecture. You do not want an agent to freely access production, secrets, or resources belonging to another agent. That is why agent runtimes typically execute actions inside controlled environments, with RBAC/IAM-style permissions defining what each agent can read, modify, execute, or access. The goal is simple: give the agent enough freedom to do the job, without giving it the keys to everything.

✔️
And that’s really what changed with agentic systems: instead of hard-coding one workflow from start to finish, we can let several ReAct loops investigate, delegate, come back with results, and keep pushing the task forward. And that the real magic…

Ability #5 – See & Hear

Our LLM has one last flaw: it talks, talks a lot, sometimes too much, but all it does is talk. Images, audio, and documents had to be translated into text before the model could process them. That is what multimodality changes.

Before
image → vision model → text ↘
audio → speech model → text → LLM → text
document → parser → text    ↗

Now
text     ↘
image    →  Multimodal model → text / image / audio
audio    ↗
document ↗

Before, each modality had to go through its own pipeline to be converted into text, and every conversion meant some information was lost.

# Before: three pipelines AND three failure modes

# Image: OCR → text → LLM
text = ocr.extract(image)
answer = llm.complete("What's in this chart?\n" + text)
# Lost: layout, colors, spatial relationships 📉

# Audio: transcription → text → LLM
text = whisper.transcribe(audio)
answer = llm.complete("Summarize this meeting:\n" + text)
# Lost: tone, speaker identity, hesitations 📉

# Document: text extraction → LLM
text = pdf_parser.extract("insurance_contract.pdf")
# You get: "Dental coverage • • • •"
# Lost: which bullet means "included", which column maps to which tier,
#       the visual hierarchy, the color coding 📉

Now, you pass the raw modality and the model perceives it natively and not through a text proxy:

# Now: one call WITH native perception
response = llm.complete(
  input=[
    image("chart.png"),  # sees layout, colors, spatial structure
    audio("meeting.mp3"), # hears tone & speaker identity
    document("insurance.pdf"), # reads the viz table & hierarchy
    ]
)

Why is this important?

Generated by DALL.E

Imagine this table is buried in a PDF from an insurance company you are working with as an AI Engineer. Your LLM has to understand that a colorful dot at the intersection of a row and a column means that a specific guarantee is included in a specific plan. If the layout is lost, the meaning is lost with it. A pure OCR pass may recover the text and symbols, but not the semantic relationships between them. That is exactly what multimodality is about.

What’s actually happening under the hood? Each modality is converted into tokens that the attention mechanism can process alongside text tokens. An image becomes a grid of patch embeddings. Audio becomes a sequence of spectral features. A PDF can be rendered visually and processed alongside extracted text. The model sees what a human sees and not what a text extractor outputs.

The key idea is simple: the transformer does not care where the tokens came from. Once encoded as tokens, text, image patches, and audio all pass through the same attention layers.

✔️
The multimodal shift isn’t a feature addition. It’s a redefinition of the LLM function as a more diverse one.

Wrapping up

We started by approximating a generative LLM as a function as simple as this: f(text_x) → text_y, predicting the next token, one token at a time.

But look at what we unpacked across these five abilities. Some of them came from the model itself like a bigger context window, constrained decoding, native multimodality, emergent reasoning. The model got genuinely better, and scaling is largely why. Others came from what we built around it: tool calling, RAG pipelines, agent loops, MCP. That’s the harness (the hot word of the moment!), everything in the AI system except the model itself.

And harness engineering is quickly proving to be as important as model engineering. Two teams can deploy the exact same LLM and achieve completely different outcomes depending on how their harness is designed. So yes, the stochastic parrot is still there 🦜. We just figured out how to put it to work. 😃

👉 The harness is one thing, but who actually built the model underneath it? Go open the tomb🏺: Opening the Tomb: Who’s Really Buried Inside ChatGPT?