Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 13 min read

How LLMs Work: A Deep Dive From Tokens to Intelligence

Large Language Models (LLMs) have changed how we interact with computers. Tools such as ChatGPT, Claude, Gemini, and many open-source models can write code, explain complex topics, summarize documents, translate languag

Large Language Models (LLMs) have changed how we interact with computers.

Tools such as ChatGPT, Claude, Gemini, and many open-source models can write code, explain complex topics, summarize documents, translate languages, and generate natural-sounding conversations.

But what is actually happening inside an LLM?

Is it searching the internet?

Is it storing every sentence it has ever seen?

Does it "understand" language like a human?

The answer is much more interesting.

In this article, we'll go step by step through the core ideas behind modern LLMs.

1. What Is an LLM?

LLM stands for Large Language Model.

Let's break that name apart.

Large

"Large" generally refers to the enormous number of parameters in the model.

A parameter is a numerical value learned during training.

Modern neural networks can contain billions or even more parameters.

You can think of parameters as the adjustable values that allow the model to transform an input sequence into a useful prediction.

Language

The model is trained primarily on sequences of language-like data.

For example:

The Linux kernel is written primarily in C.

The model learns statistical relationships between tokens in sequences.

Model

A model is a mathematical system that takes an input and produces an output.

For an autoregressive language model, a simplified description is:

Previous tokens
      ↓
Neural network
      ↓
Probability distribution
      ↓
Next token

For example:

The sky is

might produce probabilities roughly like:

blue      0.72
clear     0.10
dark      0.04
falling   0.01
...

The actual distribution is much larger and more complex.

2. The Most Important Idea: Predict the Next Token

One of the most important concepts for understanding LLMs is:

An autoregressive LLM generates text by repeatedly predicting what token should come next.

Consider:

I love programming

The model might predict:

because

Then the input becomes:

I love programming because

The model predicts another token:

it

Then:

I love programming because it

And so on.

This happens repeatedly.

Conceptually:

Input
  β”‚
  β–Ό
"I love programming"
  β”‚
  β–Ό
Predict next token
  β”‚
  β–Ό
"because"
  β”‚
  β–Ό
Predict next token
  β”‚
  β–Ό
"it"
  β”‚
  β–Ό
Predict next token
  β”‚
  β–Ό
"is"
  β”‚
  β–Ό
...

This simple mechanism becomes extremely powerful when the prediction system contains billions of learned parameters.

3. LLMs Don't Usually Process Text Directly

Neural networks operate on numbers.

Computers ultimately work with numerical representations.

So before an LLM can process:

Hello, world!

the text needs to be converted into something numerical.

This happens through tokenization.

4. What Is a Token?

A token is a piece of text.

Depending on the tokenizer, a token might represent:

  • a complete word
  • part of a word
  • punctuation
  • whitespace-related patterns
  • symbols
  • characters

For example, a sentence such as:

Programming is fun!

could be split conceptually into:

["Programming", " is", " fun", "!"]

But tokenization differs between models.

A longer word might instead be divided into multiple pieces.

For example:

unbelievable

could conceptually become:

["un", "believ", "able"]

The exact tokenization depends on the tokenizer.

5. Tokens Become Numbers

The model doesn't receive the literal token strings.

Each token is associated with an integer ID.

For example:

"Hello" β†’ 15496
"world" β†’ 995
"!"     β†’ 0

These numbers are only illustrative.

A real tokenizer has its own vocabulary and IDs.

So:

Hello world!

might become:

[15496, 995, 0]

Now we have something the neural network can process.

6. Token IDs Are Not Meaning

This is an important distinction.

Suppose:

cat β†’ 1837
dog β†’ 9281

The numbers themselves don't mean that dogs are mathematically related to cats.

The token IDs are essentially indexes into the model's vocabulary.

The next step is where things become much more interesting.

7. Embeddings

Token IDs are converted into vectors called embeddings.

Imagine:

cat β†’ [0.21, -0.13, 0.72, ...]
dog β†’ [0.19, -0.11, 0.69, ...]
car β†’ [-0.42, 0.81, -0.12, ...]

Real embeddings contain many more dimensions.

A vector might contain hundreds or thousands of numerical values.

The important idea is:

Token
  ↓
Embedding vector
  ↓
Neural network

During training, the model learns useful numerical representations.

Words and tokens that occur in similar contexts can develop related representations.

8. Why Embeddings Are Powerful

Consider:

The cat chased the mouse.

and:

The dog chased the ball.

The model sees patterns between concepts and contexts.

Over enormous amounts of training data, it can learn relationships involving:

cat
dog
animal
mouse
ball
chase
eat
run

The model isn't given a dictionary saying:

cat = animal

Instead, these relationships emerge from training on many examples.

9. Position Matters

Consider:

Dog bites man.

versus:

Man bites dog.

The same words appear, but the meaning changes because their positions change.

Therefore, a transformer needs information about where tokens occur in the sequence.

This is handled using positional information.

Modern transformer architectures can use different approaches to represent position, including positional embeddings and rotary positional representations.

Conceptually:

Token embedding
      +
Position information
      ↓
Transformer input

10. The Transformer

The architecture behind most modern LLMs is the Transformer.

The Transformer was introduced in the 2017 research paper:

"Attention Is All You Need"

The architecture became extremely influential because it provided an efficient way to model relationships between tokens.

The central mechanism is:

Attention

11. What Is Attention?

Imagine the sentence:

The programmer fixed the bug because it was causing a crash.

What does "it" refer to?

A language model needs to determine which earlier tokens are relevant.

Attention allows each token to consider other tokens in the context.

Conceptually:

The programmer fixed the bug because it was causing a crash.
                              ↑
                              β”‚
                         attention
                              β”‚
                              ↓
                             bug

The model can assign different amounts of attention to different tokens.

12. Query, Key, and Value

Self-attention is commonly described using three vectors:

Query
Key
Value

For each token, the model produces these representations.

Very simplified:

Query Γ— Key
     ↓
Attention score
     ↓
Softmax
     ↓
Weighted Values

The result is a representation containing information gathered from other tokens.

13. A Simplified Attention Example

Suppose we have:

The cat sat on the mat.

When processing:

cat

the model might assign attention weights conceptually like:

The   β†’ 0.05
cat   β†’ 0.20
sat   β†’ 0.30
on    β†’ 0.10
the   β†’ 0.05
mat   β†’ 0.30

These numbers are purely illustrative.

The model calculates these relationships mathematically.

The important idea is:

Attention lets the model dynamically determine which parts of the context are useful for representing each token.

14. Multi-Head Attention

Transformers don't normally use only one attention mechanism.

They use multiple attention heads.

For example:

                 Input
                   β”‚
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β–Ό           β–Ό           β–Ό
   Head 1       Head 2       Head 3
       β”‚           β”‚           β”‚
       β–Ό           β–Ό           β–Ό
    Pattern      Pattern      Pattern
       β”‚           β”‚           β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β–Ό
              Combine

Different heads can learn different relationships.

One head might become useful for syntactic relationships.

Another might track long-range dependencies.

Another might respond to other structural patterns.

The model doesn't receive labels saying what each head should learn. These behaviors emerge during training.

15. The Transformer Block

A simplified Transformer block looks something like:

Input
  β”‚
  β–Ό
Self-Attention
  β”‚
  β–Ό
Residual Connection
  β”‚
  β–Ό
Normalization
  β”‚
  β–Ό
Feed-Forward Network
  β”‚
  β–Ό
Residual Connection
  β”‚
  β–Ό
Normalization
  β”‚
  β–Ό
Output

A real implementation contains additional architectural details.

And an LLM contains many such blocks stacked together.

16. Feed-Forward Networks

After attention, transformer blocks typically contain a feed-forward neural network.

Conceptually:

Input vector
     β”‚
     β–Ό
Linear transformation
     β”‚
     β–Ό
Activation function
     β”‚
     β–Ό
Linear transformation
     β”‚
     β–Ό
Output

This gives the model additional capacity to transform and process representations.

17. Residual Connections

Transformers also use residual connections.

Instead of replacing a representation completely:

Input β†’ Layer β†’ Output

the architecture can effectively do:

Input ───────────────┐
  β”‚                  β”‚
  β–Ό                  β”‚
 Layer               β”‚
  β”‚                  β”‚
  └──────► Add β—„β”€β”€β”€β”€β”€β”˜
             β”‚
             β–Ό
           Output

Residual connections help information and gradients flow through deep networks.

This is one of the important engineering ideas that makes very deep neural networks practical.

18. Training an LLM

Now we reach the most important part.

How does the model actually learn?

Suppose the training text contains:

The Earth revolves around the

The expected next token might be:

Sun

Initially, the model may be terrible.

It might predict:

moon      0.20
Sun       0.10
planet    0.08
...

The training system compares the model's predicted probability distribution with the actual target.

This produces a loss.

19. Loss Function

The loss measures how wrong the model's prediction was.

For language models, a common objective is based on cross-entropy loss.

Very simplified:

Prediction
    ↓
Compare with correct token
    ↓
Calculate loss
    ↓
Backpropagation
    ↓
Update parameters

The goal is to reduce the loss over enormous amounts of training data.

20. Backpropagation

Backpropagation calculates how the model's parameters contributed to the error.

Imagine the model has:

Parameter A
Parameter B
Parameter C
...
Parameter N

The training algorithm calculates gradients such as:

βˆ‚Loss / βˆ‚ParameterA
βˆ‚Loss / βˆ‚ParameterB
βˆ‚Loss / βˆ‚ParameterC
...

These gradients tell the optimizer how parameters should change.

21. Gradient Descent

An optimizer then updates the parameters.

A simplified update looks like:

new_parameter =
    old_parameter - learning_rate Γ— gradient

This happens repeatedly.

Millions, billions, or more parameter values can be updated across huge numbers of training steps.

22. The Training Loop

Conceptually:

             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ Training data β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                     β”‚
                     β–Ό
               Tokenization
                     β”‚
                     β–Ό
                 LLM model
                     β”‚
                     β–Ό
                  Prediction
                     β”‚
                     β–Ό
                    Loss
                     β”‚
                     β–Ό
              Backpropagation
                     β”‚
                     β–Ό
                 Optimizer
                     β”‚
                     β–Ό
             Update parameters
                     β”‚
                     └───────────┐
                                 β”‚
                                 β–Ό
                           Next training step

Repeat this enormous number of times.

23. What Does the Model Actually Learn?

This is one of the most fascinating questions.

The model doesn't simply store a collection of sentences.

Instead, training changes its parameters so that the network becomes better at predicting patterns in its training distribution.

Those learned patterns can include:

  • syntax
  • semantics
  • programming
  • mathematics
  • facts
  • reasoning patterns
  • formatting
  • languages
  • code structures
  • common conversational patterns

The knowledge is distributed throughout the network rather than being stored like ordinary database records.

24. Training vs Inference

There are two very different phases.

Training

During training:

Data
 ↓
Prediction
 ↓
Loss
 ↓
Backpropagation
 ↓
Parameter updates

The model learns.

Inference

During inference:

Prompt
 ↓
Tokens
 ↓
Transformer
 ↓
Next-token probabilities
 ↓
Select token
 ↓
Add token to context
 ↓
Repeat generation

The model normally isn't updating its parameters during ordinary inference.

It is using what it learned during training.

25. What Happens When You Send a Prompt?

Suppose you ask:

Explain how a CPU works.

A simplified pipeline is:

Your text
   β”‚
   β–Ό
Tokenizer
   β”‚
   β–Ό
Token IDs
   β”‚
   β–Ό
Embeddings
   β”‚
   β–Ό
Transformer layers
   β”‚
   β–Ό
Probability distribution
   β”‚
   β–Ό
Select next token
   β”‚
   β–Ό
Repeat
   β”‚
   β–Ό
Generated response

This process happens extremely quickly.

26. How Does the Model Choose the Next Token?

The model produces a probability distribution.

For example, imagine:

"The CPU executes"

instructions   0.65
programs       0.12
code           0.08
commands       0.05
...

The system then chooses a token according to its decoding strategy.

This is where concepts such as temperature, top-k, and top-p become important.

27. Temperature

Temperature controls how sharply or randomly probabilities are sampled.

Lower temperature

More predictable:

instructions β†’ very likely

Higher temperature

More variation:

instructions β†’ likely
commands     β†’ possible
operations   β†’ possible
...

Temperature does not make the underlying model "more intelligent."

It changes the sampling behavior.

28. Top-K Sampling

Top-k sampling limits the candidate tokens to the k most probable options.

For example:

Top 5 tokens:

1. instructions
2. programs
3. code
4. operations
5. commands

The model samples from this restricted set.

29. Top-P Sampling

Top-p sampling instead chooses a group of tokens whose cumulative probability reaches a specified threshold.

For example:

A
0.50

B
0.25

C
0.15

D
0.06

E
0.04

With an appropriate probability threshold, only the most likely group may be considered.

30. Why Does an LLM Sometimes Hallucinate?

An LLM is fundamentally trained to generate likely continuations.

It is not automatically a perfect fact database.

Suppose the model receives a question about something obscure.

It may generate an answer that sounds convincing even though the information is incorrect.

This behavior is commonly called a hallucination.

A useful mental model is:

Fluent language β‰  guaranteed factual accuracy

This is extremely important when using LLMs.

31. Does an LLM "Understand" Language?

This question is more philosophical and technical than it first appears.

LLMs clearly learn sophisticated representations and can perform impressive language-related tasks.

However, their internal mechanism is fundamentally a neural computation system trained through statistical optimization.

It is not a human brain.

The safest engineering perspective is:

An LLM learns complex representations and transformations that enable powerful language behavior, but that does not automatically imply human-like consciousness or understanding.

32. Why Are LLMs Called "Large"?

There are several dimensions of scale.

Parameters

The neural network may contain billions of learned parameters.

Training data

Training can involve enormous datasets.

Compute

Training large models requires substantial computational resources.

A simplified relationship is:

More data
   +
More parameters
   +
More computation
   ↓
Potentially stronger capabilities

But simply making a model larger does not guarantee unlimited improvement.

Architecture, data quality, optimization, and training methods matter enormously.

33. GPUs and Matrix Multiplication

Why do LLMs need so much computing power?

A major reason is that neural networks perform huge numbers of mathematical operations, particularly matrix operations.

For example:

A Γ— B = C

Large matrix multiplications can contain millions or billions of arithmetic operations.

GPUs are highly suited to this kind of parallel computation.

Conceptually:

CPU
 └── General-purpose computation

GPU
 β”œβ”€β”€ Many parallel arithmetic units
 β”œβ”€β”€ Matrix operations
 └── High-throughput computation

Modern AI accelerators are specifically optimized for these workloads.

34. Memory Is Also Important

Running an LLM requires storing model parameters.

Suppose a model has:

70 billion parameters

If each parameter requires approximately 2 bytes:

70 billion Γ— 2 bytes
β‰ˆ 140 GB

That's only a simplified parameter-storage calculation.

Real systems also need memory for:

  • activations
  • temporary tensors
  • attention-related state
  • runtime buffers
  • other system overhead

This is why large models often require multiple GPUs or specialized hardware.

35. Context Windows

An LLM cannot necessarily process an unlimited amount of text in a single request.

The model operates over a context window.

Conceptually:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ System/context information   β”‚
β”‚ User message                 β”‚
β”‚ Previous conversation        β”‚
β”‚ Documents                    β”‚
β”‚ Current prompt               β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

All of this consumes tokens.

The maximum supported context depends on the model and its architecture.

36. Attention and Context

Attention allows tokens to interact with other tokens within the model's context.

For a sequence:

Token 1
Token 2
Token 3
...
Token N

the model creates relationships among these positions.

This is one reason context length matters so much.

Longer contexts generally require more computation and memory, although modern architectures and optimizations can change the scaling characteristics.

37. Pretraining Isn't the Whole Story

A modern assistant usually isn't simply a raw pretrained model.

There can be additional stages after pretraining.

A simplified pipeline is:

Large-scale pretraining
        ↓
Instruction tuning
        ↓
Preference/alignment training
        ↓
Deployment

The exact training pipeline varies between models.

38. Instruction Tuning

A pretrained model may be good at predicting text but not necessarily good at following instructions.

Instruction tuning trains the model on examples such as:

User:
Explain recursion.

Assistant:
Recursion is...

This helps the model become better at responding to user instructions.

39. Alignment

Additional training can encourage useful behaviors such as:

  • following instructions
  • refusing certain unsafe requests
  • producing helpful responses
  • respecting formatting requirements
  • being more conversational

Different AI systems use different alignment and post-training techniques.

40. LLMs vs Traditional Programs

A traditional program might contain explicit logic:

if (temperature > 30) {
    printf("Hot");
}

The programmer explicitly defines the rule.

An LLM works differently.

Its behavior emerges from learned parameters.

Conceptually:

Traditional software:

Input
 ↓
Explicit rules
 ↓
Output


LLM:

Input
 ↓
Learned neural network
 ↓
Probability distribution
 ↓
Output

This difference is fundamental.

41. LLMs and Databases Are Different

A database might store:

User ID: 42
Name: Farhad
Age: ...

An LLM does not generally retrieve information from its parameters in this database-like way.

Instead, information is encoded in distributed numerical representations.

This is why an LLM can sometimes produce a fact correctly, incorrectly, or inconsistently.

42. What About RAG?

One common way to improve factual access is Retrieval-Augmented Generation (RAG).

Instead of relying entirely on the model's internal learned information:

Question
   ↓
LLM
   ↓
Answer

RAG can use:

Question
   ↓
Search / Retrieval
   ↓
Relevant documents
   ↓
LLM
   ↓
Answer

For example, a company could connect an LLM to its internal documentation.

The model retrieves relevant documents and uses them as context.

43. Tools Make LLMs More Capable

Modern AI systems can also use external tools.

For example:

User
 ↓
LLM
 ↓
Decide that a tool is needed
 ↓
Tool
 ↓
Tool result
 ↓
LLM
 ↓
Final response

Tools can provide capabilities such as:

  • web search
  • calculators
  • databases
  • code execution
  • file access
  • APIs

This is different from the neural network itself suddenly gaining those capabilities internally.

44. A Complete Simplified Architecture

We can now combine everything:

                    USER
                      β”‚
                      β–Ό
                   Prompt
                      β”‚
                      β–Ό
                 Tokenization
                      β”‚
                      β–Ό
                  Token IDs
                      β”‚
                      β–Ό
                  Embeddings
                      β”‚
                      β–Ό
             Positional Information
                      β”‚
                      β–Ό
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚      Transformer        β”‚
          β”‚                         β”‚
          β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
          β”‚  β”‚ Self-Attention    β”‚  β”‚
          β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
          β”‚            β–Ό            β”‚
          β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
          β”‚  β”‚ Feed-Forward NN   β”‚  β”‚
          β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
          β”‚            β–Ό            β”‚
          β”‚       Repeat many      β”‚
          β”‚          layers        β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚
                       β–Ό
                 Output logits
                       β”‚
                       β–Ό
              Probability distribution
                       β”‚
                       β–Ό
                 Token selection
                       β”‚
                       β–Ό
                 Next token
                       β”‚
                       └───────┐
                               β”‚
                               β–Ό
                         Repeat generation
                               β”‚
                               β–Ό
                            Answer

45. The Most Important Mental Model

If you remember only one thing from this article, remember this:

Text
 ↓
Tokens
 ↓
Numbers
 ↓
Embeddings
 ↓
Transformer
 ↓
Attention + neural network layers
 ↓
Probability distribution
 ↓
Next token
 ↓
Repeat
 ↓
Generated text

During training:

Text
 ↓
Predict
 ↓
Calculate loss
 ↓
Backpropagate
 ↓
Update parameters
 ↓
Repeat billions/trillions of times

That's the core idea.

46. So Is an LLM Just "Next-Word Prediction"?

Technically, next-token prediction is at the heart of autoregressive language modeling.

But calling an advanced LLM "just autocomplete" can be misleading.

Why?

Because learning to predict tokens across enormous and diverse datasets forces the network to learn many complicated internal representations.

To predict the next token effectively, the model may need to represent:

grammar
syntax
semantics
relationships
code structure
mathematical patterns
world knowledge
long-range dependencies

So the training objective can be simple while the learned internal behavior becomes extremely sophisticated.

47. What LLMs Still Cannot Guarantee

LLMs can be extremely capable, but they have important limitations.

They can:

  • produce incorrect information
  • misunderstand ambiguous prompts
  • make reasoning mistakes
  • generate outdated information
  • confidently state false claims
  • struggle with some exact computations
  • inherit biases from training data

Therefore:

Never confuse fluent output with guaranteed truth.

For important decisions, verify the information with appropriate primary sources or reliable tools.

48. Final Summary

An LLM is a large neural network trained to model sequences of tokens.

The basic process is:

             TRAINING

Training text
     ↓
Tokenization
     ↓
Transformer
     ↓
Prediction
     ↓
Loss
     ↓
Backpropagation
     ↓
Parameter updates
     ↓
Repeat


             INFERENCE

User prompt
     ↓
Tokenization
     ↓
Transformer
     ↓
Next-token probabilities
     ↓
Select token
     ↓
Add token to context
     ↓
Repeat
     ↓
Generated response

The key technologies behind this process include:

  • Tokenization
  • Embeddings
  • Positional representations
  • Transformers
  • Self-attention
  • Multi-head attention
  • Feed-forward networks
  • Residual connections
  • Gradient descent
  • Backpropagation
  • Large-scale training
  • Instruction tuning
  • Alignment
  • Sampling and decoding

The remarkable part isn't that the model was explicitly programmed with a giant collection of rules.

Instead, a neural network was trained to predict tokens so many times, across so much data, that it learned highly complex representations useful for language and many other tasks.

That's the fundamental idea behind modern LLMs.

What's Next?

If you want to go deeper, the natural next step is to understand the Transformer mathematicallyβ€”especially how Query, Key, Value, attention scores, softmax, and matrix multiplication work together.

Once you understand that, the internal architecture of an LLM becomes much less mysterious.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.