AI & Machine Learning

Large Language Model Architecture: How LLMs Actually Work

A technical deep dive into LLM architecture: tokenization, embeddings, transformer blocks, attention mechanisms, and why these systems behave the way they do—including hallucinations and limitations.

Khalid Aboubakr
28 min read
LlmTransformersAttention MechanismGptNeural NetworksDeep LearningNlp

Why Understanding LLM Architecture Matters

If you're integrating LLMs into production systems, you need more than API documentation. You need to understand why these systems hallucinate, why they struggle with certain tasks, and why prompt engineering works. That understanding comes from architecture.

This isn't a simplified "neural networks are like brains" explanation. It's a technical walkthrough for engineers who need to make informed decisions about when and how to use LLMs.

The High-Level Pipeline

When you send text to an LLM, here's what actually happens:

Input Text → Tokenization → Embedding → Transformer Blocks → Output Probabilities → Token Selection → Decoded Text

Each stage has architectural implications that affect behavior. Let's examine each.

Stage 1: Tokenization

LLMs don't see text—they see sequences of integers representing tokens. Tokenization is the conversion process.

How Tokenization Works:

Modern LLMs use subword tokenization (typically Byte-Pair Encoding or SentencePiece). The algorithm:

  1. Starts with a vocabulary of individual characters
  2. Iteratively merges the most frequent adjacent pairs
  3. Stops when vocabulary reaches target size (typically 32K-100K tokens)
"Understanding" → ["Under", "stand", "ing"] → [15496, 9122, 278]
"AI" → ["AI"] → [15836]
"الذكاء" → ["ال", "ذ", "كاء"] → [various token IDs]

Why This Matters for Engineers:

  • Token limits are not word limits: A 4096 token limit might be 3000 words in English but fewer in other languages or code.
  • Tokenization affects reasoning: The model processes "understanding" as three tokens, not as a semantic unit. This can affect how it handles morphologically complex words.
  • Cost is per-token: API pricing is token-based. Understanding tokenization helps predict costs.
  • Rare words get split more: Domain-specific terminology often becomes many tokens, reducing effective context window.

Common Tokenization Artifacts:

# Leading spaces matter tokenize("hello") != tokenize(" hello") # Different token sequences # Numbers are split character by character in some tokenizers tokenize("2024") → ["2", "0", "2", "4"] # Why LLMs struggle with arithmetic

Stage 2: Embeddings

Each token ID maps to a high-dimensional vector (embedding). For GPT-4 class models, this is typically 4096-12288 dimensions.

What Embeddings Encode:

  • Semantic similarity (similar meanings → similar vectors)
  • Syntactic role (parts of speech cluster together)
  • Relationships (analogies like king - man + woman ≈ queen)

The Embedding Matrix:

Vocabulary size: 100,000 tokens
Embedding dimension: 4096

Embedding matrix: 100,000 × 4096 = 409M parameters just for embeddings

Positional Encoding:

Transformers have no inherent sense of position. Position information is added to embeddings:

  • Absolute positional encoding: Fixed patterns based on position index
  • Rotary positional encoding (RoPE): Rotations in embedding space that encode relative positions
  • ALiBi: Attention bias based on distance between tokens

RoPE and ALiBi allow for better extrapolation beyond training context lengths, which is why modern LLMs can handle longer contexts than their training data contained.

Stage 3: Transformer Blocks

The core of an LLM is a stack of transformer blocks (typically 32-96 layers). Each block contains:

  1. Multi-Head Self-Attention: Where tokens "look at" other tokens
  2. Feed-Forward Network: Where individual token representations are transformed
  3. Layer Normalization: Stabilizes training
  4. Residual Connections: Allows gradients to flow through deep networks

Self-Attention: The Core Mechanism

Self-attention computes how much each token should attend to every other token.

For each token, we compute three vectors:

  • Query (Q): "What am I looking for?"
  • Key (K): "What do I contain?"
  • Value (V): "What information do I provide?"

The attention computation:

Attention(Q, K, V) = softmax(QK^T / √d_k) × V

Where:

  • QK^T produces attention scores (how relevant is each token to each other)
  • √d_k scaling prevents softmax from saturating
  • Softmax normalizes to probability distribution
  • Multiplication with V produces weighted combination of values

Multi-Head Attention:

Instead of single attention, we run multiple parallel attention mechanisms ("heads"), each learning different patterns:

  • Some heads learn syntactic relationships
  • Some learn semantic similarity
  • Some learn positional patterns
  • Some learn domain-specific relationships
8-64 attention heads per layer
Each head: d_model / num_heads dimensions
Concatenated and projected back to d_model

The Feed-Forward Network:

After attention, each token passes through a two-layer MLP:

FFN(x) = activation(xW₁ + b₁)W₂ + b₂

This is where most parameters live and where "knowledge" is stored. Research suggests these networks act as key-value memories, storing factual associations.

Training vs. Inference

Training (Pre-training):

LLMs are trained to predict the next token given all previous tokens:

Input:  "The capital of France is"
Target: "Paris"

Loss = -log P("Paris" | "The capital of France is")

This is repeated trillions of times across internet-scale text:

  • GPT-3: ~300 billion tokens
  • LLaMA 2: ~2 trillion tokens
  • GPT-4: Estimated 10+ trillion tokens

Inference:

During inference, the model generates one token at a time:

  1. Process input, get probability distribution over vocabulary
  2. Sample or select next token
  3. Append to input, repeat

Autoregressive generation is sequential: Each new token requires processing the entire sequence. This is why:

  • Generation is slower than encoding
  • Context length directly impacts latency
  • KV-caching is critical for performance

KV-Cache:

To avoid recomputing attention for all previous tokens, we cache the Key and Value matrices:

Without cache: O(n²) per token generated
With cache: O(n) per token generated

Cache size: 2 × layers × d_model × context_length × batch_size

This is why LLM inference is memory-bound, not compute-bound, for long sequences.

Why LLMs Hallucinate

Understanding architecture explains hallucination:

1. Training Objective: LLMs are trained to predict plausible next tokens, not true ones. "The first person on Mars was Neil Armstrong" is plausible-sounding even if wrong.

2. No Knowledge Verification: There's no mechanism to verify claims against a knowledge base. The model pattern-matches against training data patterns, not facts.

3. Confidence Calibration: Softmax produces confident-looking probabilities even for uncertain outputs. A model might output "definitely" while being statistically uncertain.

4. Distributional Artifacts: If the training data contained more instances of X than Y, the model will favor X even when Y is correct for the specific context.

Architectural Implications:

  • Retrieval-Augmented Generation (RAG): Inject relevant documents into context to ground responses
  • Chain-of-thought prompting: Force intermediate reasoning steps that can be verified
  • Temperature tuning: Lower temperature for factual tasks, higher for creative
  • Ensemble verification: Use multiple queries or models to cross-check

Architectural Limitations

Fixed Context Window: Attention is O(n²) in context length. Even with optimizations, there are hard limits to how much context a model can effectively use.

No True Reasoning: LLMs approximate reasoning through pattern matching on reasoning examples in training data. Novel reasoning chains that weren't in training data are unreliable.

No Memory: Each request is independent. The model cannot learn from interactions unless explicitly retrained.

Tokenization Artifacts: Arithmetic, character-level manipulation, and non-English languages can be impaired by tokenization choices.

Prompt Sensitivity: Small changes in prompt wording can dramatically change outputs. This is a feature of the attention mechanism, not a bug, but it makes systems fragile.

Practical Architecture Decisions

When integrating LLMs into systems, architecture knowledge informs decisions:

Context Window Management:

Available context: 128K tokens
System prompt: 2K tokens
Retrieved documents: 10K tokens
Conversation history: X tokens
User query: 500 tokens
Response budget: 4K tokens

Available for history: 128K - 2K - 10K - 500 - 4K = 111.5K tokens

Latency Budgeting:

Time to first token: ~200-500ms (depends on prompt length)
Time per output token: ~20-50ms
Total for 500 token response: ~10-25 seconds

For real-time applications, consider streaming responses.

Cost Estimation:

GPT-4 Turbo: $10/million input tokens, $30/million output tokens

1000 requests/day × 2000 input tokens × 500 output tokens
= 2M input + 0.5M output daily
= $20 + $15 = $35/day = ~$1000/month

This scales fast. Understand your token usage patterns.

The Transformer Revolution: Why This Architecture Won

The transformer architecture (introduced in "Attention Is All You Need," 2017) replaced recurrent networks because:

  1. Parallelization: Unlike RNNs, transformers process all positions simultaneously during training
  2. Long-range dependencies: Attention directly connects any two positions, no vanishing gradients
  3. Scalability: Performance scales predictably with compute and data

This enabled training on unprecedented scales, leading to emergent capabilities not seen in smaller models.

Related Articles

AI & Machine Learning20 min read

Deep Learning (Goodfellow et al.): A Critical Review

A practitioner's critical analysis of the canonical Deep Learning textbook—what it does well, where it falls short, and what working engineers should know before reading it.

Backend Design19 min read

API Design: Choosing Between REST, GraphQL, and gRPC

Compare REST, GraphQL, and gRPC APIs with performance benchmarks and use cases. Learn which API style fits your project based on real production experience.