Large Language Model Architecture: How LLMs Actually Work
A technical deep dive into LLM architecture: tokenization, embeddings, transformer blocks, attention mechanisms, and why these systems behave the way they do—including hallucinations and limitations.
Why Understanding LLM Architecture Matters
If you're integrating LLMs into production systems, you need more than API documentation. You need to understand why these systems hallucinate, why they struggle with certain tasks, and why prompt engineering works. That understanding comes from architecture.
This isn't a simplified "neural networks are like brains" explanation. It's a technical walkthrough for engineers who need to make informed decisions about when and how to use LLMs.
The High-Level Pipeline
When you send text to an LLM, here's what actually happens:
Input Text → Tokenization → Embedding → Transformer Blocks → Output Probabilities → Token Selection → Decoded Text
Each stage has architectural implications that affect behavior. Let's examine each.
Stage 1: Tokenization
LLMs don't see text—they see sequences of integers representing tokens. Tokenization is the conversion process.
How Tokenization Works:
Modern LLMs use subword tokenization (typically Byte-Pair Encoding or SentencePiece). The algorithm:
- Starts with a vocabulary of individual characters
- Iteratively merges the most frequent adjacent pairs
- Stops when vocabulary reaches target size (typically 32K-100K tokens)
"Understanding" → ["Under", "stand", "ing"] → [15496, 9122, 278]
"AI" → ["AI"] → [15836]
"الذكاء" → ["ال", "ذ", "كاء"] → [various token IDs]
Why This Matters for Engineers:
- Token limits are not word limits: A 4096 token limit might be 3000 words in English but fewer in other languages or code.
- Tokenization affects reasoning: The model processes "understanding" as three tokens, not as a semantic unit. This can affect how it handles morphologically complex words.
- Cost is per-token: API pricing is token-based. Understanding tokenization helps predict costs.
- Rare words get split more: Domain-specific terminology often becomes many tokens, reducing effective context window.
Common Tokenization Artifacts:
# Leading spaces matter tokenize("hello") != tokenize(" hello") # Different token sequences # Numbers are split character by character in some tokenizers tokenize("2024") → ["2", "0", "2", "4"] # Why LLMs struggle with arithmetic
Stage 2: Embeddings
Each token ID maps to a high-dimensional vector (embedding). For GPT-4 class models, this is typically 4096-12288 dimensions.
What Embeddings Encode:
- Semantic similarity (similar meanings → similar vectors)
- Syntactic role (parts of speech cluster together)
- Relationships (analogies like king - man + woman ≈ queen)
The Embedding Matrix:
Vocabulary size: 100,000 tokens
Embedding dimension: 4096
Embedding matrix: 100,000 × 4096 = 409M parameters just for embeddings
Positional Encoding:
Transformers have no inherent sense of position. Position information is added to embeddings:
- Absolute positional encoding: Fixed patterns based on position index
- Rotary positional encoding (RoPE): Rotations in embedding space that encode relative positions
- ALiBi: Attention bias based on distance between tokens
RoPE and ALiBi allow for better extrapolation beyond training context lengths, which is why modern LLMs can handle longer contexts than their training data contained.
Stage 3: Transformer Blocks
The core of an LLM is a stack of transformer blocks (typically 32-96 layers). Each block contains:
- Multi-Head Self-Attention: Where tokens "look at" other tokens
- Feed-Forward Network: Where individual token representations are transformed
- Layer Normalization: Stabilizes training
- Residual Connections: Allows gradients to flow through deep networks
Self-Attention: The Core Mechanism
Self-attention computes how much each token should attend to every other token.
For each token, we compute three vectors:
- Query (Q): "What am I looking for?"
- Key (K): "What do I contain?"
- Value (V): "What information do I provide?"
The attention computation:
Attention(Q, K, V) = softmax(QK^T / √d_k) × V
Where:
- QK^T produces attention scores (how relevant is each token to each other)
- √d_k scaling prevents softmax from saturating
- Softmax normalizes to probability distribution
- Multiplication with V produces weighted combination of values
Multi-Head Attention:
Instead of single attention, we run multiple parallel attention mechanisms ("heads"), each learning different patterns:
- Some heads learn syntactic relationships
- Some learn semantic similarity
- Some learn positional patterns
- Some learn domain-specific relationships
8-64 attention heads per layer
Each head: d_model / num_heads dimensions
Concatenated and projected back to d_model
The Feed-Forward Network:
After attention, each token passes through a two-layer MLP:
FFN(x) = activation(xW₁ + b₁)W₂ + b₂
This is where most parameters live and where "knowledge" is stored. Research suggests these networks act as key-value memories, storing factual associations.
Training vs. Inference
Training (Pre-training):
LLMs are trained to predict the next token given all previous tokens:
Input: "The capital of France is"
Target: "Paris"
Loss = -log P("Paris" | "The capital of France is")
This is repeated trillions of times across internet-scale text:
- GPT-3: ~300 billion tokens
- LLaMA 2: ~2 trillion tokens
- GPT-4: Estimated 10+ trillion tokens
Inference:
During inference, the model generates one token at a time:
- Process input, get probability distribution over vocabulary
- Sample or select next token
- Append to input, repeat
Autoregressive generation is sequential: Each new token requires processing the entire sequence. This is why:
- Generation is slower than encoding
- Context length directly impacts latency
- KV-caching is critical for performance
KV-Cache:
To avoid recomputing attention for all previous tokens, we cache the Key and Value matrices:
Without cache: O(n²) per token generated
With cache: O(n) per token generated
Cache size: 2 × layers × d_model × context_length × batch_size
This is why LLM inference is memory-bound, not compute-bound, for long sequences.
Why LLMs Hallucinate
Understanding architecture explains hallucination:
1. Training Objective: LLMs are trained to predict plausible next tokens, not true ones. "The first person on Mars was Neil Armstrong" is plausible-sounding even if wrong.
2. No Knowledge Verification: There's no mechanism to verify claims against a knowledge base. The model pattern-matches against training data patterns, not facts.
3. Confidence Calibration: Softmax produces confident-looking probabilities even for uncertain outputs. A model might output "definitely" while being statistically uncertain.
4. Distributional Artifacts: If the training data contained more instances of X than Y, the model will favor X even when Y is correct for the specific context.
Architectural Implications:
- Retrieval-Augmented Generation (RAG): Inject relevant documents into context to ground responses
- Chain-of-thought prompting: Force intermediate reasoning steps that can be verified
- Temperature tuning: Lower temperature for factual tasks, higher for creative
- Ensemble verification: Use multiple queries or models to cross-check
Architectural Limitations
Fixed Context Window: Attention is O(n²) in context length. Even with optimizations, there are hard limits to how much context a model can effectively use.
No True Reasoning: LLMs approximate reasoning through pattern matching on reasoning examples in training data. Novel reasoning chains that weren't in training data are unreliable.
No Memory: Each request is independent. The model cannot learn from interactions unless explicitly retrained.
Tokenization Artifacts: Arithmetic, character-level manipulation, and non-English languages can be impaired by tokenization choices.
Prompt Sensitivity: Small changes in prompt wording can dramatically change outputs. This is a feature of the attention mechanism, not a bug, but it makes systems fragile.
Practical Architecture Decisions
When integrating LLMs into systems, architecture knowledge informs decisions:
Context Window Management:
Available context: 128K tokens
System prompt: 2K tokens
Retrieved documents: 10K tokens
Conversation history: X tokens
User query: 500 tokens
Response budget: 4K tokens
Available for history: 128K - 2K - 10K - 500 - 4K = 111.5K tokens
Latency Budgeting:
Time to first token: ~200-500ms (depends on prompt length)
Time per output token: ~20-50ms
Total for 500 token response: ~10-25 seconds
For real-time applications, consider streaming responses.
Cost Estimation:
GPT-4 Turbo: $10/million input tokens, $30/million output tokens
1000 requests/day × 2000 input tokens × 500 output tokens
= 2M input + 0.5M output daily
= $20 + $15 = $35/day = ~$1000/month
This scales fast. Understand your token usage patterns.
The Transformer Revolution: Why This Architecture Won
The transformer architecture (introduced in "Attention Is All You Need," 2017) replaced recurrent networks because:
- Parallelization: Unlike RNNs, transformers process all positions simultaneously during training
- Long-range dependencies: Attention directly connects any two positions, no vanishing gradients
- Scalability: Performance scales predictably with compute and data
This enabled training on unprecedented scales, leading to emergent capabilities not seen in smaller models.
Related Articles
- Artificial Intelligence for Engineers: Foundations Without the Hype - The broader AI context
- Deep Learning (Goodfellow et al.): A Critical Review - Deeper theoretical foundations
- API Design: Choosing Between REST, GraphQL, and gRPC - Exposing LLM capabilities as services
- Building Resilient Distributed Systems - Patterns for reliable LLM integration
Related Articles
AI & Machine Learning22 min read
Artificial Intelligence for Engineers: Foundations Without the Hype
A grounded explanation of what AI actually is, how it differs from classical ML and deep learning, and what working engineers need to understand beyond the marketing noise.
AI & Machine Learning20 min read
Deep Learning (Goodfellow et al.): A Critical Review
A practitioner's critical analysis of the canonical Deep Learning textbook—what it does well, where it falls short, and what working engineers should know before reading it.
Backend Design19 min read
API Design: Choosing Between REST, GraphQL, and gRPC
Compare REST, GraphQL, and gRPC APIs with performance benchmarks and use cases. Learn which API style fits your project based on real production experience.
Software Architecture19 min read
Building Resilient Distributed Systems: Patterns for Fault Tolerance
Build resilient distributed systems with circuit breakers, retries, and timeouts. Production patterns for handling failures, cascading errors, and maintaining availability.