You type a sentence into ChatGPT, Claude, or Gemini and a paragraph comes back that sounds like a person wrote it. The interesting question is the next one: how do LLMs work, actually, under the hood? The honest answer is that a large language model is a very large statistical engine trained to predict the next chunk of text, one piece at a time. That single trick, scaled up enough, produces code, essays, translations, and arguments. This guide walks through every stage of that pipeline in plain English: tokens, embeddings, attention, training, inference, and the reasons these systems still get things wrong. By the end you will understand what is happening between your prompt and the answer, and you will have better reasons to check an answer rather than trust its tone.
Table of Contents
- What Is an LLM, Really
- Tokens and Embeddings: How LLMs See Text
- The Transformer and Self-Attention
- How LLMs Are Trained
- Inference: From Prompt to Reply
- Why LLMs Hallucinate
- Context Windows, RAG, and Tool Use
- FAQ
What Is an LLM, Really
A large language model is a neural network trained to represent patterns in language. For many generative models, predicting the next token is central to pretraining. Chat behavior also reflects later training and the application around the model. This guide focuses on the common autoregressive transformer approach, not every possible language-model architecture.
The modelâs learned information is encoded in numerical parameters rather than an ordinary searchable table of sentences. That does not mean it cannot memorize text: research on training-data extraction has demonstrated verbatim recovery in some settings. At generation time, the model computes candidate next-token scores from the available context.
Three Things an LLM Is Not
The model is not a conventional database. Its generated answer is not automatically a retrieved record. An application can connect it to search or databases, but those are additional mechanisms.
A fluent answer is not a transparent view of how it was produced. Generating intermediate steps can provide more computation and help with some tasks, but an explanation in the output is not guaranteed to be a faithful account of the internal process.
The product is more than one checkpoint. Providers can change models, instructions, retrieval, and tools. Different behavior does not necessarily mean the underlying weights alone changed.
Tokens and Embeddings: How LLMs See Text
Before an LLM can do anything with your prompt, it has to turn your text into numbers. This happens in two steps: tokenisation and embedding. Both are simpler than they sound and both shape how the model behaves more than people realise.
Tokenisation: Slicing Text Into Pieces
A tokenizer divides text into units called tokens. Depending on the tokenizer and the exact text, a token may represent part of a word, punctuation, whitespace, or another unit. âTokenisationâ might be split into fragments, but you need the actual tokenizer to know how âPudgy Catâ is encoded. The cat does not have a universal token count.
Byte-pair encoding is one common approach; it builds reusable units by merging frequent pairs. Other schemes exist, including unigram tokenization. SentencePiece supports both approaches. Subword and byte-based methods can represent unfamiliar strings through smaller pieces, but vocabularies and handling of unusual characters differ.
For cost and length estimates, count tokens with the tokenizer used by the actual model. English prose, code, whitespace, and different scripts can tokenize differently. Word-count rules of thumb are approximations, and equal byte length does not guarantee that a Python file uses more tokens than prose.
Embeddings: Turning Tokens Into Vectors
Tokens are mapped into numerical vectors called embeddings. Their dimensions depend on the architecture. These are a starting representation, not a container holding everything the model knows about a word; later layers transform representations using the surrounding context.
Learned representations can capture useful similarities, but their geometry is not a simple dictionary with every synonym neatly next door. Context matters. âBankâ in a sentence about a river needs different treatment from âbankâ in a sentence about a loan, even if the starting token is the same.
The Transformer and Self-Attention
The transformer was introduced in âAttention Is All You Needâ in 2017 and became highly influential in language modeling. It is not the only architecture: Mamba research explores a state-space alternative. Public information also does not justify asserting the exact internal design of every proprietary model.
What Attention Actually Does
For each token in your prompt, the model needs to figure out which other tokens matter. Take the sentence “The cat sat on the mat because it was tired.” The word “it” refers to the cat, not the mat. A human knows this instantly. The model has to compute it.
Attention computes weighted combinations of value vectors using relationships between queries and keys. In a causal text generator, a position cannot freely inspect future tokens. In the cat-and-mat example, context can help represent the pronounâs relationship, but no particular attention weight should be treated as a guaranteed grammatical explanation.
Multi-Head Attention and Stacking
Multiple attention heads use different learned projections, allowing information to be combined in different ways. It is tempting to assign each head a neat job title, but the behavior is more distributed and variable. Head counts are architecture choices, not one fixed range for every modern model.
Transformer blocks also use feed-forward transformations, normalization, and residual connections. Repeating blocks builds contextual representations. The number and layout vary, and the feed-forward portion is not necessarily a small afterthought in the computation.
How LLMs Are Trained
Training adjusts the modelâs parameters. A useful broad distinction is pretraining, which learns from large data collections, and post-training, which can shape instruction-following and other behavior. Real pipelines can contain more stages than this simplified division.
Pretraining: The Big Crawl
Pretraining data can include text and code from multiple sources, depending on the model. For next-token training, the system computes a loss from its predictions; backpropagation calculates gradients, and an optimizer uses them to update parameters. It is not simply a binary âwrong answerâ switch. Dataset composition, filtering, and scale vary.
Large training runs require substantial computing resources. Their costs depend on model size, data, hardware, and training decisions, so a single unsupported price tag for every frontier model would be misleading. Training a small experimental model and training a frontier system are very different projects.
Fine-Tuning and RLHF
A model that has only done pretraining is a fluent text-completion engine. Ask it a question and it might just continue the question in the style of a forum thread. To turn it into a useful assistant, providers add a second phase.
One influential approach, described in the 2022 InstructGPT paper, combines supervised examples with preference comparisons and reinforcement learning. The comparisons help train a reward model, which guides further optimization. Other post-training methods exist. No single technique alone explains every polite phrase, refusal, or summary produced by a chat application.
Researchers continue to investigate how training objectives and architectures affect capability. The practical lesson is to distinguish the pretraining task from the full system that answers you. A slogan about prediction does not settle every question about reasoning or understanding.
Inference: From Prompt to Reply
During ordinary inference, the deployed modelâs weights are generally fixed. The prompt and accumulated context influence the output without requiring a new weight update for each conversation. Generation can nevertheless be nondeterministic, and application-level memory is separate from training.
The Generation Loop
The system processes the prompt, computes scores for possible next tokens, chooses a token using its decoding procedure, and continues. Generation ends at a stopping condition or limit. This describes the basic autoregressive loop; a product can also pause for tools or perform other steps around it.
Implementations avoid repeating all work from scratch. A transformerâs key-value cache can reuse previously computed attention information; PagedAttention research examines how to manage that cache efficiently. New tokens still require computation, but neither âreprocess the entire prompt every timeâ nor âactivate every parameterâ is a universal description of inference.
Temperature and Sampling
Greedy decoding chooses the highest-scoring next token. Sampling instead draws from a distribution, often after adjustments. Which method is used depends on the system and task; not every service always samples.
Temperature changes how concentrated the sampling distribution is. Lower settings usually concentrate choices, while higher settings spread probability more widely. Top-k and top-p narrow the candidates in different ways. Low temperature is not a promise of reproducibility across hardware or implementations, and higher temperature is not a guarantee of creativity or quality.
Why LLMs Hallucinate
An LLM can generate plausible but false text. Fluent form does not guarantee factual grounding, and failure can arise even when relevant information is present. Models can learn to express uncertainty or abstain, but those behaviors are imperfect. The useful question is whether the specific answer is supported, not whether it sounds hesitant or confident.
Common Hallucination Failure Modes
Fake citations: when asked for sources, the model generates strings that look like real academic references because that pattern is overwhelmingly common in its training data. The format is correct, the authors and titles are invented.
Wrong dates and quantities: a model can produce a plausible number that does not match the evidence. Verify it against an appropriate source rather than assuming confident specificity reflects a lookup.
Phantom features: ask an LLM how to do something in a software tool and it might invent a plausible API call that does not exist, because the pattern “library has method that does this” is strong.
How to Reduce Hallucinations
One way to improve grounding is to retrieve relevant material and provide it to the model. Retrieval-augmented generation research combines retrieval with generation. It can help, but bad retrieval or misuse of a document can still produce errors. Our RAG guide explains the application pattern in more detail.
Context Windows, RAG, and Tool Use
Applications can supply documents, conversation history, and tool results alongside your words. Those additions can matter as much as the visible question. Three useful pieces of the system are context management, retrieval, and tools.
The Context Window
A context window limits how much tokenized material a model can work with under a particular configuration. The application decides what to include, summarize, or retrieve. Material outside the immediate context is not necessarily deleted from the productâs storage, but the model cannot rely on seeing it unless the system supplies it again.
Retrieval Augmented Generation
A retrieval system selects material from an external collection and makes it available for generation. It may use vector similarity, keyword search, or a combination. This can provide fresher or task-specific information without retraining the model. It is a useful design option, not a requirement followed by every serious deployment.
Tool Use and MCP
A tool-enabled application can let a model request a search, calculation, or database operation, execute that request, and return the result as context. MCP was introduced in November 2024 as a protocol for connecting AI applications with external systems; tool use did not begin with it. See our MCP explainer for that connection layer.
FAQ
How do LLMs work in one sentence?
An LLM is a giant neural network that has been trained to predict the next token in a sequence, and “chatting” is just that prediction loop running over and over with your prompt as the seed.
What is the difference between an LLM and AI?
AI is the broader field, including image recognition, robotics, search, and game-playing systems. An LLM is one specific type of AI: a large language model trained on text to generate text. ChatGPT, Claude, and Gemini are LLM-based products.
Do LLMs actually understand language?
âUnderstandâ can mean different things, from using language successfully to having human-like experience. Models learn representations that support useful behavior, but that alone does not settle the philosophical question. Some systems also process images or audio, so saying that all such models lack sensory inputs is too broad.
Why do LLMs make stuff up?
Generating a well-formed answer and generating a correct answer are different requirements. Training and grounding can reduce errors, but neither guarantees correctness. Check the evidence for claims that matter, especially citations, numbers, and instructions.
Can I train my own LLM?
You can train a small model as a learning exercise, or adapt an existing model if your hardware, data, and method permit it. Fine-tuning is not automatically cheap for every setup. Running a model locally is a separate task; our local AI guide is a starting point.
The Takeaway
The useful picture has several parts: tokenized inputs, learned representations, a generation process, and an application that may add retrieval or tools. Next-token prediction explains an important mechanism, not every behavior of the complete product. Understanding those pieces helps you ask better questions and verify answers. A confident paragraph still owes you evidence, however nicely it is formatted.




