ModelRefs / What is a Large Language Model (LLM)? A Plain-English Guide
What is a Large Language Model (LLM)? A Plain-English Guide
LLMs power ChatGPT, Claude, and Gemini. Here is how they actually work, what they can and cannot do, and how to start using them today.
A large language model, or LLM, is a neural network trained on enormous amounts of text. It learns the statistical patterns of language so well that it can write essays, answer questions, summarise documents, translate between languages, and even generate working code.
How LLMs actually work
Under the hood, an LLM does one thing extremely well: predict the next token (roughly, the next word or piece of a word) given everything it has seen so far. Repeat that prediction thousands of times and you get a paragraph. Repeat it millions of times across a conversation and you get something that feels like reasoning.
The magic ingredient is the Transformer architecture, introduced in the 2017 paper Attention Is All You Need. Transformers use a mechanism called self-attention that lets the model weigh how much every word in the input relates to every other word — which is why an LLM can keep track of a name mentioned twelve sentences ago.
Why "large"?
Modern LLMs are large in three ways:
- Parameters — the tunable weights inside the network. GPT-4-class models have hundreds of billions.
- Training data — typically trillions of tokens scraped from the public web, books, and code.
- Compute — training a frontier model can cost tens of millions of dollars in GPU time.
What LLMs are great at
- Drafting and editing text
- Summarising long documents
- Translation
- Code generation and explanation
- Conversational interfaces
What LLMs are bad at
- Doing exact arithmetic
- Knowing facts after their training cutoff
- Citing sources reliably (they will sometimes hallucinate)
- Reasoning about images or audio (unless they are multimodal)
Start exploring
The fastest way to build intuition is to try one yourself. Head to the ModelRefs Playground to chat with a frontier model in seconds, no API key required.