TL;DR
LLM architecture is the internal design of a large language model, the path a piece of text takes from tokens to the next predicted word. Text is split into tokens, and each token becomes a vector of numbers. A stack of transformer blocks then refines those vectors, and self-attention lets every token weigh the words around it. Most generative models use a decoder-only layout, and many newer ones add mixture-of-experts layers that use only part of the model per token. The context window and the KV cache decide how much text a model can hold and how much GPU memory each request needs. For an enterprise, these choices set the cost, speed and hardware footprint of every model on the shortlist.
Key Takeaways LLM architecture describes how a model turns tokens into vectors, refines them through stacked transformer blocks and predicts the next token. Self-attention lets each token weigh every earlier token, which is powerful but grows expensive as prompts get longer. Decoder-only models dominate text generation, while encoder-only models still suit classification, extraction and search embeddings. Mixture-of-experts models store many parameters but activate only a fraction per token, so memory and compute must be judged separately. The KV cache can rival the model weights in memory at long context lengths, and grouped-query attention shrinks it. Reading a model card through an architecture lens predicts latency, memory and cost before any pilot starts. Watch on YouTube
World Models vs LLMs: What’s the Real Difference
A quick explainer on how world models differ from next-token predictors, and why researchers think the architecture will keep changing.
Reading One Sentence the Way a Model Does Read this sentence slowly. The contract renewed automatically because it had no termination clause. By the word “it”, your brain has already linked the pronoun to “contract” without effort.
A language model does something similar, but with a highlighter. At every word it glances back across the whole sentence, picks out the words that matter for this one, and blends them into its understanding.
Repeat that habit across dozens of layers and billions of weights and you have most of modern LLM architecture. It also explains why long prompts cost more and why some models need far more memory than others. Even your cloud bill can be read off the spec sheet.
What Is LLM Architecture? LLM architecture is the blueprint inside a large language model. It covers how text is split into tokens and how those tokens become numbers.
It also covers how a stack of transformer blocks refines them and how the model picks the next token. Because these choices come before any training, two models trained on similar data can still behave very differently.
The term is often stretched to cover everything around the model as well, such as retrieval, guardrails, orchestration and APIs. That outer layer is system design, and our guide to generative AI architecture covers those patterns. This article stays inside the model itself, which is where the cost, speed and memory profile of every deployment is set.
How One Request Moves Through the Model Follow a single prompt through the network. A tokenizer cuts the text into tokens and maps each one to an ID from a fixed vocabulary. An embedding table turns each ID into a long vector, and position information is mixed in so the model knows the word order.
Those vectors then pass through a stack of identical transformer blocks. Llama 3 8B has 32 such blocks, and the 405B version has 126, according to the Llama 3 technical report . After the last block, an output layer scores every entry in the vocabulary, and the model samples the next token from those scores.
Next, the new token is appended to the input and the loop runs again. Generation is therefore a long series of single steps, one token at a time. That detail matters later, because it is why the KV cache exists and why output length drives latency.
Why Architecture Sets Cost, Speed and Memory Every component in that path has a price. For instance, the number of layers and the width of each one set how much arithmetic each token needs. The attention design sets how much memory each conversation holds, and the context window sets how long a single request can grow.
This is why two models with similar benchmark scores can differ several times over in serving cost. A team that reads the architecture section of a model card can estimate GPU memory and latency before running a pilot. A team that reads only the parameter count usually finds out later, on the invoice.
Tokenization and Embeddings: How Text Becomes Numbers A neural network cannot read letters, so the first job of any LLM is to turn text into numbers it can compute with. Two steps do this work: tokenization first splits text into pieces, and embeddings then turn each piece into a vector.
Tokens Are Not Words Modern tokenizers split text into subword pieces. For example, a common word may be one token, while a rare product code or a word in another language may break into several. Llama 3 uses a vocabulary of 128,000 tokens, built from 100,000 tiktoken tokens plus 28,000 more for non-English text.
The tokenizer has real cost effects. Meta reports that the Llama 3 tokenizer packs 3.94 characters into each English token, up from 3.17 in Llama 2. As a result, fewer tokens for the same document means a lower bill on per-token pricing and more room left in the context window.
For multilingual workloads the gap can be larger, since a tokenizer trained mostly on English splits other scripts into more pieces. For that reason, test your own documents through each candidate’s tokenizer before comparing price sheets. A cheaper price per token can disappear if the same contract needs a third more tokens.
Embeddings Turn IDs Into Vectors Each token ID points to a row in an embedding table. That row is a learned vector, 4,096 numbers long in Llama 3 8B, which then becomes the token’s starting representation. Over training, tokens used in similar ways end up with vectors that sit close together.
This vector width, often called the hidden size or model dimension, travels through the whole network. Wider vectors can carry more nuance. However, every layer then does more arithmetic per token.
The same idea of meaning as geometry powers the search embeddings stored in a vector database , although those usually come from separate, smaller models.
Positional Encoding Gives Words an Order Attention on its own treats a sentence like a bag of words. Without help, “the bank paid the client” and “the client paid the bank” would look alike.
Positional encoding fixes this by adding order information to each token. The original Transformer added fixed sine and cosine patterns to the embeddings.
Most current open models use rotary position embedding, or RoPE. It rotates the query and key vectors by an angle tied to each position, so attention sees how far apart two tokens are. The RoFormer paper describes it, and Llama 3 raised its RoPE base frequency to 500,000 to support longer contexts.
Position handling is one reason a model’s context window cannot simply be stretched after release. A model trained on short sequences has never seen the rotation angles of very distant positions. Extending the window takes extra training, which is a topic for our LLM training guide rather than this one.
Kanerika Service
LLM Development Services
Kanerika profiles your token volumes and context needs, then shortlists and sizes the model architecture that fits your budget.
Explore LLM Development Inside a Transformer Block The transformer block is the repeating unit of LLM architecture. Each block has two main parts. An attention layer lets tokens exchange information, and a feed-forward layer then processes each token on its own.
The design comes from the 2017 paper Attention Is All You Need , which dropped the recurrence that older recurrent neural networks relied on. Its base model stacked 6 layers with a model width of 512 and 8 attention heads. Today’s models still keep the same skeleton but repeat it far more often and make it much wider.
The sketch below shows the logic of one block and of the full model in simplified Python. Real implementations add details such as dropout during training, caching and fused GPU kernels, but the flow is the same.
def transformer_block(x):
# x holds one vector per token, shape [tokens, d_model]
h = x + attention(normalize(x)) # tokens share information
out = h + feed_forward(normalize(h)) # each token is refined on its own
return out
def llm_next_token_scores(token_ids):
x = embed(token_ids) # token IDs become vectors
for block in blocks: # 32 blocks in Llama 3 8B
x = transformer_block(x)
logits = output_head(normalize(x)) # one score per vocabulary entry
return logits[-1] # scores for the next tokenFeed-Forward Layers Hold Most of the Weights Attention gets the headlines, yet the feed-forward network is where most parameters live. In Llama 3 8B, each block’s feed-forward layer expands a 4,096-wide vector to 14,336 and back, using a gated SwiGLU design. By our arithmetic on the published dimensions, that layer holds about 176 million weights per block, against roughly 42 million for attention.
That works out to about four-fifths of each block’s weights. It also explains where mixture-of-experts designs make their change, since they replace this feed-forward layer with several smaller experts. More on that below.
Residual Connections and Normalization Keep Deep Stacks Stable Look again at the “x +” in the sketch above. That is a residual connection, and it means each sub-layer adds a correction to its input instead of replacing it. Because of this, information from early layers can flow straight to later ones, which lets very deep stacks train without the signal fading.
Normalization does a related job, rescaling each vector before a sub-layer reads it so numbers stay stable across 80 or 126 layers. Neither part is glamorous, but a stack as deep as Llama 3 405B would not train reliably without them. The broader basics of layered networks are covered in our deep learning glossary entry.
How Self-Attention Works Self-attention is the step that lets each token gather context from the others. It is the highlighter from the opening example, written as matrix arithmetic. Each token asks a question, every other token holds a possible answer, and the model blends those answers by relevance.
Queries, Keys and Values Each token vector is projected three ways. The query describes what this token is looking for, and the key describes what each token has to offer. A third projection, the value, carries the information that will actually be passed along.
The model compares one token’s query with every key using a dot product, which gives a relevance score for each pair. It divides the scores by the square root of the head size to keep them in a stable range. A softmax then turns them into weights that add up to one, and the output is the weighted sum of the values.
A Worked Example: Resolving One Pronoun Return to the contract sentence from the opening. When the model processes “it”, its query is compared with the keys of the five tokens before it and with its own key. The weights below are illustrative, chosen to show the shape of a trained head rather than measured from a real model.
Table 1: Illustrative attention weights for the token “it”
Earlier token Attention weight What the weight means contract 0.61 Most likely referent, so its value dominates the blend renewed 0.12 The action the subject took automatically 0.08 Modifier, minor context because 0.07 Signals that a reason follows it (itself) 0.07 Keeps some of its own meaning The 0.05 Little information to add
After this step, the vector for “it” therefore carries most of the meaning of “contract”. Later layers can then link “had no termination clause” to the right noun. In practice, one head rarely does all of this alone, which is where multiple heads come in.
Multi-Head, Multi-Query and Grouped-Query Attention Multi-head attention runs several attention operations side by side, each on a slice of the vector. One head might track grammar while another tracks which entity a pronoun refers to. The original Transformer used 8 heads of 64 dimensions each, and Llama 3 405B uses 128.
Normally, each head keeps its own keys and values, and those must be stored during generation. Multi-query attention shares one set of keys and values across all heads, which cuts memory traffic but can cost some quality. Grouped-query attention sits in between, with a few shared key-value groups.
Llama 3 uses grouped-query attention with 8 key-value heads in every model size. According to Meta, the goal was faster inference and a smaller KV cache during decoding.
For an enterprise buyer, the KV head count on a model card is a strong clue. It predicts how many users a GPU can serve at once.
Causal Masking Stops the Model Peeking Ahead Look at the example again and notice what “it” could not see. The words “termination clause” come after the pronoun, and a generative model is not allowed to read ahead. A causal mask blocks every token from attending to later positions, because during generation those words do not exist yet.
Even so, the model still resolves the sentence, just later. When it reaches “clause”, that token can attend back to both “it” and “contract”. This one-way flow is the defining trait of decoder-only models.
Why Long Prompts Make Attention Expensive In full attention every token is scored against every earlier token, so the work grows with the square of the sequence length. Doubling a prompt from 50,000 to 100,000 tokens roughly quadruples the attention work during prompt processing. The FlashAttention paper states it plainly, noting that time and memory for self-attention are quadratic in sequence length.
Engineers have answered with exact but faster kernels such as FlashAttention. Designs such as sliding-window attention also limit how far back some layers look.
These help, but the economics stay the same. Long prompts are slower and costlier per request, so stuffing every document into a prompt is rarely the cheapest design.
Encoder-Only vs Decoder-Only vs Encoder-Decoder Models The original Transformer had two halves. An encoder read the whole input in both directions, while a decoder wrote the output one token at a time, consulting the encoder through cross-attention. Later model families kept one half or both, and that choice still shapes what each model is good at.
Table 2: The three transformer layouts compared
Layout How tokens see each other Typical enterprise jobs Well-known examples Encoder-only Both directions, every token sees the full input Classification, entity extraction, search embeddings BERT and its descendants Decoder-only One direction, each token sees only earlier tokens Chat, drafting, summarization, code, agents Llama 3, Mixtral, DeepSeek-V3, GPT-1 to GPT-3 Encoder-decoder Encoder reads both ways, decoder writes one way and attends to the encoder Translation, structured rewriting, text-to-text tasks Original Transformer, T5
BERT showed how strong a bidirectional encoder could be for understanding tasks. T5 framed every language problem as text in, text out, using the full encoder-decoder design. Both still remain useful, especially where the output is a label or a vector rather than free text.
On the decoder side, OpenAI’s first GPT paper describes a 12-layer decoder-only transformer, and GPT-2 and GPT-3 kept that design. OpenAI has not published the architecture of its newer models.
Why Generative Models Settled on Decoder-Only A decoder-only model learns one simple objective, predicting the next token, and that objective works on almost any text. The same model can answer questions, write code and follow instructions without a separate encoder. One stack is also simpler to scale and to serve.
Talk to Kanerika
Talk Through Your Model Choice With Our Architects
Bring your workload, token volumes and latency targets to Kanerika’s engineers and compare the architecture options before you commit to a model.
Book a Meeting → The trade-off is that a decoder reads only leftward. For pure understanding tasks, a small encoder can match it at a fraction of the cost.
That is why many enterprise pipelines quietly run both. A compact encoder classifies or embeds documents at volume, and a decoder handles the steps that need fluent language. Our SLMs vs LLMs comparison looks at that split from the sizing side.
Mixture of Experts: Swapping Out the Feed-Forward Layer The feed-forward layer holds about four-fifths of each block’s weights, so it is the natural place to make a model sparse. A mixture-of-experts (MoE) layer therefore replaces that one network with several smaller experts plus a router. For each token, the router picks a few experts and only those run, while attention and the rest of the block stay unchanged.
Mixtral 8x7B is the classic example, with 8 experts per layer and every token routed to 2 of them. Mistral’s launch post reports 46.7B total parameters but only 12.9B used per token. According to Mistral, it therefore runs at the speed and cost of a 12.9B model.
Table 3: Published total and active parameters for dense and MoE models
Model Design Total parameters Active per token Context window Llama 3.1 405B Dense 405B 405B 128K tokens Mixtral 8x7B MoE, 8 experts, 2 active 46.7B 12.9B 32K tokens DeepSeek-V3 MoE, 256 routed plus 1 shared expert, 8 routed active 671B 37B 128K tokens Llama 4 Scout MoE, 16 experts 109B 17B 10M tokens
Figures come from the DeepSeek-V3 technical report and Meta’s Llama 4 announcement . Meta also lists Llama 4 Maverick at 400B total and 17B active across 128 experts. Scout and Maverick therefore share one active count, yet their stored size differs almost fourfold.
Why a Sparse Model Needs Two Numbers on Its Spec Sheet Active parameters set the compute per token, which drives speed and the cost of each response. Total parameters, by contrast, set how much memory the weights need, because every expert must sit in GPU memory even when it is idle. So the two numbers answer different questions.
That makes MoE attractive for high-volume serving on large GPU clusters but less attractive on a single small server. Real latency and memory still depend on hardware, precision and the serving stack. The Switch Transformer paper also notes the costs that come with routing, including complexity, communication between devices and training instability.
Routing is also learned rather than hand-assigned, so experts do not map neatly to topics such as finance or law. For routers, load balancing and when sparse models make sense, read our guide to MoE architecture . Models such as DeepSeek have since made the design common in open-weight releases.
Context Windows and the KV Cache The context window is the maximum number of tokens a model can consider at once, counting both the prompt and the reply. It has grown fast. Llama 3 supports up to 128K tokens, and Meta lists 10 million for Llama 4 Scout.
What Sets the Context Window Three things set the limit. The position scheme must handle distant positions, the model must have been trained on long sequences, and the hardware must hold the attention state. DeepSeek-V3, for example, was extended in two stages, first to 32K tokens and then to 128K.
Raising the limit is therefore a training and engineering decision, not a setting a buyer can change. When a vendor quotes a context length, ask how it was reached and how the model was tested at that length.
A Bigger Window Is Not Better Recall Even so, a model can accept 128K tokens and still miss a fact buried in the middle of them. Researchers behind the Lost in the Middle study found accuracy was often highest when relevant information sat at the start or end of the input. It dropped when the model had to use information from the middle, even for models built for long contexts.
Therefore, the practical lesson is to test recall at the lengths you plan to use, with your own documents. Often the better design is to send less rather than more, selecting only the passages that matter. That is the job of retrieval, covered in our retrieval-augmented generation guide, and of good context engineering .
Prefill and Decode Generation happens in two phases. First, in prefill, the model processes the whole prompt in parallel, which is compute-heavy and sets the time to the first token. Then, in decode, it produces the reply one token at a time, and each step is limited mostly by how fast memory can be read.
During decode, the new token needs the keys and values of every earlier token in every layer. Recomputing them at each step would waste enormous effort, so the model stores them in the KV cache and adds one entry per new token. The cache makes generation practical, but it grows with every token and every concurrent user.
A Worked KV Cache Calculation The size of the cache follows directly from the architecture.
To estimate it, multiply two (one key and one value) by the layer count and the KV head count. Then multiply by the head size and the bytes per stored number. The calculation below uses the dimensions Meta publishes for Llama 3 and assumes 16-bit storage.
kv_bytes_per_token = 2 * layers * kv_heads * head_dim * bytes_per_value
# Llama 3 8B: 32 layers, 8 KV heads, head_dim 128 (4,096 / 32), 16-bit
2 * 32 * 8 * 128 * 2 = 131,072 bytes (128 KiB per token)
128 KiB * 8,192 tokens = 1 GiB per conversation
same model, 32 KV heads = 4 GiB per conversation (no GQA)
# Llama 3 70B: 80 layers, 8 KV heads, head_dim 128 (8,192 / 64), 16-bit
2 * 80 * 8 * 128 * 2 = 327,680 bytes (320 KiB per token)
320 KiB * 131,072 tokens = 40 GiB for one full 128K-token requestNow put that next to the weights. A 70B model stored at 16 bits needs about 130 GiB for its parameters alone, and one maxed-out 128K conversation adds another 40 GiB. So just four such users at once would need more memory for their caches than for the model itself.
This is why grouped-query attention matters so much in practice. It cuts the 8B cache to a quarter of what full multi-head attention would need.
DeepSeek-V3 goes further with multi-head latent attention, which compresses keys and values. Serving engines also manage the cache in pages, an approach described in the PagedAttention paper and compared in our vLLM vs Ollama guide.
What LLM Architecture Means for Enterprise Model Selection All of this comes down to a question every platform team faces. Which model can do the job at an acceptable cost, speed and risk? The architecture section of a model card answers much of it before any benchmark is run.
Latency Comes From Two Different Places Users feel two delays. Time to first token comes from prefill, so it rises with prompt length and with the cost of full attention. Tokens per second comes from decode, so it depends on active parameters, cache size and memory bandwidth.
A model that feels slow on long documents may need shorter prompts rather than a bigger GPU. On the other hand, a model that types slowly may need fewer active parameters, which is where MoE or a smaller dense model helps. Our explainer on AI inference vs training covers why inference, not training, usually dominates the long-run bill.
Memory Is Weights Plus Cache Plan GPU memory as two buckets. On one side, the weights are a fixed cost set by total parameters and precision. On the other, the cache is a variable cost set by context length, KV heads and the number of people using the model at the same moment.
Consequently, teams that size only for the weights tend to hit memory errors the first week real users arrive. Teams running a private LLM on their own hardware feel this most, because they cannot borrow capacity from a provider at peak load.
The table below pulls these threads together. Each line on a model card maps to a cost you will pay in production.
Table 4: How to read an LLM spec sheet
Spec on the model card What it tells you What it costs you Total parameters Size of the stored weights GPU memory, about 2 bytes per parameter at 16-bit Active parameters Arithmetic per token Speed and cost of each response Layers and hidden size Depth and width of the stack Latency per token and cache size Attention heads and KV heads MHA, GQA or MQA design Cache per token, so users per GPU Context length Longest input accepted Prefill time and cache memory, not proof of recall Vocabulary and tokenizer How text is split Tokens per document, so the bill and window use Position scheme How word order is encoded How well quality holds at long lengths
Match the Layout to the Job Start from the task rather than the leaderboard. For example, high-volume classification, routing and embedding jobs rarely need a large decoder, and an encoder or small model is cheaper. By contrast, open-ended drafting, reasoning and agent work need a capable decoder-only model.
Checklist
Generative AI Checklist for Secure AI Adoption & Governance
Picking an architecture is only the first decision. Kanerika’s downloadable checklist helps teams confirm the data rules, security controls and acceptance criteria are in place before the chosen model goes into production.
Get the Checklist → Parameter count alone is also a poor guide. The Chinchilla study found that a 70B model trained on four times more data beat the 280B Gopher. A smaller, better-fed model can therefore win, and it then costs less on every request.
How that training is done belongs in our LLM training guide. Our list of open-source LLMs shows how many capable models now sit in the 7B to 70B range.
Likewise, long-document work needs a model tested for recall at your real lengths, plus retrieval to keep prompts lean. Busy shared assistants benefit from GQA or MLA designs, and from MoE where the cluster can hold the full weights. Our enterprise LLM guide turns these points into a selection scorecard, and our ranking of top LLMs shows how leading models compare.
Whatever the shortlist, still measure quality on your own data with a repeatable test set. Our LLM evaluation framework lays out the metrics, and the LLM hallucination guide explains why fluent next-token prediction can still produce confident errors.
Case Study
37% Faster Claims Processing With Generative AI
An Asian health and travel insurer used Kanerika’s generative AI data integration to cut claim processing time by 37% and fraud by 29%.
Read the Case Study → Common Misconceptions About LLM Architecture A few wrong assumptions turn up in almost every model evaluation. Each one leads to a bad comparison or a sizing mistake.
Treating tokens as words, when a document in another language or full of codes can need far more tokens than its word count suggests. Assuming every transformer has an encoder and a decoder, when most generative models are decoder-only. Confusing attention heads with layers, when heads run side by side inside each layer. Believing each MoE expert owns a subject, when the router learns its own, less tidy split. Comparing MoE models by total parameters alone, when active parameters drive speed and total parameters drive memory. Reading a maximum context length as a promise of accurate recall across the whole window. Assuming the KV cache is free, when at long lengths it can rival the weights in memory. Avoiding these seven traps takes no deep math. It takes one person on the evaluation team who reads the architecture table before the marketing page.
Where LLM Architecture Is Heading The transformer is not standing still, and three directions stand out. First, sparse designs are spreading, with MoE now common in large open-weight releases from DeepSeek and Meta. Second, attention is getting leaner through grouped-query and latent attention, mainly to shrink the KV cache.
Meanwhile, the third direction questions attention itself. State space models such as Mamba scale linearly with sequence length, and the authors report five times higher inference throughput than Transformers. Hybrid models that mix attention layers with state space layers are being tested as a middle path.
Further out, though, researchers are asking whether next-token prediction is enough at all. Our comparison of world models vs LLMs looks at that debate. The multimodal models guide covers how the same blocks now process images and audio.
How Kanerika Helps Enterprises Put the Right LLM Architecture to Work Kanerika treats model choice as an engineering decision with a budget attached. Our teams build on both Anthropic’s Claude and OpenAI’s models, and Kanerika is an OpenAI Select Partner. That lets us recommend the right model for each use case rather than defaulting to one vendor, including open-weight models where data must stay in-house.
A typical engagement through our LLM development services follows five stages.
Profile the workload by measuring real prompt lengths, token counts per document, request volume and latency targets. Shortlist by architecture, weighing dense against MoE, KV heads, tested context length and tokenizer fit for your languages. Benchmark the shortlist on your own documents, with recall tests at the context lengths you plan to use. Size the serving footprint as weights plus cache at peak concurrency, then choose hosting and precision. Govern and operate with evaluation, cost tracking and a review whenever a new model release changes the trade-offs. Proof in Business Numbers In the end, though, the results that count show up in business metrics. For a fast-growing ERP and CRM software provider, Kanerika built a ChatGPT-powered generative AI CRM dashboard . The client recorded a 22% uptick in KPI identification accuracy, a 14% boost in sales and revenue and a 10% increase in customer retention.
Once a model is picked, our wider generative AI services team connects the chosen model to data, retrieval and governance.
Case Study
22% Better KPI Accuracy With a Generative AI CRM Dashboard
Kanerika built a ChatGPT-powered CRM dashboard for an ERP provider, lifting sales and revenue by 14% and customer retention by 10%.
Read the Case Study → Wrapping Up LLM architecture can look like a wall of jargon, yet it follows one simple path. Tokens become vectors, transformer blocks refine them, attention lets each token borrow context, and the model predicts one token at a time.
Each step carries a cost.
Tokenizers set the bill, and attention sets the price of long prompts. MoE separates compute from memory, while the KV cache decides how many users a GPU can serve.
So read the architecture table first and test at the context lengths you will really use. Most bad model choices we see could have been caught at that stage, before any GPU was rented.
Frequently Asked Questions
What is LLM architecture? LLM architecture is the internal design of a large language model. It covers how text is split into tokens, how tokens become vectors, how stacked transformer blocks refine those vectors with attention and feed-forward layers, and how the model predicts the next token. These choices decide a model’s speed, memory needs and running cost.
What are the main components of an LLM? The main components are the tokenizer, the embedding table, positional encoding, a stack of transformer blocks and an output layer. Each transformer block combines self-attention, a feed-forward network, residual connections and normalization. Newer models may replace the feed-forward network with mixture-of-experts layers that activate only a few experts for each token.
What is the difference between a transformer and an LLM? A transformer is a neural network design built around self-attention, first described in 2017 for translation. An LLM is a specific model built with that design and trained on very large amounts of text to predict tokens. Almost every current LLM is a transformer, but not every transformer is a large language model.
Why do most generative LLMs use a decoder-only architecture? A decoder-only model learns one objective, predicting the next token from earlier tokens, and that objective works on almost any text. The same model can chat, summarize, write code and follow instructions without a separate encoder. A single stack is also simpler to scale and serve, which is why most generative models use it.
How does self-attention work in an LLM? Each token is projected into a query, a key and a value. The model compares one token’s query with every earlier token’s key to get relevance scores, scales them, and turns them into weights with a softmax. The token’s new representation is the weighted blend of the values, so it absorbs context from related words.
What is the difference between dense and mixture-of-experts LLMs? A dense model runs every weight for every token. A mixture-of-experts model replaces feed-forward layers with several experts and a router that picks a few per token. Mixtral 8x7B stores 46.7B parameters but uses 12.9B per token, so it computes like a smaller model while still needing memory for all experts.
What is a KV cache in an LLM? The KV cache stores the attention keys and values of every earlier token so the model does not recompute them for each new token. It makes generation much faster but grows with context length and with each concurrent user. For Llama 3 70B at 16-bit precision, one full 128K-token request needs about 40 GiB of cache.
Does a larger context window make an LLM more accurate? Not automatically. A larger window lets a model accept more text, but research such as the Lost in the Middle study found accuracy often drops for facts placed in the middle of long inputs. Long prompts also cost more to process. Test recall at your real lengths and send only the passages that matter.
What is grouped-query attention? Grouped-query attention lets several query heads share one set of keys and values instead of each head keeping its own. It keeps quality close to standard multi-head attention while shrinking the KV cache and speeding up generation. Llama 3 uses 8 key-value heads in every model size, which cuts cache memory sharply.
How does LLM architecture affect enterprise cost? Total parameters set the memory needed for weights, active parameters set the compute per token, and the attention design plus context length set the cache memory for each user. The tokenizer decides how many billable tokens a document becomes. Reading these specs together predicts hardware needs, latency and per-request cost before a pilot begins.