Model terms, explained
You do not need to be an ML researcher to choose a model deployment. Here is what the terms mean and which questions they help you answer.
On this page
Serving engine: the software that runs the model
The model contains the learned weights and a definition of its calculations. A serving engine is the software that loads it, executes those calculations on the hardware, and handles requests. The surrounding server typically exposes an API for your application.
Inference means running a trained model to produce an output, such as an answer. It is different from training, which changes the learned parameters. The engine, hardware and request settings can affect memory use and response time, so a model name alone does not describe a deployment.
Tokens and context: how much text the model works with
A token is a unit the model uses to represent text. It may be a word, part of a word, punctuation or another text fragment. The tokenizer determines the split. The same sentence can have different token counts in different models, and language affects that count too.
A prompt is the input you give the model: instructions, a question and any supporting context. Changing the prompt changes the input; it does not train the model.
The context window is the limit on tokens the model can work with in a request. Input instructions, conversation history, documents, tool results and the generated answer all need room within the applicable limit. A large context window does not guarantee that the model will use every detail correctly.
Ask: will our real inputs and expected answers fit, and does the model answer correctly when the relevant detail is buried in a long input?
Streaming: showing the answer as it arrives
Streaming sends pieces of the response to your application while the model is generating. The user can start reading before the full answer is ready. Without streaming, the application waits for the complete response.
Streaming changes how the answer is delivered. It does not by itself make the model generate faster. Your application also needs to handle partial text, interrupted responses and errors.
Latency, TTFT and throughput: different performance questions
- Latency
- Elapsed time for a request. State what is timed: the first visible text, the full answer, or the complete application task.
- TTFT
- Time to first token: how long the client waits for the first generated token. Queueing, processing the input and network delivery can all contribute.
- Generation speed
- How quickly output tokens arrive after generation starts, often expressed as tokens per second. It does not include all the waiting before the first token.
- Throughput
- How much work the system completes over time across requests. A server can process more requests overall while each user waits longer.
Ask: what did you measure, with which input and answer lengths, and how many requests were running at once? Compare results with the same measurement boundaries.
Fine-tuning: further training for a particular task
Fine-tuning starts with an existing model and trains it further on selected examples. It can target task behavior, terminology or an output format. Depending on the method, training updates model weights or smaller added sets of parameters called adapters.
It is different from putting instructions or documents in a prompt. For information that changes often, retrieving relevant documents at request time may be a better fit. Before training, compare against a good prompt and evaluate on examples kept out of the training data.
Frontier models: a broad description, not a fit test
Frontier models generally means models near the leading edge of capabilities at a given time. Definitions vary by context, and the leading edge moves. The term alone does not tell you whether a model is good at your task, affordable to run or available under a suitable license.
Ask: which exact model and version, measured on what task? For a deployment decision, those answers matter more than the label.
Speculative decoding: propose text, then verify it
Ordinary text generation advances one token at a time. With speculative decoding, a cheaper process proposes several possible next tokens, and the target model checks them together. Accepted tokens can save work; rejected proposals cost time too. Whether it helps depends on the model, software and workload.
Draft agreement describes how often the proposed tokens match or are accepted by the target model, depending on the measurement. It is a measure of how useful the draft is for acceleration. It does not measure whether the final answer is factually correct or whether generated code works.
Exact speculative algorithms are designed to preserve the target model's output distribution when implemented correctly. That property still needs verification in the running system, alongside a speed comparison.
Research paper: Fast Inference from Transformers via Speculative Decoding
Pruning: removing parts of the model
Pruning removes or zeros selected weights or groups of model components. Quantization instead stores values with less numerical precision. Both can reduce resource needs, but they change the model in different ways.
Zeros alone do not guarantee a smaller file or faster execution. The storage format and software must be able to take advantage of what was removed. As with quantization, check task quality and performance after the change.