Pretraining is the initial large-scale training of a model on a broad data distribution before adaptation to a particular application. For a causal language model, training usually learns to predict the next token from preceding tokens. It changes model parameters; putting documents in a prompt does not.
In a Learnastra design interview, the first decision is whether building a foundation model is necessary. Most application designs start from an existing model and spend their effort on data, evaluation, and serving. Understanding pretraining still helps explain model limitations and adaptation choices.
The Pretraining Objective
For tokens x₁ … xₜ, causal language modeling minimizes negative log-likelihood:
L = -Σ log pθ(xᵢ | x₁ … xᵢ₋₁)
- Tokenize a training sequence and create shifted input/target pairs.
- Use a causal attention mask so a position cannot inspect its future target.
- Compute a distribution over the vocabulary at each predicted position.
- Compare those probabilities with the observed targets using cross-entropy.
- Backpropagate and update parameters with the optimizer.
Training can process many sequence positions in parallel because their preceding tokens are already known. Autoregressive generation produces new output tokens sequentially. This distinction explains why training throughput is not a direct prediction of interactive serving latency. Review tokenization and attention before estimating either.
Read diagram source
flowchart LR
A[Licensed raw data] --> B[Filter and deduplicate]
B --> C[Tokenize and sample batches]
C --> D[Causal prediction and loss]
D --> E[Gradient update]
E --> F[Checkpoint and held-out evaluation]
F --> C
A lower training loss means the model fits this objective better. It does not prove that it follows instructions, answers current factual questions, or behaves safely in a product.
Data Curriculum and Quality
A data mixture specifies what is sampled; a curriculum changes that sampling or task difficulty during training. There is no universally correct percentage of web, books, code, research, or synthetic text.
| Data decision | Potential benefit | What to verify |
|---|---|---|
| Remove exact and near duplicates | Reduces repeated examples and train/test leakage | Retain useful rare material; inspect false duplicate matches |
| Add domain text | Improves coverage of relevant language and patterns | Rights, freshness, domain evaluation, general-capability regression |
| Increase code or mathematical data | May improve particular structured tasks | Transfer must be demonstrated on held-out noncoding tasks |
| Add synthetic examples | Fills identified coverage gaps | Independent verification and diversity; generated does not mean correct |
| Change the late-stage mixture | Concentrates remaining compute on selected objectives | Compare with a constant-mixture baseline; avoid treating a named phase as a guarantee |
Store document provenance, filtering decisions, and dataset versions. Split evaluation by source or entity when related documents would otherwise leak into both training and test sets. Unknown proprietary data mixtures should remain unknown in an interview answer.
Scaling laws and lifecycle cost
Scaling laws are empirical relationships between model size, training data, compute, and loss under a particular experimental setup. The Chinchilla study investigated allocation of a fixed training-compute budget; its result is not a universal instruction to stop after exactly twenty tokens per parameter. Training Compute-Optimal Large Language Models.
A smaller model trained on more tokens may be attractive when it will serve many requests. The relevant comparison is:
Lifecycle cost = data + training + adaptation + serving + operations
Worked decision: suppose an additional training run costs $120,000 and reduces expected serving cost by $0.002 per request at the same measured quality. The simple break-even point is 120,000 / 0.002 = 60 million requests. Include refreshes, hosting utilization, evaluation, and engineering effort before making the real decision. These are exercise assumptions, not vendor prices.
Computational requirements and training stability
A rough dense-transformer training estimate is C ≈ 6ND floating-point operations, with N parameters and D training tokens. It omits important architecture, attention, and systems details. For an illustrative 1-billion-parameter model trained on 20 billion tokens, it gives 1.2 × 10²⁰ FLOPs. At a measured aggregate effective throughput of 10¹⁵ FLOP/s, this portion takes about 120,000 seconds, or 33.3 hours. Peak accelerator specifications are not effective throughput.
| Risk | Investigation and response | Cost of the response |
|---|---|---|
| Loss spike | Inspect batches, gradient norms, learning rate, overflow, and data corruption; resume from a known checkpoint if needed | Lost work and checkpoint storage |
| Out-of-memory failure | Account for weights, gradients, optimizer state, activations, and communication buffers | Sharding or recomputation adds communication or compute |
| Low utilization | Inspect input starvation, padding, synchronization, and hardware failures | More complex data and distributed-training pipelines |
| Numerical instability | Validate precision, scaling, optimizer settings, and accumulation behavior | Higher precision may reduce throughput or available batch size |
BF16 and FP8 describe numerical formats, not fixed speedups. Supported kernels and scaling determine their useful behavior. FP8 does not automatically halve total training memory because many tensors and optimizer states may use other formats. Residual connections add branch outputs to a running hidden representation. Some architectures scale residual branches or their initialization to control activation and gradient growth with depth; use the model's documented recipe rather than applying an arbitrary universal scale. Checkpoint recovery should restore optimizer, scheduler, random state, and data position as well as model weights.
Interview practice
- Why is pretraining not a freshness mechanism? Parameters do not automatically change when a source document changes. Use a maintained retrieval pipeline when updates and citations are central.
- Can the model see the answer during training? The training example contains target tokens, but the causal mask must prevent a prediction position from attending to its future target.
- Why train a smaller model longer? Repeated serving savings may justify additional one-time training cost; compare quality and lifecycle cost.
- Does twice the peak GPU throughput halve training time? Only if the relevant computation dominates and the system can use that throughput. Data, memory, and communication can dominate instead.
- What does a loss spike require first? Diagnose the data and numerical behavior. Blindly lowering learning rate or rolling back may hide a recurring defect.
- How would you test a new data mixture? Hold evaluation and compute comparisons consistent, inspect important slices, and check both target gains and retained capabilities.
Recall card and closing
Objective → data → compute → evaluation → lifecycle cost. Explain what the model learns, how the data represents the target workload, and why the training investment is justified. Close by naming what the training objective cannot guarantee for the application.
Continue with fine-tuning, synthetic data, and inference fundamentals.