A visual guide to the technical building blocks behind open-weight AI models — from tensors and parameters to formats, precision, sharding, quantization, adapters and deployment.
Model weights are the learned numerical parameters of a machine-learning model. During training, an optimization process repeatedly adjusts these values so that the model becomes better at its objective.
The architecture defines how the model is structured. The weights contain the learned values used by that architecture.
Illustrative tensor values — real models contain millions or billions of parameters.
Weights are one layer in a larger technical pipeline.
The learned numerical values that encode model behavior.
How tensors and metadata are stored and distributed.
How many bits are used to represent numerical values.
Splitting very large model artifacts across multiple files.
Which base model, revision and transformations produced the artifact.
The software layer that loads and executes the weights.
Designed for safe and efficient tensor serialization and deeply integrated into the Hugging Face ecosystem.
Tensor storageMetadataFast loadingWidely used in local-inference workflows in the llama.cpp ecosystem, including quantized models.
Local inferenceQuantizationPortable runtimesA file format is not the model architecture itself. It is the container or serialization layer used to store model tensors and associated metadata.
Numerical precision affects storage size, memory demand, hardware compatibility and sometimes model quality.
Illustrative relative storage intuition only. Real memory use also depends on runtime overhead, KV cache, activations, batching and architecture.
Quantization reduces the numerical precision used for weights or activations. It can make larger models practical on less expensive or more constrained hardware.
Large models are often split into multiple weight files called shards.
Sharding makes very large artifacts easier to distribute and load, but it also increases the importance of metadata, indexes and completeness checks.
Updates model parameters directly for a new objective or domain.
Trains smaller parameter-efficient components while keeping base weights largely frozen.
Combines adapters or derived changes into a new model artifact.
Open-weight engineering is not only about downloading a checkpoint. A trustworthy deployment should preserve the origin and transformation history of the artifact.
| Field | Why it matters |
|---|---|
| Base model | Identifies the original model family. |
| Revision | Links deployment to a specific repository state. |
| Fine-tune / adapter | Shows whether the model has been adapted. |
| Quantization | Documents numerical transformation of the artifact. |
| Format | Identifies how tensors are serialized. |
| Checksum / hash | Helps verify artifact identity. |
| License | Defines permitted use and redistribution conditions. |
The same model may behave differently across runtimes and hardware. Compatibility is a technical property that changes over time.
Architecture support, weight format, precision, quantization and hardware kernels all influence portability.
| Layer | Question | Example |
|---|---|---|
| Weights | What was learned? | Model parameters |
| Architecture | How are parameters organized? | Transformer |
| Format | How are tensors stored? | Safetensors, GGUF |
| Precision | How are values represented? | BF16, FP16, INT8 |
| Sharding | How are large files split? | Multiple weight shards |
| Quantization | How is the artifact reduced? | INT8 / INT4 |
| Adapter | How is it specialized? | LoRA / PEFT |
| Runtime | How is it executed? | Transformers, vLLM, llama.cpp |
| Hardware | Where does it run? | GPU, CPU, Apple Silicon |
| Provenance | Where did it come from? | Revision + transformation history |
Compare tensor serialization and model formats.
Understand precision, memory and runtime trade-offs.
Map formats, runtimes, hardware and compatibility.