Merging adapters and exporting models
Be able to merge LoRA weights, export to GGUF/safetensors and run locally.
Prerequisites
- FQuantisationrequired
- FLoRA — Low-Rank Adaptationrequired
Intuition
A trained LoRA adapter is a few tens of MB and requires the base model. For distribution you often want one file.
The merge: W′ = W + (α/r)·BA. After that the adapter is gone and the model is an ordinary model — no extra inference cost, no PEFT dependency.
The formats:
| The format | Used by | Note |
|---|---|---|
| safetensors | transformers, vLLM, TGI | the standard; safe (no pickle) |
| GGUF | llama.cpp, Ollama, LM Studio | a single file, CPU/Metal, quantised at conversion |
| ONNX | ONNX Runtime, edge | good portability, weaker LLM support |
A warning: do not merge an adapter trained against a 4-bit base with an fp16 base straight off. The adapter learnt to compensate for the quantisation error; the result becomes slightly worse. Evaluate after the merge, not before.
Code
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B", torch_dtype=torch.bfloat16, device_map="cpu")
m = PeftModel.from_pretrained(base, "adapter/")
m = m.merge_and_unload() # W += (alpha/r) * B @ A
m.save_pretrained("model-merged", safe_serialization=True)
AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B").save_pretrained("model-merged")
# GGUF for running locally
python convert_hf_to_gguf.py model-merged --outfile model-f16.gguf
./llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M
./llama-cli -m model-Q4_K_M.gguf -p "Explain gradient descent briefly." -n 128
A verification step that is often skipped — always do it:
prompts = [...] # 20 representative prompts
before = [generate(m_with_adapter, p, temperature=0) for p in prompts]
after = [generate(m_merged, p, temperature=0) for p in prompts]
print(sum(a == b for a, b in zip(before, after)), "/", len(prompts), "identical")
# After that: run the whole eval suite on the EXPORTED format, not on the original.
A model that behaves differently after export is a common and silent fault — often because of a change of precision or a chat template that did not come along.
Mastery means
- Merges the LoRA weights into the base model
- Exports to safetensors and GGUF
- Verifies that the exported model gives the same answers
Sign in to do the exercises and build your mastery up.
Sources
- PEFT — dokumentation (Apache-2.0) — Apache-2.0
- llama.cpp (MIT) — MIT