Skip to content
AI-grafen
FAI engineeringModel training and fine-tuning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Merging adapters and exporting models

Be able to merge LoRA weights, export to GGUF/safetensors and run locally.

Prerequisites

Intuition

A trained LoRA adapter is a few tens of MB and requires the base model. For distribution you often want one file.

The merge: W′ = W + (α/r)·BA. After that the adapter is gone and the model is an ordinary model — no extra inference cost, no PEFT dependency.

The formats:

The formatUsed byNote
safetensorstransformers, vLLM, TGIthe standard; safe (no pickle)
GGUFllama.cpp, Ollama, LM Studioa single file, CPU/Metal, quantised at conversion
ONNXONNX Runtime, edgegood portability, weaker LLM support

A warning: do not merge an adapter trained against a 4-bit base with an fp16 base straight off. The adapter learnt to compensate for the quantisation error; the result becomes slightly worse. Evaluate after the merge, not before.

Code

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B", torch_dtype=torch.bfloat16, device_map="cpu")
m = PeftModel.from_pretrained(base, "adapter/")
m = m.merge_and_unload()                      # W += (alpha/r) * B @ A
m.save_pretrained("model-merged", safe_serialization=True)
AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B").save_pretrained("model-merged")
# GGUF for running locally
python convert_hf_to_gguf.py model-merged --outfile model-f16.gguf
./llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M
./llama-cli -m model-Q4_K_M.gguf -p "Explain gradient descent briefly." -n 128

A verification step that is often skipped — always do it:

prompts = [...]                                # 20 representative prompts
before = [generate(m_with_adapter, p, temperature=0) for p in prompts]
after  = [generate(m_merged, p, temperature=0) for p in prompts]
print(sum(a == b for a, b in zip(before, after)), "/", len(prompts), "identical")
# After that: run the whole eval suite on the EXPORTED format, not on the original.

A model that behaves differently after export is a common and silent fault — often because of a change of precision or a chat template that did not come along.

Mastery means

  • Merges the LoRA weights into the base model
  • Exports to safetensors and GGUF
  • Verifies that the exported model gives the same answers

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences