This tutorial walks through one complete FlagOS-Compressor path: inspect a local checkpoint, verify the quantization plan with a dry-run, generate a W8A16 artifact, and validate the output directory.
Before you begin, make sure you have:
- Python 3.10 or later
- PyTorch installed for the backend you plan to use
- FlagOS-Compressor installed. See the installation guide
- a local HuggingFace
safetensorscheckpoint directory - an empty output location with enough space for generated shards
The input checkpoint should include at least:
config.json
model.safetensors.index.json
model-00001-of-000xx.safetensors
...
Do not write output into the source model directory. Use a new empty directory for each test so old shards and old configs cannot mix with new output.
Set environment variables for the source model and output root:
MODEL_QWEN_DENSE=/data/models/qwen3.5-dense
MODEL_OUTPUT_ROOT=/data/models/flagos-compressor-testsReplace these paths with directories that exist in your environment.
Run inspect before conversion or quantization:
flagos-compressor inspect --input "$MODEL_QWEN_DENSE"Check the inspection result before continuing:
- recognized linear layer counts match the model structure
- source weight formats are recognized as BF16, FP16, FP32, MXFP4, or supported block FP8
- Dense models do not show routed expert counts
- MoE models show expected routed or shared expert groups
For a structured record, save JSON output:
flagos-compressor inspect \
--input "$MODEL_QWEN_DENSE" \
--json > qwen3.5-dense-inspect.jsonRun the same quantization command with --dry-run first:
flagos-compressor quantize \
--input "$MODEL_QWEN_DENSE" \
--output "$MODEL_OUTPUT_ROOT/qwen3.5-dense-w8a16-g128" \
--select linear \
--bits 8 \
--strategy group \
--group-size 128 \
--backend cuda \
--dry-runReview the plan before writing weights:
- selected weight counts match expectations
- W/A bits, strategy, and group size are correct
- no selector is reported as unmatched
- no fused unit or routed expert bank is split
unmatched quantizedis0for supported source-quantized inputs
If CUDA is not available, use --backend cpu for a functional check.
Remove --dry-run and run the command again:
flagos-compressor quantize \
--input "$MODEL_QWEN_DENSE" \
--output "$MODEL_OUTPUT_ROOT/qwen3.5-dense-w8a16-g128" \
--select linear \
--bits 8 \
--strategy group \
--group-size 128 \
--backend cudaDuring generation, monitor for:
- unexpected backend fallback
- shape or group-size errors
- output file creation failures
Run validation on the output directory:
flagos-compressor validate \
--input "$MODEL_OUTPUT_ROOT/qwen3.5-dense-w8a16-g128"Expected result:
Valid
Save a structured validation report when you need an audit trail:
flagos-compressor validate \
--input "$MODEL_OUTPUT_ROOT/qwen3.5-dense-w8a16-g128" \
--json > qwen3.5-dense-w8a16-g128-validate.jsonYou now have a validated W8A16 output directory containing model config, shard index, weight shards, quantization metadata, report files, and tokenizer files copied from the source model.
Use the output validation guide to continue with runtime loading checks and platform-specific validation.