Benchmarks
Compression numbers on representative real-world checkpoints.
Setup
- Corpus
openai/gpt-oss-20b— MXFP4-quantized MoE, 13.76 GB raw (FP4 nibbles + E8M0 scales + BF16 layers).nvidia/Llama-3_3-Nemotron-Super-49B-v1_5-NVFP4— NVFP4-quantized Llama derivative, 31.07 GB raw (FP4 nibbles + FP8-E4M3 scales).
- Configurations
huffman_only— PTWM with the preprocessing graph forced to passthrough; pure byte-Huffman per tensor. Isolates the contribution of the preprocessing graph.microscale_full— PTWM with the production preprocessing graph and theMICROSCALEmethod enabled.
- Hardware — CPU-only host, 16 cores, 64 GB RAM.
Compression ratios
Ratio is compressed bytes ÷ raw bytes. Lower is better.
| Model | huffman_only | microscale_full |
|---|---|---|
openai/gpt-oss-20b | 0.8804 | 0.8603 |
nvidia/Llama-3_3-Nemotron-Super-49B-v1_5-NVFP4 | 0.9089 | 0.9003 |
Combined-corpus ratio (44.83 GB raw): microscale_full lands at 0.8881.
Throughput
The encode and decode pipelines run across a rayon thread pool. Full-model
throughput through the safetensors file integration on a 32-core CPU host:
| Model | dtype | Encode | Decode |
|---|---|---|---|
mistralai/Mistral-7B-v0.1 | bf16 | 45 MB/s | 1111 MB/s |
openai/gpt-oss-20b | mxfp4 | 77 MB/s | 1635 MB/s |
nvidia/Nemotron-Super-120B-FP8 | fp8 | 26 MB/s | 627 MB/s |
Qwen/Qwen2.5-72B-Instruct | bf16 | 24 MB/s | 793 MB/s |
Decode — the model-load hot path — runs at 0.6–1.6 GB/s: every tensor in a
shard is decoded across the pool in a single call. Encode favours ratio over
speed by trial-encoding the full codec menu per plane; the opt-in
compress_safetensors_file(fast=True) mode forces a single codec for a 4–8×
encode speedup at parity ratio on float models.
See pipeline validation for the full dtype, scale, memory, and edge-case matrix (including the fast-mode and thread-count dials that bound peak encode memory).
Key findings
- Microscaling formats are already near-incompressible at the byte level. Trained FP4 / NVFP4 nibble-packed bytes are close to uniform; no amount of plane-aware coding extracts much further.
- The remaining lossless headroom is in the scale plane. The
MICROSCALEmethod's Order-1 ScaleAC codec compresses E8M0 scales to ≈ 0.16 of original. The whole-model effect is bounded by the scale-mass fraction (≈ 5%), which is why the on-disk gain over byte-Huffman is small in absolute terms. - The PTWM preprocessing graph adds value over pure byte-Huffman,
but only modestly on microscaling formats.
huffman_only(graph forced to passthrough) hits 0.8804 on gpt-oss-20b vs 0.8603 with the full graph — a 2.0 percentage-point delta, most of which is the scale codec; the rest is dtype-aware nibble routing.
On-disk vs. algorithmic size
The numbers above report on-disk .ptwm directory size, including
the multi-tensor container header, tensor index, per-plane codec
dispatch records, and payload hashes. Container overhead runs roughly
1.6 percentage points on gpt-oss-20b, the price of features the raw
entropy coders lack (random access, hash verification, multi-tensor
packaging).
For a pure entropy-coder comparison, subtract the container overhead:
microscale_full on gpt-oss-20b drops from 0.8603 to roughly 0.844
algorithmically. The on-disk number is reported here because it
matches what users measure with du.
Preprocessing ablations
Effect of preprocessing-graph variants on the byte-Huffman ratio across
dtypes. Higher ratio is better (compressed bytes ÷ raw bytes is the
inverse of huf_ratio).
| Category | Model | Dtype | Ablation | Variant | Bytes | N tensors | huf_ratio | Plane entropies (bits/byte) |
|---|---|---|---|---|---|---|---|---|
| audio | facebook/encodec_24khz | fp32 | A1 | fp32_reorder_on | 92,815,360 | 98 | 1.0701 | 8.00 / 8.00 / 7.97 / 3.25 |
| audio | facebook/encodec_24khz | fp32 | A1 | fp32_reorder_off | 92,815,360 | 98 | 1.0690 | 8.00 / 8.00 / 7.97 / 3.30 |
| vision | microsoft/resnet-50 | fp32 | A1 | fp32_reorder_on | 102,011,648 | 54 | 1.0588 | 8.00 / 8.00 / 7.97 / 2.68 |
| vision | microsoft/resnet-50 | fp32 | A1 | fp32_reorder_off | 102,011,648 | 54 | 1.0579 | 8.00 / 8.00 / 7.97 / 2.75 |
| nlp | distilbert/distilbert-base-uncased | fp32 | A1 | fp32_reorder_on | 267,687,144 | 40 | 1.0379 | 8.00 / 8.00 / 7.93 / 2.59 |
| nlp | distilbert/distilbert-base-uncased | fp32 | A1 | fp32_reorder_off | 267,687,144 | 40 | 1.0372 | 8.00 / 8.00 / 7.97 / 2.61 |
| nlp | HuggingFaceTB/SmolLM2-135M | bf16 | A1 | bf16_reorder_on | 268,959,744 | 211 | 1.0665 | 7.97 / 2.66 |
| nlp | HuggingFaceTB/SmolLM2-135M | bf16 | A1 | bf16_reorder_off | 268,959,744 | 211 | 1.0652 | 7.97 / 2.74 |
| quantised | RedHatAI/Qwen2-1.5B-Instruct-FP8 | bf16 | A1 | bf16_reorder_on | 466,747,392 | 1 | 1.0000 | 7.97 / 2.59 |
| quantised | RedHatAI/Qwen2-1.5B-Instruct-FP8 | bf16 | A1 | bf16_reorder_off | 466,747,392 | 1 | 1.0000 | 7.97 / 2.67 |
| quantised | RedHatAI/Qwen2-1.5B-Instruct-FP8 | fp8-e4m3fn | A3 | e4m3_passthrough | 1,310,195,712 | 196 | 1.0000 | 6.64 |
| quantised | RedHatAI/Qwen2-1.5B-Instruct-FP8 | fp8-e4m3fn | A3 | e4m3_reorder_only | 1,310,195,712 | 196 | 1.0000 | 6.64 |
| quantised | RedHatAI/Qwen2-1.5B-Instruct-FP8 | fp8-e4m3fn | A3 | e4m3_nibble_split | 1,310,195,712 | 196 | 0.5048 | 2.70 / 3.98 |
| quantised | RedHatAI/Qwen2-1.5B-Instruct-quantized.w8a8 | bf16 | A1 | bf16_reorder_on | 934,498,304 | 58 | 1.0006 | 7.97 / 2.60 |
| quantised | RedHatAI/Qwen2-1.5B-Instruct-quantized.w8a8 | bf16 | A1 | bf16_reorder_off | 934,498,304 | 58 | 1.0005 | 7.97 / 2.68 |
The FP8-E4M3 row (e4m3_nibble_split) demonstrates the dtype-aware
nibble routing: splitting the byte into two 4-bit planes drops entropy
from 6.64 to ~3.3 bits/byte per plane, yielding a 2× compression ratio
on a dtype that is otherwise incompressible byte-wise.
Plane entropy
Per-plane entropy on facebook/encodec_24khz (fp32). The reorder
strategy moves the IEEE-754 exponent bits into a dedicated plane,
exposing the trained-weight exponent distribution (≈ 3.25 bits/byte)
that is otherwise scattered across byte boundaries.
| Strategy | Plane | Entropy (bits/byte) |
|---|---|---|
| raw | all | 7.467 |
| shuffle | byte0 | 8.000 |
| shuffle | byte1 | 8.000 |
| shuffle | byte2 | 7.971 |
| shuffle | byte3 | 3.301 |
| reorder | mantissa_lo | 8.000 |
| reorder | mantissa_mid | 8.000 |
| reorder | mantissa_hi | 7.971 |
| reorder | exponent | 3.251 |
Method comparison (synthetic)
Compression ratio and throughput on synthetic uniform-random data across the four high-level PTWM methods. Ratios above 1.0 reflect compression; values at or below 1.0 reflect the method passing through (or in the FP8 cases, the raw input being smaller than the container's per-tensor overhead at 500 kB).
| Dtype | Bytes | Method | Compressed | Ratio | Compress (MB/s) | Decompress (MB/s) |
|---|---|---|---|---|---|---|
| float32 | 2,000,000 | IDENTITY | 2,000,357 | 0.9998 | 533.4 | 612.6 |
| float32 | 2,000,000 | HUFFMAN | 1,673,915 | 1.1948 | 661.0 | 1155.0 |
| float32 | 2,000,000 | RANS | 1,661,813 | 1.2035 | 378.3 | 1068.7 |
| float32 | 2,000,000 | ZSTD | 1,706,934 | 1.1717 | 417.8 | 1055.2 |
| bfloat16 | 1,000,000 | IDENTITY | 1,000,141 | 0.9999 | 543.8 | 933.3 |
| bfloat16 | 1,000,000 | HUFFMAN | 673,395 | 1.4850 | 462.2 | 753.1 |
| bfloat16 | 1,000,000 | RANS | 660,441 | 1.5141 | 265.3 | 654.8 |
| bfloat16 | 1,000,000 | ZSTD | 705,920 | 1.4166 | 331.2 | 750.2 |
| float16 | 1,000,000 | IDENTITY | 1,000,141 | 0.9999 | 665.9 | 1180.4 |
| float16 | 1,000,000 | HUFFMAN | 847,603 | 1.1798 | 438.5 | 836.5 |
| float16 | 1,000,000 | RANS | 845,157 | 1.1832 | 263.1 | 656.9 |
| float16 | 1,000,000 | ZSTD | 848,612 | 1.1784 | 349.2 | 892.6 |
| fp8-e4m3fn | 500,000 | IDENTITY | 1,000,105 | 0.4999 | 283.9 | 1077.2 |
| fp8-e4m3fn | 500,000 | HUFFMAN | 410,981 | 1.2166 | 222.9 | 220.9 |
| fp8-e4m3fn | 500,000 | RANS | 406,754 | 1.2292 | 96.4 | 176.6 |
| fp8-e5m2 | 500,000 | IDENTITY | 1,000,105 | 0.4999 | 346.2 | 913.3 |
| fp8-e5m2 | 500,000 | HUFFMAN | 355,703 | 1.4057 | 206.3 | 264.2 |
| fp8-e5m2 | 500,000 | RANS | 351,351 | 1.4231 | 96.7 | 172.1 |
| int8 | 500,000 | IDENTITY | 500,105 | 0.9998 | 1266.2 | 2276.0 |
| int8 | 500,000 | HUFFMAN | 500,105 | 0.9998 | 624.8 | 986.3 |
| int8 | 500,000 | RANS | 500,105 | 0.9998 | 236.1 | 1762.7 |
| int16 | 1,000,000 | IDENTITY | 1,000,141 | 0.9999 | 752.2 | 981.8 |
| int16 | 1,000,000 | HUFFMAN | 1,000,141 | 0.9999 | 428.8 | 894.3 |
| int16 | 1,000,000 | RANS | 1,000,141 | 0.9999 | 241.3 | 765.0 |
| int32 | 2,000,000 | IDENTITY | 2,000,357 | 0.9998 | 991.4 | 694.8 |
| int32 | 2,000,000 | HUFFMAN | 1,938,945 | 1.0315 | 715.3 | 574.7 |
| int32 | 2,000,000 | RANS | 1,943,599 | 1.0290 | 399.7 | 613.4 |
| int64 | 4,000,000 | IDENTITY | 4,001,221 | 0.9997 | 1110.3 | 1019.3 |
| int64 | 4,000,000 | HUFFMAN | 2,199,601 | 1.8185 | 819.7 | 882.6 |
| int64 | 4,000,000 | RANS | 2,232,699 | 1.7916 | 561.0 | 809.5 |
Codec comparison (real models)
Per-codec ratio and throughput on real fp32 models, holding the preprocessing graph fixed at reorder+split for the Huffman and rANS rows and varying the ZSTD prefilter across the rest.
| Model | Codec | Compressed | Ratio | Compress (MB/s) | Decompress (MB/s) |
|---|---|---|---|---|---|
facebook/encodec_24khz | Huffman (reorder+split) | 86,736,647 | 1.0701 | 205.3 | 1375.9 |
facebook/encodec_24khz | rANS (reorder+split) | 86,451,110 | 1.0736 | 119.5 | 928.8 |
facebook/encodec_24khz | ZSTD (reorder+split) | 79,498,276 | 1.1675 | 158.1 | 1276.6 |
facebook/encodec_24khz | ZSTD (shuffle) | 79,409,482 | 1.1688 | 67.5 | 541.2 |
facebook/encodec_24khz | ZSTD (bitshuffle) | 80,272,706 | 1.1563 | 204.7 | 562.6 |
facebook/encodec_24khz | ZSTD (bytedelta) | 92,818,496 | 1.0000 | 420.9 | 699.9 |
microsoft/resnet-50 | Huffman (reorder+split) | 96,346,850 | 1.0588 | 168.1 | 1282.5 |
microsoft/resnet-50 | rANS (reorder+split) | 96,062,980 | 1.0619 | 131.9 | 938.8 |
microsoft/resnet-50 | ZSTD (reorder+split) | 86,884,859 | 1.1741 | 148.4 | 1133.0 |
microsoft/resnet-50 | ZSTD (shuffle) | 86,536,651 | 1.1788 | 59.7 | 707.5 |
microsoft/resnet-50 | ZSTD (bitshuffle) | 87,549,748 | 1.1652 | 222.4 | 804.0 |
microsoft/resnet-50 | ZSTD (bytedelta) | 102,013,376 | 1.0000 | 559.2 | 1094.1 |
Random-access tradeoff
Per-tensor compression preserves random access but pays a small ratio
penalty against monolithic compression. On distilbert-base-uncased
the gap is negligible for ZSTD (1.1811 vs 1.1803) and reversed for
Huffman and rANS, where the per-tensor coder benefits from
dtype-aware preprocessing while the monolithic stream cannot
recognise tensor boundaries.
| Coder | Strategy | Compressed | Ratio | Compress (MB/s) |
|---|---|---|---|---|
| huffman | per_tensor | 257,921,169 | 1.0379 | 336.2 |
| huffman | monolithic | 267,687,148 | 1.0000 | 260.5 |
| rans | per_tensor | 257,426,023 | 1.0399 | 276.4 |
| rans | monolithic | 267,687,148 | 1.0000 | 203.9 |
| zstd | per_tensor | 226,638,984 | 1.1811 | 278.2 |
| zstd | monolithic | 226,799,862 | 1.1803 | 209.3 |
Source data for distilbert/distilbert-base-uncased
(267,687,144 raw bytes, 40 fp32 tensors).