Glyd
Lossless AI compression: 33% less GPU memory, bit for bit. Rust, C ABI, CLI; v0.26.0, October 2026. Every number below is measured on public data and reproducible from the repository.
What would it save you? Quick startFastest or slowest? The table everyone uses
The 8.7 GB real-data corpus on AWS Graviton3, ratio · compress MB/s · decompress MB/s. One core, the way zstd's README reports it, then eight cores (Glyd's output decodes in parallel; a zstd or LZ4 frame decodes on one thread).
| One core | Ratio | Compress MB/s | Decompress MB/s |
|---|---|---|---|
| LZ4 | 2.72 | 504 | 1,394 |
| Glyd default | 2.82 | 332 | 3,309 |
| zstd -3 | 3.86 | 313 | 1,425 |
| Glyd --max | 3.98 | 233 | 1,571 |
| Eight cores | Ratio | Compress MB/s | Decompress MB/s |
|---|---|---|---|
| Glyd default | 2.82 | 2,143 | 22,707 |
| zstd -3 -T8 | 3.85 | 1,969 | 1,422 |
| Glyd --max | 3.94 | 1,512 | 10,413 |
| Glyd --max -r | 4.71 | 643 | 4,040 |
| LZ4 | 2.72 | 503 | 1,392 |
| zstd -19 -T8 | 4.66 | 13 | 1,326 |
| Glyd --ultra | 4.66 | 14.5 | 9,620 |
| Glyd --ultra -r | 5.22 | 20 | 3,927 |
Reads: Glyd is the fastest, 3–7× zstd on a server. Writes: Glyd is slower, 0.77× zstd -3 (record mode 0.33×); LZ4 is the write-speed king on one core. Size: Glyd is the smallest wherever the data has structure, and ties elsewhere.
Against zstd, xz and brotli on 24 kinds of data
Every codec's own CLI on one thread, every decode compared byte for byte with its input. Green: Glyd's file is smaller.
| Data | vs zstd -3 (--max) | vs zstd -19, xz, brotli -11 (--ultra) |
|---|---|---|
| Records: logs, dumps, JSON (-r) | 12–61% fewer bytes | 3–38% fewer (JSON: 9% more) |
| Containers: gzip, zip, Office, PDF, PNG, JPEG | up to 60% fewer | 14–63% fewer |
Glyd wins where the data has structure: each field of a record becomes a column, and the deflate, snappy, zstd or JPEG inside a container is opened and re-created bit for bit, Parquet pages included. On plain text and executables it is zstd-class: within 2% at the fast tier, 1–15% larger than xz and brotli -11 at their strongest. Every codec's speed, the strong-tier and ratio-against-speed charts, and lz4, bzip2, zpaq and JPEG XL.
The store: compression across objects, not just inside them
Inside one object every codec sits on the same floor. The redundancy of object storage is between objects — builds, snapshots, dumps and releases that are near-copies of earlier ones. glyd-store bucket/ --put ... keeps each object as a delta against the stored object it most resembles when that pays; a read is at most five decodes. Measured on a 39 GB bucket, each object arriving in order, every one read back and compared:
| Family | Raw | zstd -3, each object alone | Glyd store | Gain |
|---|---|---|---|---|
| Linux 6.10 releases (15) | 22.5 GB | 3,237 MB | 232 MB | 13.9× smaller |
| Ubuntu 24.04 cloud images (6 builds) | 6.6 GB | 1,880 MB | 361 MB | 5.2× |
| Wikipedia dumps (2 months, 3 tables) | 0.7 GB | 140 MB | 71 MB | 2.0× |
| GitHub events (12 hours) | 9.4 GB | 875 MB | 670 MB | 1.3× (nothing is a version of anything) |
| The bucket | 39.2 GB | 6,132 MB (6.4×) | 1,334 MB (29.4×) | 4.6× smaller |
Put runs at 620 MB/s end to end on ten cores; delete, compact and verify are there. A petabyte of such data in S3 Standard: $43K a year with zstd -3, $9.4K with the store. Chunk-level dedup, the backup approach, gains 1–4× on the same pairs.
Base mode: a version stored for the change, not the whole
Most stored bytes are versions — nightly dumps, snapshots, images, source trees. glyd --base old new parses every 32 MB of the new version with the old one's matching region as history and writes a stream that decodes with the same base. Measured on consecutive versions of real objects against zstd 1.5.7's own --patch-from, same machine, every rebuild byte-exact:
| Old → new | zstd -3 --patch-from | zstd -19 --patch-from | Glyd --max --base | Glyd --ultra --base |
|---|---|---|---|---|
| Wikipedia page-table dumps, a month apart (108 MB) | 3.84 MB · 409 MB/s | 1.30 MB · 2 MB/s | 1.79 MB · 720 MB/s | 1.23 MB · 6 MB/s |
| Ubuntu 24.04 cloud root filesystem, builds 16 days apart (1.1 GB) | 8.82 MB · 654 MB/s | 5.61 MB · 39 MB/s | 5.33 MB · 1,590 MB/s | 4.60 MB · 11 MB/s |
| Linux 6.10 → 6.10.1 source tar (1.5 GB) | 3.26 MB · 560 MB/s | 2.58 MB · 30 MB/s | 3.03 MB · 1,700 MB/s | 2.04 MB · 3 MB/s |
Compressed alone those versions are 33, 287 and 200 MB: a version costs 1–5% of what it did. Against zstd's fast patch Glyd stores 1.1–2.1× less at 1.8–3× the speed; --ultra --base stores 5–21% less than zstd -19's patch on every pair, at the plain --ultra speed. Content is found wherever it moved in the old version (a coarse map of the base picks each unit's region). Chunk-level dedup, the backup approach, gains only 1–4× on the same pairs.
| 15 Linux 6.10 point releases (1.5 GB each) | Glyd --max --base | zstd -3 --patch-from | Stored one by one |
|---|---|---|---|
| each release against the one before it | 228 MB | 260 MB | 3,000 MB (Glyd --max) · 3,236 MB (zstd -3) |
| each release against 6.10 (any version is two reads) | 246 MB | 265 MB |
A step costs 1.8 MB, 0.12% of the tree; the delta against a base 14 releases old is 3.6 MB, so a series needs no rebasing. A terabyte of such trees in S3 Standard: $37 a year compressed one by one with --max ($40 with zstd -3), $2.8–3.0 with base mode ($3.2–3.3 with zstd's patch).
Model weights and checkpoints
A safetensors file is opened tensor by tensor; against a base (--base, and the store, which finds a checkpoint's predecessor by its tensors), a checkpoint is stored for what changed. The Ryzen box, all cores, every file decoded and compared (report):
| File | zstd -3 | zstd -19 | Glyd -9 |
|---|---|---|---|
| Pythia-410M, step 143000 (fp32, 1,621 MB) | 959 MB · 0.8 s | 809 MB · 89 s | 701 MB · 3.3 s |
| Qwen2.5-0.5B (bf16, 988 MB) | 769 MB · 0.3 s | 750 MB · 29 s | 663 MB · 2.5 s |
| A checkpoint against another | zstd -19 --patch-from | Glyd --max --base | |
| Pythia step 143000 against step 142000 | 806 MB · 148 s | 501 MB · 4.7 s | |
| Qwen2.5-0.5B-Instruct against Qwen2.5-0.5B | 748 MB · 57 s | 558 MB · 3.4 s |
The smallest size a coder that sees each tensor on its own can reach is 682–684 MB for a Pythia checkpoint and 650 MB for Qwen2.5: Glyd is within 2–3% of it.
Training checkpoints with the optimizer's state (torch.save; Qwen2.5-0.5B with AdamW, 5.93 GB each) go in at 83% of their size alone and 77% against the checkpoint before; zstd -19 stores 92%, and its --patch-from stops at 2 GB.
| Qwen2.5-7B-Instruct on an RTX 4080 SUPER (16 GB) | GPU memory | 1 sequence | 32 at once | Prompt of 128 tokens | Of 4096 |
|---|---|---|---|---|---|
| bf16 | 15.25 GB | 43.4 tokens/s | 1,154 tokens/s | 29 ms | 645 ms |
Glyd, mma | 10.61 GB | 55.7 | 1,519 | 29 ms | 702 ms |
The weights stay compressed in GPU memory and are rebuilt bit for bit as the model runs (gpu/): about a third of a bf16 model's memory back, and generation faster than bf16; perplexity as bf16's (17.0052 against 17.0015 on Wikipedia text).
Record mode: the data's own structure, not just its bytes
Byte-level matching (zstd, LZ4, Glyd's own levels) is a plateau: on real data zstd -19, xz and Glyd --ultra land within a few percent of each other. The redundancy of a log, a dump or a telemetry export sits in the same field of every record. glyd -r detects delimited lines, SQL dumps and JSON lines, turns them into one typed stream per field (integer, decimal and date-time deltas, dictionaries with recency ranks, text), compresses those, and rebuilds the bytes exactly. Logs whose lines vary in shape take a template: the line's text with a hole where every token holding a digit was, and the tokens as typed columns keyed by template and slot. Anything else is left as it is.
| Telemetry, 128 MB slices (ratio, higher is smaller) | LZ4 | zstd -3 | zstd -19 | Glyd --max -r | Glyd --ultra -r |
|---|---|---|---|---|---|
| Alibaba cluster machine usage, CSV | 2.8 | 4.5 | 6.9 | 12.6 | 13.7 |
| the same rows as JSON lines | 8.3 | 15.7 | 28.7 | 54.6 | 58.9 |
| NOAA daily weather, CSV | 3.8 | 7.0 | 12.0 | 19.6 | 23.1 |
| the same rows as JSON lines | 10.5 | 18.7 | 31.6 | 47.0 | 58.0 |
| NYC taxi trips, CSV export | 3.3 | 5.6 | 8.4 | 8.9 | 9.4 |
2.5–3.5× fewer bytes than the fast tier (LZ4, zstd -3), 1.5–2× fewer than zstd -19, written at 340–690 MB/s on 10 cores.
| Application and system logs, 128 MB slices (loghub 2.0) | LZ4 | zstd -3 | zstd -19 | Glyd --max -r | Glyd --ultra -r |
|---|---|---|---|---|---|
| HDFS (Hadoop file system) | 5.6 | 10.4 | 16.0 | 22.0 | 27.5 |
| Spark (application logs) | 7.9 | 14.5 | 25.2 | 47.0 | 53.6 |
| BGL (supercomputer RAS log) | 6.1 | 11.0 | 22.4 | 15.9 | 28.9 |
| Android (system log) | 5.0 | 12.9 | 23.0 | 17.9 | 25.4 |
1.4–3.3× fewer bytes than zstd -3 and 1.1–2.1× fewer than zstd -19 with --ultra -r; --max -r writes at 260–460 MB/s and reads back at 1,200–1,400 MB/s. Sources and the script that fetches them: scripts/download_ext_corpus.sh.
The cold level: below the floor every LZ codec shares
On real data zstd -19, xz -9 and Glyd --ultra land within 5% of each other: that is the floor of byte matching. Context mixing — every bit predicted from many contexts at once and coded at the mixed probability, no parse — goes below it at 30–100× the CPU. glyd --cold is that level: eleven predictors with paq-style bit histories, two mixers, two SSE stages, 32 MB units coded in parallel at 1.2–1.3 MB/s per core each way. Measured on 64 MB slices against the strongest tools, every decode byte-checked:
| Data | zstd -19 | xz -9 | Glyd --ultra (-r) | zpaq -m5 | Glyd --cold (-r) | vs zstd -19 |
|---|---|---|---|---|---|---|
| GitHub Archive JSON events | 14.6 | 14.8 | 15.9 | 22.8 · 0.4 MB/s | 22.5 · 1.3 MB/s per core | 1.54× smaller |
| NASA access log | 15.7 | 15.4 | 26.4 | 31.7 · 0.3 MB/s | 31.1 · 1.2 MB/s per core | 1.98× smaller |
| enwiki page_props SQL dump | 6.2 | 6.4 | 8.6 | 11.1 · 0.4 MB/s | 11.6 · 1.2 MB/s per core | 1.87× smaller |
| webster (text, 41 MB) | 4.8 | 4.9 | 4.8 | 7.3 · 0.35 MB/s | 7.1 · 1.2 MB/s per core | 1.47× smaller |
| HDFS log, 128 MB (-r) | 16.0 | 27.5 | 33.7 | 2.1× smaller | ||
| Spark log, 128 MB (-r) | 25.2 | 53.6 | 65.2 | 2.6× smaller |
The zpaq -m5 class within 3% either way, at 3–4× its speed per core; reads cost what writes cost. A terabyte is about 210 core-hours each way ($8 on Graviton3); 1.5–2× fewer bytes than zstd -19 saves $3–5 a year per raw terabyte in S3 Standard-IA, so the level pays for data kept two years or more and read a few times at most, or moved more than it is read.
The 8.7 GB corpus on AWS (Graviton3 and Sapphire Rapids, 8 threads)
| Data | Glyd --max | Glyd --max -r | zstd -3 | Glyd --ultra | Glyd --ultra -r | zstd -19 |
|---|---|---|---|---|---|---|
| Whole corpus | 3.94 | 4.71 | 3.85 | 4.66 | 5.22 | 4.66 |
| JSON events (2.6 GB) | 13.7 | 13.7 | 10.5 | 16.7 | 16.7 | 15.1 |
| Access and pageview logs (1.3 GB) | 5.09 | 6.16 | 4.89 | 6.63 | 7.74 | 6.83 |
| SQL dumps (3.9 GB) | 4.88 | 8.13 | 4.97 | 6.99 | 10.1 | 7.10 |
| Parquet (1.0 GB, already compressed) | 1.01 | 1.01 | 1.01 | 1.02 | 1.02 | 1.02 |
| Compress MB/s, Graviton3 | 1,512 | 643 | 1,969 | 14.5 | 20 | 13 |
| Decompress MB/s, Graviton3 | 10,413 | 4,040 | 1,422 | 9,620 | 3,927 | 1,326 |
--max -r stores 18% less than zstd -3 and 1% less than zstd -19 while writing 50× faster than zstd -19; Glyd's output decodes in parallel, so with 8 cores it reads 3–7× faster than a zstd frame. JSON events are not record-shaped (their bytes are hashes, ids and free text); the 128 MB long-distance matcher (--max --long, --ultra, the store) is what wins there.
A terabyte in S3 for a year
Compressed once by the CLI, read back once a month in the region, S3 Standard at list price, instance CPU seconds at the on-demand price (Graviton3):
| Codec | Stored | Year 1, 1 read/month | 10 reads/month | 100 reads/month |
|---|---|---|---|---|
| Raw | 1.000 TB | $276 | $276 | $278 |
| Glyd --max -r | 0.212 TB | $61.7 | $82.4 | $289 |
| Glyd --max | 0.254 TB | $71.6 | $81.3 | $178 |
| zstd -3 | 0.260 TB | $73.3 | $84.1 | $191 |
| Glyd --ultra -r | 0.191 TB | $82.8 | $103 | $311 |
| Glyd --ultra | 0.214 TB | $99.7 | $110 | $210 |
| zstd -19 | 0.216 TB | $102 | $114 | $232 |
Where zstd still wins, plainly: a hundred CPU-billed reads a month of record-mode data (its reads spend 2× zstd's CPU rebuilding the columns), single-core compression speed on match-dense data (--max at 0.89× zstd -3 on logs and table dumps on Graviton3; 1.01–1.05× on the rest, v0.14.5), small objects with dictionaries (1–3% larger, 1.4–2× slower per object), and data made of hashes.
How the numbers are made
- Correctness first: 106 tests, a million-mutation fuzz, files from every earlier release decoded unchanged, and
scripts/verify_roundtrip.shthrough the CLI on both machines with corrupted copies rejected or decoded exactly; base-mode streams refuse any other base. - Fair comparison: zstd -3, zstd -19 and LZ4 on the same machines and thread counts, each on its fastest API, buffers kept; compressed bytes include every codec's own framing.
- A real workflow: compress, upload to S3, download, decompress, sha256 every file, at list prices.
- The floors are stated: JSON API events are 20% hashes, 10% random ids and 30% free text once compressed; Parquet is already zstd inside. Neither moves. Where zstd wins — write speed, small objects per second, reads at very high read rates — is written down too.
Full report with every table · Design notes · Changelog
The codec is BSD-3-Clause or GPL-2.0 at your option (the same licenses as zstd); the store (glyd-store) is under the Business Source License 1.1. Contact: suryakoritala1324@gmail.com.