Glyd

Lossless AI compression: 33% less GPU memory, bit for bit. Rust, C ABI, CLI; v0.26.0, October 2026. Every number below is measured on public data and reproducible from the repository.

What would it save you? Quick start

Fastest or slowest? The table everyone uses

The 8.7 GB real-data corpus on AWS Graviton3, ratio · compress MB/s · decompress MB/s. One core, the way zstd's README reports it, then eight cores (Glyd's output decodes in parallel; a zstd or LZ4 frame decodes on one thread).

One coreRatioCompress MB/sDecompress MB/s
LZ42.725041,394
Glyd default2.823323,309
zstd -33.863131,425
Glyd --max3.982331,571
Eight coresRatioCompress MB/sDecompress MB/s
Glyd default2.822,14322,707
zstd -3 -T83.851,9691,422
Glyd --max3.941,51210,413
Glyd --max -r4.716434,040
LZ42.725031,392
zstd -19 -T84.66131,326
Glyd --ultra4.6614.59,620
Glyd --ultra -r5.22203,927

Reads: Glyd is the fastest, 3–7× zstd on a server. Writes: Glyd is slower, 0.77× zstd -3 (record mode 0.33×); LZ4 is the write-speed king on one core. Size: Glyd is the smallest wherever the data has structure, and ties elsewhere.

Against zstd, xz and brotli on 24 kinds of data

Every codec's own CLI on one thread, every decode compared byte for byte with its input. Green: Glyd's file is smaller.

Compression benchmark: Glyd --max bytes against zstd -3 on logs, SQL dumps, JSON, gzip, zip, jar, Office documents, PDF, PNG, JPEG, text, executables and Parquet
Datavs zstd -3 (--max)vs zstd -19, xz, brotli -11 (--ultra)
Records: logs, dumps, JSON (-r)12–61% fewer bytes3–38% fewer (JSON: 9% more)
Containers: gzip, zip, Office, PDF, PNG, JPEGup to 60% fewer14–63% fewer

Glyd wins where the data has structure: each field of a record becomes a column, and the deflate, snappy, zstd or JPEG inside a container is opened and re-created bit for bit, Parquet pages included. On plain text and executables it is zstd-class: within 2% at the fast tier, 1–15% larger than xz and brotli -11 at their strongest. Every codec's speed, the strong-tier and ratio-against-speed charts, and lz4, bzip2, zpaq and JPEG XL.

The store: compression across objects, not just inside them

Inside one object every codec sits on the same floor. The redundancy of object storage is between objects — builds, snapshots, dumps and releases that are near-copies of earlier ones. glyd-store bucket/ --put ... keeps each object as a delta against the stored object it most resembles when that pays; a read is at most five decodes. Measured on a 39 GB bucket, each object arriving in order, every one read back and compared:

FamilyRawzstd -3, each object aloneGlyd storeGain
Linux 6.10 releases (15)22.5 GB3,237 MB232 MB13.9× smaller
Ubuntu 24.04 cloud images (6 builds)6.6 GB1,880 MB361 MB5.2×
Wikipedia dumps (2 months, 3 tables)0.7 GB140 MB71 MB2.0×
GitHub events (12 hours)9.4 GB875 MB670 MB1.3× (nothing is a version of anything)
The bucket39.2 GB6,132 MB (6.4×)1,334 MB (29.4×)4.6× smaller

Put runs at 620 MB/s end to end on ten cores; delete, compact and verify are there. A petabyte of such data in S3 Standard: $43K a year with zstd -3, $9.4K with the store. Chunk-level dedup, the backup approach, gains 1–4× on the same pairs.

Base mode: a version stored for the change, not the whole

Most stored bytes are versions — nightly dumps, snapshots, images, source trees. glyd --base old new parses every 32 MB of the new version with the old one's matching region as history and writes a stream that decodes with the same base. Measured on consecutive versions of real objects against zstd 1.5.7's own --patch-from, same machine, every rebuild byte-exact:

Old → newzstd -3 --patch-fromzstd -19 --patch-fromGlyd --max --baseGlyd --ultra --base
Wikipedia page-table dumps, a month apart (108 MB)3.84 MB · 409 MB/s1.30 MB · 2 MB/s1.79 MB · 720 MB/s1.23 MB · 6 MB/s
Ubuntu 24.04 cloud root filesystem, builds 16 days apart (1.1 GB)8.82 MB · 654 MB/s5.61 MB · 39 MB/s5.33 MB · 1,590 MB/s4.60 MB · 11 MB/s
Linux 6.10 → 6.10.1 source tar (1.5 GB)3.26 MB · 560 MB/s2.58 MB · 30 MB/s3.03 MB · 1,700 MB/s2.04 MB · 3 MB/s

Compressed alone those versions are 33, 287 and 200 MB: a version costs 1–5% of what it did. Against zstd's fast patch Glyd stores 1.1–2.1× less at 1.8–3× the speed; --ultra --base stores 5–21% less than zstd -19's patch on every pair, at the plain --ultra speed. Content is found wherever it moved in the old version (a coarse map of the base picks each unit's region). Chunk-level dedup, the backup approach, gains only 1–4× on the same pairs.

15 Linux 6.10 point releases (1.5 GB each)Glyd --max --basezstd -3 --patch-fromStored one by one
each release against the one before it228 MB260 MB3,000 MB (Glyd --max) · 3,236 MB (zstd -3)
each release against 6.10 (any version is two reads)246 MB265 MB

A step costs 1.8 MB, 0.12% of the tree; the delta against a base 14 releases old is 3.6 MB, so a series needs no rebasing. A terabyte of such trees in S3 Standard: $37 a year compressed one by one with --max ($40 with zstd -3), $2.8–3.0 with base mode ($3.2–3.3 with zstd's patch).

Model weights and checkpoints

A safetensors file is opened tensor by tensor; against a base (--base, and the store, which finds a checkpoint's predecessor by its tensors), a checkpoint is stored for what changed. The Ryzen box, all cores, every file decoded and compared (report):

Filezstd -3zstd -19Glyd -9
Pythia-410M, step 143000 (fp32, 1,621 MB)959 MB · 0.8 s809 MB · 89 s701 MB · 3.3 s
Qwen2.5-0.5B (bf16, 988 MB)769 MB · 0.3 s750 MB · 29 s663 MB · 2.5 s
A checkpoint against anotherzstd -19 --patch-fromGlyd --max --base
Pythia step 143000 against step 142000806 MB · 148 s501 MB · 4.7 s
Qwen2.5-0.5B-Instruct against Qwen2.5-0.5B748 MB · 57 s558 MB · 3.4 s

The smallest size a coder that sees each tensor on its own can reach is 682–684 MB for a Pythia checkpoint and 650 MB for Qwen2.5: Glyd is within 2–3% of it.

Training checkpoints with the optimizer's state (torch.save; Qwen2.5-0.5B with AdamW, 5.93 GB each) go in at 83% of their size alone and 77% against the checkpoint before; zstd -19 stores 92%, and its --patch-from stops at 2 GB.

Qwen2.5-7B-Instruct on an RTX 4080 SUPER (16 GB)GPU memory1 sequence32 at oncePrompt of 128 tokensOf 4096
bf1615.25 GB43.4 tokens/s1,154 tokens/s29 ms645 ms
Glyd, mma10.61 GB55.71,51929 ms702 ms

The weights stay compressed in GPU memory and are rebuilt bit for bit as the model runs (gpu/): about a third of a bf16 model's memory back, and generation faster than bf16; perplexity as bf16's (17.0052 against 17.0015 on Wikipedia text).

Record mode: the data's own structure, not just its bytes

Byte-level matching (zstd, LZ4, Glyd's own levels) is a plateau: on real data zstd -19, xz and Glyd --ultra land within a few percent of each other. The redundancy of a log, a dump or a telemetry export sits in the same field of every record. glyd -r detects delimited lines, SQL dumps and JSON lines, turns them into one typed stream per field (integer, decimal and date-time deltas, dictionaries with recency ranks, text), compresses those, and rebuilds the bytes exactly. Logs whose lines vary in shape take a template: the line's text with a hole where every token holding a digit was, and the tokens as typed columns keyed by template and slot. Anything else is left as it is.

Telemetry, 128 MB slices (ratio, higher is smaller)LZ4zstd -3zstd -19Glyd --max -rGlyd --ultra -r
Alibaba cluster machine usage, CSV2.84.56.912.613.7
the same rows as JSON lines8.315.728.754.658.9
NOAA daily weather, CSV3.87.012.019.623.1
the same rows as JSON lines10.518.731.647.058.0
NYC taxi trips, CSV export3.35.68.48.99.4

2.5–3.5× fewer bytes than the fast tier (LZ4, zstd -3), 1.5–2× fewer than zstd -19, written at 340–690 MB/s on 10 cores.

Application and system logs, 128 MB slices (loghub 2.0)LZ4zstd -3zstd -19Glyd --max -rGlyd --ultra -r
HDFS (Hadoop file system)5.610.416.022.027.5
Spark (application logs)7.914.525.247.053.6
BGL (supercomputer RAS log)6.111.022.415.928.9
Android (system log)5.012.923.017.925.4

1.4–3.3× fewer bytes than zstd -3 and 1.1–2.1× fewer than zstd -19 with --ultra -r; --max -r writes at 260–460 MB/s and reads back at 1,200–1,400 MB/s. Sources and the script that fetches them: scripts/download_ext_corpus.sh.

The cold level: below the floor every LZ codec shares

On real data zstd -19, xz -9 and Glyd --ultra land within 5% of each other: that is the floor of byte matching. Context mixing — every bit predicted from many contexts at once and coded at the mixed probability, no parse — goes below it at 30–100× the CPU. glyd --cold is that level: eleven predictors with paq-style bit histories, two mixers, two SSE stages, 32 MB units coded in parallel at 1.2–1.3 MB/s per core each way. Measured on 64 MB slices against the strongest tools, every decode byte-checked:

Datazstd -19xz -9Glyd --ultra (-r)zpaq -m5Glyd --cold (-r)vs zstd -19
GitHub Archive JSON events14.614.815.922.8 · 0.4 MB/s22.5 · 1.3 MB/s per core1.54× smaller
NASA access log15.715.426.431.7 · 0.3 MB/s31.1 · 1.2 MB/s per core1.98× smaller
enwiki page_props SQL dump6.26.48.611.1 · 0.4 MB/s11.6 · 1.2 MB/s per core1.87× smaller
webster (text, 41 MB)4.84.94.87.3 · 0.35 MB/s7.1 · 1.2 MB/s per core1.47× smaller
HDFS log, 128 MB (-r)16.027.533.72.1× smaller
Spark log, 128 MB (-r)25.253.665.22.6× smaller

The zpaq -m5 class within 3% either way, at 3–4× its speed per core; reads cost what writes cost. A terabyte is about 210 core-hours each way ($8 on Graviton3); 1.5–2× fewer bytes than zstd -19 saves $3–5 a year per raw terabyte in S3 Standard-IA, so the level pays for data kept two years or more and read a few times at most, or moved more than it is read.

The 8.7 GB corpus on AWS (Graviton3 and Sapphire Rapids, 8 threads)

DataGlyd --maxGlyd --max -rzstd -3Glyd --ultraGlyd --ultra -rzstd -19
Whole corpus3.944.713.854.665.224.66
JSON events (2.6 GB)13.713.710.516.716.715.1
Access and pageview logs (1.3 GB)5.096.164.896.637.746.83
SQL dumps (3.9 GB)4.888.134.976.9910.17.10
Parquet (1.0 GB, already compressed)1.011.011.011.021.021.02
Compress MB/s, Graviton31,5126431,96914.52013
Decompress MB/s, Graviton310,4134,0401,4229,6203,9271,326

--max -r stores 18% less than zstd -3 and 1% less than zstd -19 while writing 50× faster than zstd -19; Glyd's output decodes in parallel, so with 8 cores it reads 3–7× faster than a zstd frame. JSON events are not record-shaped (their bytes are hashes, ids and free text); the 128 MB long-distance matcher (--max --long, --ultra, the store) is what wins there.

A terabyte in S3 for a year

Compressed once by the CLI, read back once a month in the region, S3 Standard at list price, instance CPU seconds at the on-demand price (Graviton3):

CodecStoredYear 1, 1 read/month10 reads/month100 reads/month
Raw1.000 TB$276$276$278
Glyd --max -r0.212 TB$61.7$82.4$289
Glyd --max0.254 TB$71.6$81.3$178
zstd -30.260 TB$73.3$84.1$191
Glyd --ultra -r0.191 TB$82.8$103$311
Glyd --ultra0.214 TB$99.7$110$210
zstd -190.216 TB$102$114$232

Where zstd still wins, plainly: a hundred CPU-billed reads a month of record-mode data (its reads spend 2× zstd's CPU rebuilding the columns), single-core compression speed on match-dense data (--max at 0.89× zstd -3 on logs and table dumps on Graviton3; 1.01–1.05× on the rest, v0.14.5), small objects with dictionaries (1–3% larger, 1.4–2× slower per object), and data made of hashes.

How the numbers are made

Full report with every table · Design notes · Changelog

The codec is BSD-3-Clause or GPL-2.0 at your option (the same licenses as zstd); the store (glyd-store) is under the Business Source License 1.1. Contact: suryakoritala1324@gmail.com.