Structure-first compression, measured: the dependency axis closed, the niche narrowed
Two benchmarks on a disclosed corpus continue the entry of 2026-07-14, negatives first: XOR dependency coding refuted in 1D and 2D, a historical codec version that could not decode its own output, and a positive that did not carry over to real scans — leaving born-digital rasters with bit-exact repetition as the only niche. With corrections to the earlier entry.
Koch Laboratory — structure-first compression, measured: the dependency axis closed, the niche narrowed
The entry of 2026-07-14 ended by announcing a defined benchmark with a disclosed corpus and reference baselines, because the compression ratio of neither direction had been measured. This entry reports two measurement sessions: Benchmark 1 (2026-07-15; 10 files, 154 measurements) and Benchmark 2 (2026-07-17; realistic images, 51 measurements). The falsifiable criteria were set before each session, and every decode was checked against the original with SHA-256. Statements of the entry of 2026-07-14 that the measurements overturned are collected under “Corrections to the entry of 2026-07-14”.
The balance in three sentences. The dependency-coding axis is refuted in 1D and in 2D and is closed. The historical version of the PGA codec could not restore the data it had itself encoded. After repair, the codec beats the classical baselines only on a born-digital raster whose repetitions are rigid and bit-exact — on real scans and rendered pages it loses.
Method. Problem → state of the art → criterion set before measurement → method → result with boundary conditions → next step. Compression ratio = output size / input size: lower is better, above 1 means expansion. The whole container counts — headers, metadata, boundary bits, checksum — not the payload alone. Criterion labels (F1–F7) follow the private measurement protocols, whose hashes are given in section 9.
Publication note. For directions still open we publish the problem, the state of the art, the criterion and the result — not the construction. Where “Construction: withheld.” appears, the technical detail is retained as filing material. The dependency-coding axis is closed as falsified, so its rule is disclosed here.
1. Direction 1 — dependency coding in 1D: rule disclosed, hypothesis refuted
Research question. Does a reversible representation of a stream as layers of relations between neighbouring bits yield, after a classical final coder, an output smaller than the same coder without the transform, and smaller than zlib-9 and zstd-19?
Rule (disclosed, since the axis is closed). The first layer is the XOR of every pair of neighbouring bits, b[i] ⊕ b[i+1]; each further layer is built from the previous one by the same rule. Inverting L layers requires L boundary bits.
Analytic note. For the measured depths L = 1, 2, 4, 8 the iteration yields exactly b[i] ⊕ b[i+L] (binomial coefficients modulo 2): layer L is an XOR difference at distance L, and eight layers are the XOR difference of neighbouring bytes. At the measured depths the transform therefore lies within classical delta coding.
State of the art (published). Delta and predictive coding (including XOR delta), RLE, DEFLATE (zlib, gzip), zstd.
Criteria (set before measurement). F1: the SHA-256 of the reconstruction equals that of the original on every decode. F2: layer L plus the final coder yields a smaller output than the same container without the transform (L = 0), for L ∈ {1, 2, 4, 8}. F3: the output is smaller than zlib-9 and than zstd-19.
Method. A corpus of 10 files (section 8); L ∈ {0, 1, 2, 4, 8}; final coder RLE or zlib-9; 100 measurements in total. The measurements used a canonical re-implementation that stores all boundary bits — the earlier reference implementation stored only the first of them and so could not decode more than one layer.
Result. Compression ratio with zlib-9 as the final coder (whole container), and the baselines on the raw data:
| file | L = 0 (control) | L = 1 | L = 2 | L = 4 | L = 8 | zlib-9 | zstd-19 |
|---|---|---|---|---|---|---|---|
| technical text | 0.2929 | 0.2972 | 0.3083 | 0.3261 | 0.3541 | 0.2903 | 0.2778 |
| synthetic form raster | 0.1411 | 0.1439 | 0.1581 | 0.1714 | 0.1720 | 0.1411 | 0.1069 |
| telemetry (CSV) | 0.1623 | 0.1624 | 0.1648 | 0.1694 | 0.2050 | 0.1622 | 0.0740 |
| sparse file | 0.0028 | 0.0029 | 0.0030 | 0.0033 | 0.0034 | 0.0026 | 0.0014 |
| long bit runs | 0.0067 | 0.0065 | 0.0065 | 0.0065 | 0.0065 | 0.0065 | 0.0094 |
| random data | 1.0005 | 1.0005 | 1.0005 | 1.0005 | 1.0005 | 1.0003 | 1.0001 |
| executable slice | 0.8641 | 0.8701 | 0.8762 | 0.8838 | 0.8937 | 0.8640 | 0.8509 |
| all zeros | 0.0012 | 0.0012 | 0.0012 | 0.0012 | 0.0012 | 0.0011 | 0.0001 |
| alternating bits | 0.0012 | 0.0012 | 0.0012 | 0.0012 | 0.0012 | 0.0011 | 0.0001 |
| periodic file | 0.0034 | 0.0035 | 0.0036 | 0.0034 | 0.0028 | 0.0033 | 0.0004 |
With zlib-9 the result worsens monotonically with the number of layers on five files: the technical text, the synthetic form raster, the telemetry, the sparse file and the executable slice. The exceptions are all synthetic. On the long-run file one layer saves 48 B against the same container without the transform (1,707 vs 1,755 B), which merely equals zlib-9 alone (1,707 B). On the periodic file eight layers save 175 B (725 vs 900 B), while zstd-19 codes the same file to 97 B. On random data and in the two degenerate cases (all zeros, alternating bits) the output is constant to within 1 B. With RLE, the transform turned expansion into compression on one file only — the alternating-bit file (16.00 → 0.0629 at one layer), which zlib-9 codes to 0.0011 with no transform at all. Where layers reduced the RLE output on other files, it stayed larger than the input (e.g. form raster 1.51, telemetry 5.74, technical text 6.33). Criteria outcome. F1: met, 100/100. F2: with zlib-9 met only on two synthetic files (long runs, periodic file); with RLE only where the output is still larger than the input, and on the alternating bits. F3: unmet on 10 of 10 files — no output with the transform was smaller than both zlib-9 and zstd-19. Hypothesis refuted. Mechanism (measured). On the text and the telemetry the first layers push the density of ones towards 0.5: technical text 0.3932 → 0.4607 (L = 1) → 0.4805 (L = 2), telemetry 0.4261 → 0.5119 (L = 1). The transform thus blurs the structure the final coder sees instead of exposing it. Where the density falls, the output still does not improve: on the synthetic form raster one layer lowers the density from 0.1918 to 0.1365, yet the zlib-9 ratio rises from 0.1411 to 0.1439. The same pattern recurs in 2D (section 6). Next step. None — the axis is closed (section 7). Status: falsified by measurement (F3 unmet on 10/10 files; F1 met, 100/100); axis closed.
2. Direction 1 — the full layer pyramid: reversible, but quadratic
Research question. Is the historical implementation, which builds the full pyramid of layers down to a single bit, feasible? Criteria. As in section 1 (F1–F3). Method. The historical implementation, unmodified, run on prefixes of the technical text from 128 to 1,024 B. Result.
| input | output | output / input | encoding time |
|---|---|---|---|
| 128 B | 515,415 B | ×4,027 | 0.086 s |
| 256 B | 2,055,749 B | ×8,030 | 0.257 s |
| 512 B | 8,285,223 B | ×16,182 | 0.969 s |
| 1,024 B | 33,310,203 B | ×32,529 | 3.844 s |
The pyramid for n input bits holds n(n−1)/2 dependency bits; the measured output tracks n(n−1)/2 bytes to within 2% (98.1–99.3%), i.e. about one byte per dependency bit. Size and time grow quadratically. Extrapolated to the 14.2 MB executable from the entry of 2026-07-14: about 6.5 PB of output and about 23 years of computation on the measurement machine. F1 met (4/4); F2 and F3 refuted — the output is roughly 4,000 to 32,500 times the size of the input. An observation on the construction: the decoder of this implementation reads the first layer only (together with one boundary bit it suffices for reconstruction), so storing the layers from the second upwards is redundant by design. Next step. None — implementation withdrawn. Status: reversible (4/4), infeasible beyond toy sizes; withdrawn.
3. Direction 2 — the PGA codec in its historical version: reconstruction failed on 10 of 10 files
Research question. Does the codec version on which the 56% expansion reported on 2026-07-14 was measured restore data bit-exactly? Criteria (set before measurement). F1 for the PGA codec, and F4: the size of the whole container (all sections and metadata) is smaller than the input on the structured corpus, with F1 met. Method. The historical version, unmodified; encoding and decoding of each of the 10 corpus files; SHA-256 of the output against the original. Result: reconstruction criterion refuted. Encoding succeeded on 10 of 10 files, decoding on 0 of 10: on 4 files it aborted with an error, on 6 it produced data whose SHA-256 did not match the original. Code analysis found two format defects: the residual layer did not carry the data it was meant to restore, and the map section was not unambiguously parseable. The consequence reaches beyond the benchmark: the historical 22.2 MB artefact from the entry of 2026-07-14 is not fully reproducible either. Its size was measured correctly (14.2 → 22.2 MB), but it was not a valid lossless encoding. F4 cannot be assessed while F1 is unmet. Next step. Repair of both defects in a measurement variant (section 4). Status: negative result — the historical version fails the reconstruction criterion (0/10); the correction of the entry of 2026-07-14 is given below.
4. Direction 2 — after repair: a conditional positive that did not generalise to realistic data
Research question. After the format repair, does the codec restore data bit-exactly and achieve a positive size balance — first on the Benchmark 1 corpus, then on realistic images? State of the art (published). For bilevel images: CCITT G4 (ITU-T T.6, two-dimensional run coding relative to a reference line) and JBIG2 (ITU-T T.88), which itself codes text through symbol dictionaries with pattern matching. These, not gzip and zstd, are the real competition for this class of data. G4 was measured in Benchmark 2; JBIG2 was not — no encoder was available in the measurement environment. Criteria (set before measurement). F4 (Benchmark 1): container smaller than the input on the structured corpus, with F1 met. F5 (Benchmark 2): compression ratio smaller than zstd-19 and than CCITT G4 on the realistic corpus. Method. A measurement variant with both defects repaired, in two forms: faithful to the historical format, and with a map section stripped of redundancy (construction withheld). Benchmark 2 uses the second form on seven bilevel images (section 8), with the true raster geometry. Result — Benchmark 1. F1 restored: 10/10 in both forms. F4 as worded in the protocol proved weak: the format-faithful form has a container smaller than the input on 4 of 5 structured files (on the technical text 1.098), the form with the redundancy-free map on 5 of 5. The comparison with the baselines decides. The format-faithful form was larger than zstd-19 on every one of the 10 files; on the synthetic form raster the map took 97% of its container. The form with the redundancy-free map produced the program’s first measured positive — on one file: the synthetic form raster at 0.0495 against 0.1069 (zstd-19) and 0.1411 (zlib-9), i.e. 2.2× better than the best baseline. On the other 9 files the baselines won (e.g. telemetry 0.2981 against 0.0740, technical text 0.7949 against 0.2778). On high-entropy data the codec expands: random data 1.176 (format-faithful form: 1.547), executable slice 1.115 (1.519). Status after Benchmark 1: a conditional positive — synthetic raster, glyphs aligned to the block grid, no noise, no comparison with G4 or JBIG2. Result — Benchmark 2: F5 refuted.
| image | PGA | zlib-9 | zstd-19 | CCITT G4 | best |
|---|---|---|---|---|---|
fax2d — CCITT fax, 1728×1082 |
0.1888 | 0.1374 | 0.1241 | 0.1209 | G4 |
g3test — CCITT fax, 1728×1103 |
0.2629 | 0.1806 | 0.1624 | 0.1756 | zstd-19 |
jim___ah — dense test image, 664×813 |
0.2443 | 0.2083 | 0.1910 | 0.6301 | zstd-19 |
| text page, sans-serif font | 0.0489 | 0.0440 | 0.0364 | 0.0376 | zstd-19 |
| text page, serif font | 0.0348 | 0.0376 | 0.0302 | 0.0346 | zstd-19 |
| form page | 0.0351 | 0.0263 | 0.0133 | 0.0372 | zstd-19 |
| synthetic form raster (control from Benchmark 1) | 0.0492 | 0.1411 | 0.1069 | 0.4030 | PGA |
The codec loses to zstd-19 on 6 of 6 realistic images and to CCITT G4 on 4 of 6. The only image on which any of the program’s own methods comes out best remains the synthetic raster aligned to the block grid. F1: 7/7. A methodological note: in Benchmark 1 the codec rasterised its input without knowing its true geometry, and the raster width matched the true width of the synthetic form raster by coincidence; Benchmark 2 uses the true geometry in every measurement. Next step. Establish what the only positive depends on (section 5). Status: reversibility restored (10/10, 7/7); generalisation refuted (F5); the positive is confined to the synthetic raster.
5. Direction 2 — the mechanism of the positive: rigid, bit-exact repetition
Research question. What does the only positive depend on — the alignment of patterns to the block grid, their bit-exactness, or both? Criteria. A sensitivity test announced in protocol 1, without a success threshold (what is measured is the degradation curve against G4 and zstd-19 on the same data); F7 (set before measurement): choosing the grid offset restores the compression ratio of the aligned raster. Method. The synthetic form raster shifted horizontally by less than one block width, and made noisy by flipping random bits with a probability of 0.1 to 2%; every result verified with SHA-256 (13/13). Result — shift. Depending on the size of the shift, a compression ratio of 0.0492 to 0.0880, with a dictionary of up to 46 entries instead of 16; G4 unchanged (0.403). Result — noise.
| flipped bits | PGA | dictionary entries | zstd-19 | CCITT G4 |
|---|---|---|---|---|
| 0 | 0.0492 | 16 | 0.1069 | 0.4030 |
| 0.1% | 0.0713 | 20 | 0.1327 | 0.4158 |
| 0.5% | 0.1290 | 78 | 0.1969 | 0.4652 |
| 1% | 0.1991 | 561 | 0.2555 | 0.5195 |
| 2% | 0.3851 | 2,656 | 0.3434 | 0.6180 |
The lead over zstd-19 shrinks from 2.2× on clean data to 1.9×, 1.5× and 1.3×, and reverses at 2% flipped bits (0.3851 against 0.3434). The codec’s ratio worsens by a factor of 7.8, G4’s by a factor of 1.5. Conclusion: the positive requires rigid, bit-exact repetition of patterns — the kind that born-digital rasters have and scans do not. Result — F7: confirmed technically. A raster deliberately shifted against the block grid: 0.1416; after choosing the grid offset: 0.0492, exactly the value of the aligned raster; cost: a few bits; F1 met. This holds for a shift of the whole raster; it was not measured on rendered pages or scans, where the positions of characters vary individually. Construction: withheld. Next step. The cross-document deduplication hypothesis (section 7). Status: boundary conditions measured; F7 confirmed technically.
6. Direction 3 — spatial dependency in 2D: hypothesis refuted, with a predictor ablation
Research question. Does the dependency rule, carried over from 1D to spatial neighbours in a raster, carry a signal that 1D lacked? State of the art (published). Context modelling of bilevel images in JBIG and JBIG2: a context template of 10–16 neighbouring pixels and an adaptive arithmetic coder. Here a minimal variant was measured on purpose, to decide whether the axis carries any signal at all. Criteria (set before measurement). F6(a): the majority predictor (majority of three neighbours: left, upper, upper-left) beats the trivial predictors (left neighbour alone, upper neighbour alone). F6(b): the output is smaller than the best baseline for the given image. Method. Residual = pixel ⊕ prediction; causal decoder; final coder zlib-9 or zstd-19; 7 images × 4 variants = 28 measurements; F1: 28/28. Result.
| image | left + zlib-9 | upper + zlib-9 | majority + zlib-9 | majority + zstd-19 | best baseline |
|---|---|---|---|---|---|
fax2d |
0.1359 | 0.1646 | 0.1691 | 0.1567 | 0.1209 (G4) |
g3test |
0.1800 | 0.2178 | 0.2355 | 0.2154 | 0.1624 (zstd-19) |
jim___ah |
0.2085 | 0.3058 | 0.2849 | 0.2642 | 0.1910 (zstd-19) |
| text page, sans-serif font | 0.0450 | 0.0456 | 0.0554 | 0.0469 | 0.0364 (zstd-19) |
| text page, serif font | 0.0387 | 0.0409 | 0.0465 | 0.0376 | 0.0302 (zstd-19) |
| form page | 0.0272 | 0.0264 | 0.0316 | 0.0165 | 0.0133 (zstd-19) |
| synthetic form raster | 0.1412 | 0.1610 | 0.2447 | 0.2189 | 0.1069 (zstd-19) |
F6(a) refuted: with the same coder (zlib-9) the majority predictor was worse than the left neighbour on 7 of 7 images, and worse than both trivial predictors on 6 of 7 (exception: on jim___ah it beat the upper neighbour). F6(b) refuted: no variant beat the best baseline on any image (0/7). The best variant beat G4 alone on 3 of 7 images (jim___ah, the form page, the synthetic raster) and zstd-19 on the raw data on none.
Mechanism. The residual is sparser than the original and still does not compress better. On jim___ah it has 0.165 ones per pixel against 0.557 in the original, and with zlib-9 yields practically the same output (0.2085 against 0.2083). With the same coder the left-neighbour residual was at most marginally smaller than the original (fax2d: 0.1359 against 0.1374). An interpretation consistent with both measurements: a coder of the LZ family gains from repetitions — identical glyphs and rows — not from bit density; the dependency transform changes the density (towards 0.5 in 1D, downwards in 2D) but adds no repetitions and breaks some of those present. The gain from sparsity does not make up for the lost matches.
Next step. None in this form: further work would mean context modelling with an arithmetic coder, i.e. convergence on the state of the art (JBIG/JBIG2), not a new axis.
Status: falsified (F6a, F6b); axis closed.
7. Program conclusion: the dependency axis closed, the niche narrowed
- The dependency-coding axis is closed as falsified, in 1D (sections 1–2) and in 2D (section 6), with final coders of the RLE and LZ classes. Recorded mechanism: the transform changes bit density but adds none of the repetitions the coder relies on.
- The PGA codec — the niche narrowed. The program moved from “a better general-purpose compressor” (refuted in the 2021–2025 iterations), through “structured corpora” (documents, scans, telemetry — refuted by measurement for scans, rendered pages, text and telemetry), to a narrow class of data: born-digital rasters with rigid, bit-exact repetition (programmatically generated forms, print spool output, documents filled from a template). There the codec was about 2.2× better than the best classical baseline — measured with full reconstruction — and about 2.6× after a further improvement of the container; the latter is an estimate from section sizes, not yet measured end to end. Construction: withheld.
- Boundary conditions of the niche, stated plainly. The lead was measured on a single synthetic raster; the comparison covered G4, zlib-9 and zstd-19, not JBIG2, which uses symbol dictionaries and is the closest state of the art; robustness to noise is weak (section 5). This is a scoping hypothesis, not a demonstrated lead on real data.
- Next hypothesis: cross-document deduplication. A shared dictionary for a series of pages from the same template (e.g. series of invoices or reports); expectation: the dictionary cost amortises over the series, and the compression ratio of the series falls well below that of a single page. JBIG2 provides symbol dictionaries shared across pages, so without a comparison with it the result of this hypothesis will not be conclusive. The criterion will be set before measurement, as in both sessions.
8. Corpus, baselines, reproducibility
Benchmark 1 (2026-07-15), 10 files. Five structured: technical text (project documentation, 18,309 B), a synthetic form raster at 1 bit/pixel (512 KiB; table lines and a small fixed set of glyphs placed on the block grid, no noise), synthetic telemetry in CSV format (418,929 B), a sparse file (256 KiB; zeros and a 16-byte record header every 4,096 B), a file of long runs of 0x00 and 0xFF bytes (256 KiB). Two high-entropy control files: pseudo-random data (256 KiB) and the first 1 MiB of the compiled Windows executable from the entry of 2026-07-14. Three edge cases of 256 KiB each: all zeros, alternating bits (byte 0xAA), a periodic sequence (bytes 0–63). The synthetic files are deterministic (fixed generator seed).
Benchmark 2 (2026-07-17), 7 bilevel images. Three real images from the public libtiff test-image archive pics-3.8.0 (download.osgeo.org/libtiff/pics-3.8.0.tar.gz): fax2d (1728×1082) and g3test (1728×1103) — classic CCITT fax test images — and jim___ah (664×813; dense: 56% black pixels, not a single blank row). Three pages rendered with TrueType system fonts at about 200 dpi (1656×2336): text in a sans-serif font, text in a serif font, and a form with ruled lines and labels in a monospaced font; line and character positions deliberately not aligned to the block grid. Control: the synthetic form raster from Benchmark 1. The rasters are stored as raw bits, 1 bit/pixel, row by row, black = 1.
Baselines. zlib-9 (DEFLATE, as in gzip), zstd-19, CCITT G4 (in Benchmark 2; measured as a complete TIFF file). JBIG2 — not measured: no encoder in the measurement environment.
Environment. Windows 11 (AMD64), Python 3.12.10, numpy 2.5.0, Pillow 10.4.0, zstandard 0.25.0.
Verification. Benchmark 1: 154 measurements, 144 decodes matching the SHA-256 of the original; the 10 mismatches are the historical version of the PGA codec (section 3). Benchmark 2: 51 measurements, including 48 decodes — 48/48 matching; the remaining 3 measurements concerned a construction detail of the container and their result is withheld.
Reproducibility. Anyone can repeat the baseline results on the libtiff images: the archive is public, and the geometry and raster format are given above. The results of the program’s own methods require the withheld construction; the hashes in section 9 bind them to the protocols.
9. Evidence: SHA-256 hashes and timestamps
The measurement protocols remain private: they contain the withheld construction. Their SHA-256 hashes, published here, make it possible to check later that a document produced in the future is the same one that was stamped in July 2026. The evidence lists contain the SHA-256 hashes of the raw result files and the corpus manifests, and the commit identifiers of the measurement code (the first list also the hash of the 2025 whitepaper).
| document | SHA-256 | OpenTimestamps stamp |
|---|---|---|
| measurement protocol no. 1 (Benchmark 1) | 0830cc94c042c05103e45a47de2bdae0db2411f683fcd309e6630c5700552375 |
2026-07-15, Bitcoin attestation |
| evidence list for Benchmark 1 | 8b85e99708febb32682a100f5b2336bd0ffc8cf0e0ef4c7671aa91576de6dc56 |
2026-07-15, Bitcoin attestation |
| measurement protocol no. 2 (Benchmark 2) | c478b509fda7419a60b0ad2149d28f4dec83aae5fb074ffbeab0cd3bb3aba1de |
2026-07-17 |
| evidence list for Benchmark 2 | 7b2a7a3fa8a49c3adda2e6604cdf462106f1ab4fc11a26c4246ba8e955324b09 |
2026-07-17 |
The hashes of both protocols match the entries in the evidence lists. Where a protocol states a number or a description that does not match the raw result files, this entry gives the value from the result files: (1) protocol no. 1 states that with zlib-9 the output worsens monotonically with the number of layers on every file — according to the data this holds for 5 of 10 files (section 1); (2) protocol no. 1 assesses F4 against zstd-19, although F4 is worded against the input size — section 4 gives both readings; (3) protocol no. 2 reports 30/30 verified decodes in the 2D test — the result file contains 28, and the session total of 48/48 is correct; (4) protocol no. 2 states that the majority predictor was worse than the left neighbour on 6 of 7 images — with the same coder it was worse on 7 of 7, and worse than both trivial predictors on 6 of 7; (5) protocol no. 2 describes jim___ah as a handwritten scan — 56% black pixels and no blank rows point rather to a dense or halftoned image, so we describe it neutrally.
Corrections to the entry of 2026-07-14
- Direction 1, criterion (b). Entry: “measured ratio on defined corpora against gzip/zstd — the measurement program has not yet been run”. Now: measured on 2026-07-15 and refuted (section 1); the axis is closed, in its 2D variant too (section 6).
- Direction 1, construction. Entry: “Construction (dependency rule, layer construction, container format): withheld”. The rule and the layer construction are now disclosed — XOR of neighbouring bits, iterated layer by layer — because the axis is falsified and closed.
- Direction 1, reversibility. The claim stands (100/100 and 4/4), with one precision: an earlier reference implementation stored only the first boundary bit and could not decode more than one layer; the measurements used a corrected re-implementation.
- Direction 2, reconstruction criterion. Entry: “Bit-exact reconstruction (met, hash-verified)”. For the PGA codec this was not true: the verification concerned an earlier proof of concept. The first full validation of the PGA codec (2026-07-15) gave 0/10; only the repaired variant is reversible (sections 3–4).
- Negative result of +56%. The size was measured correctly (14.2 → 22.2 MB), but the artefact is not fully reproducible and so was not a valid lossless encoding. The conclusion — expansion on high-entropy data — is confirmed by the repaired variant with SHA-256 verification: on the first 1 MiB of the same file +52% (format-faithful form) and +12% (redundancy-free map), on random data +55% and +18%.
- Direction 2, state of the art. The entry compared against gzip/zstd. The proper state of the art for bilevel images is CCITT G4 and JBIG2; G4 was measured in Benchmark 2, JBIG2 not yet.
- Direction 2, status. Entry: “POC implemented / in validation”. Now: validated — the historical version fails the reconstruction criterion, the repaired variant meets it, and the positive does not carry over beyond the synthetic raster (sections 3–5).
- Scoping conclusion. The entry narrowed the program to structured corpora (documents, scans, telemetry). The measurements narrow it further: on scans, rendered text pages, telemetry and text the classical coders win; what remains is the niche of born-digital rasters with rigid repetition (section 7).
Program status: active. Dependency-coding axis — closed as falsified, in 1D and 2D. PGA codec — reversible in the repaired variant, with a lead measured only on a single synthetic born-digital raster; next step: cross-document deduplication on series of pages from one template.
Proof of time
This text was hashed and stamped on publication. Verifying it takes both files: the stamp proves that one exact byte string existed at that moment, and only the source below is that byte string. Any later edit changes the hash and breaks the match — which is the point.
SHA-256 of the stamped text: 9655706edbeb5b0f3b47af0d9c98c8d37e0ff8d4084f884ee4b298b02a6d4476
- Sigelith anchor portfolio-dependibit-benchmarks.public.md.beattime.json
- OpenTimestamps proof portfolio-dependibit-benchmarks.public.md.ots
- Stamped source text portfolio-dependibit-benchmarks.public.md
Verify with: ots verify <proof> --file <source>