# Koch Laboratory — Kompresja strukturalna zmierzona (DependiBit) · Forschungsportfolio / Portfolio badawcze / Research Portfolio

> **Zasada uczciwości.** Ten wpis kontynuuje wpis z 2026-07-14 i go nie zastępuje: tamten pozostaje bez zmian, a jego stwierdzenia, które pomiary zmieniły, są skorygowane tutaj, w sekcji „Korekty wpisu z 2026-07-14". Wyniki negatywne są głównym produktem tej linii badawczej i podajemy je z liczbami. Wszystkie liczby pochodzą z surowych plików wyników; gdzie prywatny protokół pomiarowy podaje inaczej, obowiązuje plik wyników, a różnica jest wskazana w sekcji 9.
> **Zasada publikacji.** Dla kierunków wciąż otwartych publikujemy problem, stan techniki, kryterium i wynik — nie konstrukcję rozwiązania. Gdzie widnieje „Konstrukcja: wstrzymana.", szczegół techniczny jest utrzymywany jako materiał zgłoszeniowy. Oś kodowania zależnościowego jest zamknięta jako sfalsyfikowana, dlatego jej reguła jest tu podana jawnie.

---
---

# 🇵🇱 WERSJA POLSKA

## Koch Laboratory — kompresja strukturalna zmierzona: oś zależności zamknięta, nisza zawężona

Wpis z 2026-07-14 kończył się zapowiedzią zdefiniowanego benchmarku z jawnym korpusem i punktami odniesienia, bo stopień kompresji żadnego z kierunków nie był wtedy zmierzony. Ten wpis podaje wyniki dwóch sesji pomiarowych: benchmarku 1 (2026-07-15; 10 plików, 154 pomiary) i benchmarku 2 (2026-07-17; obrazy realistyczne, 51 pomiarów). Kryteria falsyfikowalne ustalono przed każdą sesją, a każde odkodowanie sprawdzano skrótem SHA-256 względem oryginału. Stwierdzenia wpisu z 2026-07-14, które pomiary zmieniły, są zebrane w sekcji „Korekty wpisu z 2026-07-14".

Bilans w trzech zdaniach. Oś kodowania zależnościowego jest obalona w 1D i w 2D i zostaje zamknięta. Historyczna wersja kodeka PGA nie odtwarzała danych, które sama zakodowała. Po naprawie kodek wygrywa z klasycznymi punktami odniesienia wyłącznie na rastrze generowanym cyfrowo, w którym powtórzenia są sztywne i dokładne co do bitu — na realnych skanach i renderowanych stronach przegrywa.

**Metodyka.** Problem → stan techniki → kryterium ustalone przed pomiarem → metoda → wynik z warunkami brzegowymi → następny krok. Stopień kompresji = rozmiar wyniku / rozmiar wejścia: mniej znaczy lepiej, powyżej 1 oznacza ekspansję. Do wyniku liczony jest cały kontener — nagłówki, metadane, bity brzegowe, suma kontrolna — a nie sam ładunek. Oznaczenia kryteriów (F1–F7) odpowiadają prywatnym protokołom pomiarowym, których skróty podaje sekcja 9.

> **Zasada publikacji.** Dla kierunków wciąż otwartych publikujemy problem, stan techniki, kryterium i wynik — nie konstrukcję. Gdzie widnieje „Konstrukcja: wstrzymana.", szczegół techniczny jest utrzymywany jako materiał zgłoszeniowy. Oś kodowania zależnościowego jest zamknięta jako sfalsyfikowana, dlatego jej reguła jest tu podana jawnie.

### 1. Kierunek 1 — kodowanie zależnościowe 1D: reguła ujawniona, hipoteza obalona
**Pytanie badawcze.** Czy odwracalna reprezentacja strumienia jako warstw relacji między sąsiednimi bitami daje po klasycznym koderze końcowym wynik mniejszy niż ten sam koder bez transformacji oraz niż zlib-9 i zstd-19?
**Reguła (ujawniona, bo oś jest zamknięta).** Pierwsza warstwa to XOR każdej pary sąsiednich bitów, `b[i] ⊕ b[i+1]`; każda kolejna powstaje z poprzedniej według tej samej reguły. Odwrócenie L warstw wymaga L bitów brzegowych.
**Uwaga analityczna.** Dla zmierzonych głębokości L = 1, 2, 4, 8 iteracja daje dokładnie `b[i] ⊕ b[i+L]` (współczynniki dwumianowe modulo 2): warstwa L jest różnicą XOR na odległość L, a osiem warstw to różnica XOR sąsiednich bajtów. Przy zmierzonych głębokościach transformacja mieści się więc w klasycznym kodowaniu różnicowym.
**Stan techniki (publikowany).** Kodowanie różnicowe i predykcyjne (w tym różnica XOR), RLE, DEFLATE (zlib, gzip), zstd.
**Kryteria (ustalone przed pomiarem).** F1: SHA-256 rekonstrukcji równy SHA-256 oryginału przy każdym odkodowaniu. F2: warstwa L z koderem końcowym daje wynik mniejszy niż ten sam kontener bez transformacji (L = 0), dla L ∈ {1, 2, 4, 8}. F3: wynik mniejszy niż zlib-9 i niż zstd-19.
**Metoda.** Korpus 10 plików (sekcja 8); L ∈ {0, 1, 2, 4, 8}; koder końcowy RLE albo zlib-9; razem 100 pomiarów. Mierzono kanoniczną reimplementację z kompletem bitów brzegowych — wcześniejsza implementacja referencyjna zapisywała tylko pierwszy z nich, więc nie mogła odkodować więcej niż jednej warstwy.
**Wynik.** Stopień kompresji z koderem końcowym zlib-9 (cały kontener) oraz punkty odniesienia na surowych danych:

| plik | L = 0 (kontrola) | L = 1 | L = 2 | L = 4 | L = 8 | zlib-9 | zstd-19 |
|---|---|---|---|---|---|---|---|
| tekst techniczny | 0,2929 | 0,2972 | 0,3083 | 0,3261 | 0,3541 | 0,2903 | 0,2778 |
| syntetyczny raster formularza | 0,1411 | 0,1439 | 0,1581 | 0,1714 | 0,1720 | 0,1411 | 0,1069 |
| telemetria (CSV) | 0,1623 | 0,1624 | 0,1648 | 0,1694 | 0,2050 | 0,1622 | 0,0740 |
| plik rzadki | 0,0028 | 0,0029 | 0,0030 | 0,0033 | 0,0034 | 0,0026 | 0,0014 |
| długie serie bitów | 0,0067 | 0,0065 | 0,0065 | 0,0065 | 0,0065 | 0,0065 | 0,0094 |
| dane losowe | 1,0005 | 1,0005 | 1,0005 | 1,0005 | 1,0005 | 1,0003 | 1,0001 |
| wycinek pliku wykonywalnego | 0,8641 | 0,8701 | 0,8762 | 0,8838 | 0,8937 | 0,8640 | 0,8509 |
| same zera | 0,0012 | 0,0012 | 0,0012 | 0,0012 | 0,0012 | 0,0011 | 0,0001 |
| bity naprzemienne | 0,0012 | 0,0012 | 0,0012 | 0,0012 | 0,0012 | 0,0011 | 0,0001 |
| plik okresowy | 0,0034 | 0,0035 | 0,0036 | 0,0034 | 0,0028 | 0,0033 | 0,0004 |

Z koderem zlib-9 wynik pogarsza się monotonicznie z liczbą warstw na pięciu plikach: tekście technicznym, syntetycznym rastrze formularza, telemetrii, pliku rzadkim i wycinku pliku wykonywalnego. Wyjątki są wyłącznie syntetyczne. Na pliku długich serii jedna warstwa oszczędza 48 B względem tego samego kontenera bez transformacji (1707 wobec 1755 B), co daje dokładnie wynik samego zlib-9 (1707 B). Na pliku okresowym osiem warstw oszczędza 175 B (725 wobec 900 B), a zstd-19 koduje ten plik do 97 B. Na danych losowych i w obu przypadkach zdegenerowanych (same zera, bity naprzemienne) wynik jest stały z dokładnością do 1 B. Z koderem RLE transformacja zamieniła ekspansję w kompresję tylko na jednym pliku — pliku bitów naprzemiennych (16,00 → 0,0629 przy jednej warstwie), który zlib-9 bez żadnej transformacji koduje do 0,0011. Tam, gdzie warstwy zmniejszały wynik RLE na innych plikach, pozostawał on większy od wejścia (np. raster formularza 1,51, telemetria 5,74, tekst techniczny 6,33).
**Rozstrzygnięcie kryteriów.** F1: spełnione, 100/100. F2: ze zlib-9 spełnione tylko na dwóch plikach syntetycznych (długie serie, plik okresowy), z RLE tylko tam, gdzie wynik nadal jest większy od wejścia, oraz na bitach naprzemiennych. F3: niespełnione na 10 z 10 plików — żaden wynik z transformacją nie był mniejszy jednocześnie od zlib-9 i od zstd-19. **Hipoteza obalona.**
**Mechanizm (zmierzony).** Na tekście i telemetrii pierwsze warstwy pchają gęstość jedynek ku 0,5: tekst techniczny 0,3932 → 0,4607 (L = 1) → 0,4805 (L = 2), telemetria 0,4261 → 0,5119 (L = 1). Transformacja zaciera więc strukturę widzianą przez koder końcowy, zamiast ją eksponować. Gdzie gęstość spada, wynik i tak się nie poprawia: na syntetycznym rastrze formularza jedna warstwa obniża gęstość z 0,1918 do 0,1365, a stopień kompresji ze zlib-9 rośnie z 0,1411 do 0,1439. Ten sam wzorzec powtarza się w 2D (sekcja 6).
**Następny krok.** Brak — oś zamknięta (sekcja 7).
**Status:** sfalsyfikowane pomiarowo (F3 niespełnione na 10/10 plików; F1 spełnione 100/100); oś zamknięta.

### 2. Kierunek 1 — pełna piramida warstw: odwracalna, ale kwadratowa
**Pytanie badawcze.** Czy historyczna implementacja, która buduje pełną piramidę warstw aż do pojedynczego bitu, jest wykonalna?
**Kryteria.** Jak w sekcji 1 (F1–F3).
**Metoda.** Implementacja historyczna bez modyfikacji, uruchomiona na prefiksach tekstu technicznego o długości od 128 do 1024 B.
**Wynik.**

| wejście | wyjście | wyjście / wejście | czas kodowania |
|---|---|---|---|
| 128 B | 515 415 B | ×4027 | 0,086 s |
| 256 B | 2 055 749 B | ×8030 | 0,257 s |
| 512 B | 8 285 223 B | ×16 182 | 0,969 s |
| 1024 B | 33 310 203 B | ×32 529 | 3,844 s |

Piramida dla n bitów wejścia zawiera n(n−1)/2 bitów zależności; zmierzone wyjście odpowiada n(n−1)/2 bajtom z dokładnością do 2 % (98,1–99,3 %), czyli ok. jednemu bajtowi na bit zależności. Rozmiar i czas rosną kwadratowo. Ekstrapolacja na plik wykonywalny 14,2 MB z wpisu z 2026-07-14: ok. 6,5 PB wyjścia i ok. 23 lata obliczeń na maszynie pomiarowej. F1 spełnione (4/4); F2 i F3 obalone — wynik jest od ok. 4 tys. do ok. 32,5 tys. razy większy od wejścia. Uwaga o konstrukcji: dekoder tej implementacji czyta wyłącznie pierwszą warstwę (z jednym bitem brzegowym wystarcza ona do odtworzenia), więc zapis warstw od drugiej wzwyż jest z założenia nadmiarowy.
**Następny krok.** Brak — implementacja wycofana.
**Status:** odwracalna (4/4), niewykonalna poza rozmiarami zabawkowymi; wycofana.

### 3. Kierunek 2 — kodek PGA w wersji historycznej: odtworzenie nieudane na 10 z 10 plików
**Pytanie badawcze.** Czy wersja kodeka, na której zmierzono opisaną 2026-07-14 ekspansję o 56 %, odtwarza dane co do bitu?
**Kryteria (ustalone przed pomiarem).** F1 dla kodeka PGA oraz F4: rozmiar całego kontenera (wszystkie sekcje i metadane) mniejszy od wejścia na korpusie strukturalnym, przy spełnionym F1.
**Metoda.** Wersja historyczna bez modyfikacji; zakodowanie i odkodowanie każdego z 10 plików korpusu; SHA-256 wyniku względem oryginału.
**Wynik: kryterium odtworzenia obalone.** Kodowanie przebiegło na 10 z 10 plików, odkodowanie — na 0 z 10: na 4 plikach przerwało się błędem, na 6 dało dane niezgodne z SHA-256 oryginału. Analiza kodu wskazała dwa defekty formatu: warstwa resztkowa nie niosła danych, które miała odtwarzać, a sekcja mapy nie była jednoznacznie parsowalna. Konsekwencja wykracza poza benchmark: historyczny artefakt 22,2 MB z wpisu z 2026-07-14 również nie jest w pełni odtwarzalny. Jego rozmiar zmierzono poprawnie (14,2 → 22,2 MB), ale nie był to poprawny zapis bezstratny. F4 nie podlega ocenie, dopóki F1 nie jest spełnione.
**Następny krok.** Naprawa obu defektów w wariancie pomiarowym (sekcja 4).
**Status:** wynik negatywny — wersja historyczna nie spełnia kryterium odtworzenia (0/10); korekta wpisu z 2026-07-14 poniżej.

### 4. Kierunek 2 — po naprawie: warunkowy pozytyw, który nie uogólnił się na dane realistyczne
**Pytanie badawcze.** Czy po naprawie formatu kodek odtwarza dane co do bitu i daje dodatni bilans rozmiaru — najpierw na korpusie benchmarku 1, potem na obrazach realistycznych?
**Stan techniki (publikowany).** Dla obrazów dwupoziomowych: CCITT G4 (ITU-T T.6, dwuwymiarowe kodowanie serii względem wiersza odniesienia) oraz JBIG2 (ITU-T T.88), który sam koduje tekst przez słowniki symboli z dopasowaniem wzorców. To one, a nie gzip i zstd, są właściwą konkurencją dla tej klasy danych. G4 zmierzono w benchmarku 2; JBIG2 nie został zmierzony — w środowisku pomiarowym nie było jego kodera.
**Kryteria (ustalone przed pomiarem).** F4 (benchmark 1): kontener mniejszy od wejścia na korpusie strukturalnym, przy spełnionym F1. F5 (benchmark 2): stopień kompresji mniejszy niż zstd-19 **i** niż CCITT G4 na korpusie realistycznym.
**Metoda.** Wariant pomiarowy z naprawą obu defektów, w dwóch odmianach: wiernej formatowi historycznemu oraz z sekcją mapy pozbawioną nadmiarowości (konstrukcja wstrzymana). Benchmark 2 używa drugiej odmiany na siedmiu obrazach dwupoziomowych (sekcja 8), z prawdziwą geometrią rastra.
**Wynik — benchmark 1.** F1 przywrócone: 10/10 w obu odmianach. F4 w brzmieniu protokołu okazało się słabe: kontener mniejszy od wejścia ma odmiana wierna formatowi na 4 z 5 plików strukturalnych (na tekście technicznym 1,098), a odmiana z mapą bez nadmiarowości na 5 z 5. Rozstrzyga porównanie z punktami odniesienia. Odmiana wierna formatowi była większa od zstd-19 na każdym z 10 plików; na syntetycznym rastrze formularza 97 % jej kontenera zajmowała mapa. Odmiana z mapą bez nadmiarowości dała pierwszy zmierzony pozytyw programu — na jednym pliku: syntetyczny raster formularza 0,0495 wobec 0,1069 (zstd-19) i 0,1411 (zlib-9), czyli 2,2× lepiej od najlepszego punktu odniesienia. Na pozostałych 9 plikach punkty odniesienia wygrały (np. telemetria 0,2981 wobec 0,0740, tekst techniczny 0,7949 wobec 0,2778). Na danych wysokoentropijnych kodek rozszerza dane: dane losowe 1,176 (odmiana wierna formatowi: 1,547), wycinek pliku wykonywalnego 1,115 (1,519). Stan po benchmarku 1: pozytyw warunkowy — raster syntetyczny, glify wyrównane do siatki kafli, bez szumu, bez porównania z G4 i JBIG2.
**Wynik — benchmark 2: F5 obalone.**

| obraz | PGA | zlib-9 | zstd-19 | CCITT G4 | najlepszy |
|---|---|---|---|---|---|
| `fax2d` — faks CCITT, 1728×1082 | 0,1888 | 0,1374 | 0,1241 | **0,1209** | G4 |
| `g3test` — faks CCITT, 1728×1103 | 0,2629 | 0,1806 | **0,1624** | 0,1756 | zstd-19 |
| `jim___ah` — gęsty obraz testowy, 664×813 | 0,2443 | 0,2083 | **0,1910** | 0,6301 | zstd-19 |
| strona tekstu, font bezszeryfowy | 0,0489 | 0,0440 | **0,0364** | 0,0376 | zstd-19 |
| strona tekstu, font szeryfowy | 0,0348 | 0,0376 | **0,0302** | 0,0346 | zstd-19 |
| strona formularza | 0,0351 | 0,0263 | **0,0133** | 0,0372 | zstd-19 |
| syntetyczny raster formularza (kontrola z benchmarku 1) | **0,0492** | 0,1411 | 0,1069 | 0,4030 | PGA |

Kodek przegrywa ze zstd-19 na 6 z 6 obrazów realistycznych, z CCITT G4 na 4 z 6. Jedynym obrazem, na którym którakolwiek własna metoda programu wypada najlepiej, pozostaje syntetyczny raster wyrównany do siatki kafli. F1: 7/7. Uwaga metodyczna: w benchmarku 1 kodek rasteryzował wejście bez znajomości jego prawdziwej geometrii, a szerokość rastra zgodziła się z prawdziwą szerokością syntetycznego rastra formularza przypadkiem; benchmark 2 używa prawdziwej geometrii we wszystkich pomiarach.
**Następny krok.** Ustalić, od czego zależy jedyny pozytyw (sekcja 5).
**Status:** odwracalność przywrócona (10/10, 7/7); generalizacja obalona (F5); pozytyw ograniczony do syntetycznego rastra.

### 5. Kierunek 2 — mechanizm pozytywu: sztywna powtarzalność co do bitu
**Pytanie badawcze.** Od czego zależy jedyny pozytyw — od wyrównania wzorców do siatki kafli, od ich dokładności co do bitu, czy od obu?
**Kryteria.** Test wrażliwości zapowiedziany w protokole nr 1, bez progu sukcesu (mierzona jest krzywa degradacji wobec G4 i zstd-19 na tych samych danych); F7 (ustalone przed pomiarem): wybór przesunięcia siatki przywraca stopień kompresji rastra wyrównanego.
**Metoda.** Syntetyczny raster formularza przesuwany poziomo o mniej niż szerokość kafla oraz zaszumiany przez losową zmianę bitów z prawdopodobieństwem od 0,1 do 2 %; każdy wynik z weryfikacją SHA-256 (13/13).
**Wynik — przesunięcie.** Zależnie od wielkości przesunięcia stopień kompresji od 0,0492 do 0,0880, słownik do 46 wpisów zamiast 16; G4 bez zmian (0,403).
**Wynik — szum.**

| zmienione bity | PGA | wpisy słownika | zstd-19 | CCITT G4 |
|---|---|---|---|---|
| 0 | 0,0492 | 16 | 0,1069 | 0,4030 |
| 0,1 % | 0,0713 | 20 | 0,1327 | 0,4158 |
| 0,5 % | 0,1290 | 78 | 0,1969 | 0,4652 |
| 1 % | 0,1991 | 561 | 0,2555 | 0,5195 |
| 2 % | 0,3851 | 2656 | 0,3434 | 0,6180 |

Przewaga nad zstd-19 maleje z 2,2× na danych czystych do 1,9×, 1,5× i 1,3×, a przy 2 % zmienionych bitów odwraca się (0,3851 wobec 0,3434). Stopień kompresji kodeka pogarsza się 7,8-krotnie, G4 — 1,5-krotnie. Wniosek: pozytyw wymaga sztywnej, dokładnej co do bitu powtarzalności wzorców — takiej, jaką mają rastry generowane cyfrowo (born-digital), a nie skany.
**Wynik — F7: potwierdzone technicznie.** Raster celowo przesunięty względem siatki kafli: 0,1416; po wyborze przesunięcia siatki: 0,0492, czyli dokładnie wartość rastra wyrównanego; koszt: kilka bitów; F1 spełnione. Dotyczy to przesunięcia całego rastra; na renderowanych stronach i skanach, gdzie położenie znaków zmienia się indywidualnie, nie było to mierzone.
**Konstrukcja: wstrzymana.**
**Następny krok.** Hipoteza deduplikacji między dokumentami (sekcja 7).
**Status:** warunki brzegowe zmierzone; F7 potwierdzone technicznie.

### 6. Kierunek 3 — zależność przestrzenna 2D: hipoteza obalona, z ablacją predyktorów
**Pytanie badawcze.** Czy reguła zależności przeniesiona z 1D na sąsiadów przestrzennych w rastrze niesie sygnał, którego zabrakło w 1D?
**Stan techniki (publikowany).** Modelowanie kontekstowe obrazów dwupoziomowych w JBIG i JBIG2: szablon kontekstu z 10–16 sąsiednich pikseli i adaptacyjny koder arytmetyczny. Tu świadomie zmierzono wariant minimalny, żeby rozstrzygnąć, czy oś w ogóle niesie sygnał.
**Kryteria (ustalone przed pomiarem).** F6(a): predyktor większościowy (większość z trzech sąsiadów: lewego, górnego i lewego-górnego) jest lepszy od predyktorów trywialnych (sam lewy albo sam górny sąsiad). F6(b): wynik jest mniejszy od najlepszego punktu odniesienia dla danego obrazu.
**Metoda.** Residuum = piksel ⊕ predykcja; dekoder kauzalny; koder końcowy zlib-9 albo zstd-19; 7 obrazów × 4 warianty = 28 pomiarów; F1: 28/28.
**Wynik.**

| obraz | lewy + zlib-9 | górny + zlib-9 | większość + zlib-9 | większość + zstd-19 | najlepszy punkt odniesienia |
|---|---|---|---|---|---|
| `fax2d` | 0,1359 | 0,1646 | 0,1691 | 0,1567 | 0,1209 (G4) |
| `g3test` | 0,1800 | 0,2178 | 0,2355 | 0,2154 | 0,1624 (zstd-19) |
| `jim___ah` | 0,2085 | 0,3058 | 0,2849 | 0,2642 | 0,1910 (zstd-19) |
| strona tekstu, font bezszeryfowy | 0,0450 | 0,0456 | 0,0554 | 0,0469 | 0,0364 (zstd-19) |
| strona tekstu, font szeryfowy | 0,0387 | 0,0409 | 0,0465 | 0,0376 | 0,0302 (zstd-19) |
| strona formularza | 0,0272 | 0,0264 | 0,0316 | 0,0165 | 0,0133 (zstd-19) |
| syntetyczny raster formularza | 0,1412 | 0,1610 | 0,2447 | 0,2189 | 0,1069 (zstd-19) |

F6(a) obalone: przy tym samym koderze (zlib-9) predyktor większościowy był gorszy od lewego sąsiada na 7 z 7 obrazów i gorszy od obu predyktorów trywialnych na 6 z 7 (wyjątek: na `jim___ah` pokonał górnego sąsiada). F6(b) obalone: żaden wariant nie był lepszy od najlepszego punktu odniesienia na żadnym obrazie (0/7). Najlepszy wariant był lepszy od samego G4 na 3 z 7 obrazów (`jim___ah`, strona formularza, syntetyczny raster), a od zstd-19 na surowych danych — na żadnym.
**Mechanizm.** Residuum jest rzadsze od oryginału, a mimo to nie kompresuje się lepiej. Na `jim___ah` ma 0,165 jedynki na piksel wobec 0,557 w oryginale, a ze zlib-9 daje praktycznie ten sam wynik (0,2085 wobec 0,2083). Przy tym samym koderze residuum lewego sąsiada było najwyżej nieznacznie mniejsze od oryginału (`fax2d`: 0,1359 wobec 0,1374). Interpretacja spójna z oboma pomiarami: koder z rodziny LZ zyskuje na powtórzeniach — identycznych glifach i wierszach — a nie na gęstości bitów; transformacja zależnościowa zmienia gęstość (w 1D ku 0,5, w 2D w dół), ale powtórzeń nie dodaje, a część istniejących rozbija. Zysk z rzadkości nie rekompensuje utraty dopasowań.
**Następny krok.** Brak w tej formie: dalsza praca oznaczałaby modelowanie kontekstowe z koderem arytmetycznym, czyli zbieżność ze stanem techniki (JBIG/JBIG2), a nie nową oś.
**Status:** sfalsyfikowane (F6a, F6b); oś zamknięta.

### 7. Wniosek programu: oś zależności zamknięta, nisza zawężona
- **Oś kodowania zależnościowego — zamknięta jako sfalsyfikowana**, w 1D (sekcje 1–2) i w 2D (sekcja 6), przy koderach końcowych klasy RLE i LZ. Zapisany mechanizm: transformacja zmienia gęstość bitów, ale nie dodaje powtórzeń, na których opiera się koder.
- **Kodek PGA — nisza zawężona.** Program przeszedł od „lepszego kompresora ogólnego" (obalonego w iteracjach 2021–2025), przez „korpusy strukturalne" (dokumenty, skany, telemetria — obalone pomiarem dla skanów, renderowanych stron, tekstu i telemetrii), do wąskiej klasy danych: rastrów generowanych cyfrowo ze sztywną, dokładną co do bitu powtarzalnością (formularze generowane programowo, dane z kolejki wydruku, dokumenty wypełniane z szablonu). Tam kodek był ok. 2,2× lepszy od najlepszego klasycznego punktu odniesienia — zmierzone z pełnym odtworzeniem — i ok. 2,6× po dalszym dopracowaniu kontenera, co jest oszacowaniem z rozmiarów sekcji, jeszcze niezmierzonym od początku do końca. Konstrukcja: wstrzymana.
- **Warunki brzegowe niszy, podane wprost.** Przewagę zmierzono na jednym syntetycznym rastrze; porównanie objęło G4, zlib-9 i zstd-19, nie JBIG2, który stosuje słowniki symboli i jest najbliższym stanem techniki; odporność na szum jest słaba (sekcja 5). To hipoteza zakresowa, nie wykazana przewaga na danych rzeczywistych.
- **Następna hipoteza: deduplikacja między dokumentami.** Wspólny słownik dla serii stron z jednego szablonu (np. serii faktur lub protokołów); oczekiwanie: koszt słownika amortyzuje się w serii, a stopień kompresji serii jest znacznie niższy niż pojedynczej strony. JBIG2 przewiduje słowniki symboli wspólne dla wielu stron, więc bez porównania z nim wynik tej hipotezy nie będzie rozstrzygający. Kryterium zostanie ustalone przed pomiarem, jak w obu sesjach.

### 8. Korpus, punkty odniesienia, odtwarzalność
**Benchmark 1 (2026-07-15), 10 plików.** Pięć strukturalnych: tekst techniczny (dokumentacja projektu, 18 309 B), syntetyczny raster formularza 1 bit/piksel (512 KiB; linie tabeli i mały stały zestaw glifów ułożonych na siatce kafli, bez szumu), syntetyczna telemetria w formacie CSV (418 929 B), plik rzadki (256 KiB; zera i 16-bajtowy nagłówek rekordu co 4096 B), plik długich serii bajtów 0x00 i 0xFF (256 KiB). Dwa wysokoentropijne pliki kontrolne: dane pseudolosowe (256 KiB) i pierwszy 1 MiB skompilowanego pliku wykonywalnego Windows z wpisu z 2026-07-14. Trzy przypadki skrajne po 256 KiB: same zera, bity naprzemienne (bajt 0xAA), sekwencja okresowa (bajty 0–63). Pliki syntetyczne są deterministyczne (stałe ziarno generatora).
**Benchmark 2 (2026-07-17), 7 obrazów dwupoziomowych.** Trzy rzeczywiste obrazy z publicznego archiwum obrazów testowych libtiff `pics-3.8.0` (download.osgeo.org/libtiff/pics-3.8.0.tar.gz): `fax2d` (1728×1082) i `g3test` (1728×1103) — klasyczne faksowe obrazy testowe CCITT — oraz `jim___ah` (664×813; gęsty: 56 % czarnych pikseli, ani jednego pustego wiersza). Trzy strony renderowane fontami systemowymi TrueType przy ok. 200 dpi (1656×2336): tekst fontem bezszeryfowym, tekst fontem szeryfowym oraz formularz z liniami i etykietami fontem o stałej szerokości znaków; położenie wierszy i znaków celowo niewyrównane do siatki kafli. Kontrola: syntetyczny raster formularza z benchmarku 1. Rastry zapisane jako surowe bity, 1 bit/piksel, wiersz po wierszu, czarny = 1.
**Punkty odniesienia.** zlib-9 (DEFLATE, jak w gzip), zstd-19, CCITT G4 (w benchmarku 2; mierzony jako cały plik TIFF). JBIG2 — niezmierzony: brak kodera w środowisku pomiarowym.
**Środowisko.** Windows 11 (AMD64), Python 3.12.10, numpy 2.5.0, Pillow 10.4.0, zstandard 0.25.0.
**Weryfikacja.** Benchmark 1: 154 pomiary, 144 odkodowania zgodne z SHA-256 oryginału; 10 niezgodnych to wersja historyczna kodeka PGA (sekcja 3). Benchmark 2: 51 pomiarów, w tym 48 odkodowań — 48/48 zgodnych; pozostałe 3 pomiary dotyczyły szczegółu konstrukcji kontenera i ich wynik jest wstrzymany.
**Odtwarzalność.** Wyniki punktów odniesienia na obrazach z archiwum libtiff może powtórzyć każdy: archiwum jest publiczne, a geometria i format rastra są podane wyżej. Wyniki własnych metod wymagają konstrukcji wstrzymanej; z protokołami wiążą je skróty z sekcji 9.

### 9. Dowody: skróty SHA-256 i stemple czasu
Protokoły pomiarowe pozostają prywatne: zawierają konstrukcję wstrzymaną. Ich skróty SHA-256, opublikowane tutaj, pozwalają później sprawdzić, że dokument okazany w przyszłości jest tym samym, który ostemplowano w lipcu 2026 r. Wykazy materiału dowodowego zawierają skróty SHA-256 surowych plików wyników i manifestów korpusu oraz identyfikatory commitów kodu pomiarowego (pierwszy z nich także skrót whitepapera z 2025 r.).

| dokument | SHA-256 | stempel OpenTimestamps |
|---|---|---|
| protokół pomiarowy nr 1 (benchmark 1) | `0830cc94c042c05103e45a47de2bdae0db2411f683fcd309e6630c5700552375` | 2026-07-15, atestacja Bitcoin |
| wykaz materiału dowodowego benchmarku 1 | `8b85e99708febb32682a100f5b2336bd0ffc8cf0e0ef4c7671aa91576de6dc56` | 2026-07-15, atestacja Bitcoin |
| protokół pomiarowy nr 2 (benchmark 2) | `c478b509fda7419a60b0ad2149d28f4dec83aae5fb074ffbeab0cd3bb3aba1de` | 2026-07-17 |
| wykaz materiału dowodowego benchmarku 2 | `7b2a7a3fa8a49c3adda2e6604cdf462106f1ab4fc11a26c4246ba8e955324b09` | 2026-07-17 |

Skróty obu protokołów zgadzają się z wpisami w wykazach. Tam, gdzie protokół podaje liczbę lub opis niezgodny z surowymi plikami wyników, ten wpis podaje wartość z plików wyników: (1) protokół nr 1 stwierdza, że ze zlib-9 wynik pogarsza się monotonicznie z liczbą warstw na każdym pliku — według danych dotyczy to 5 z 10 plików (sekcja 1); (2) protokół nr 1 ocenia F4 względem zstd-19, choć F4 jest sformułowane względem rozmiaru wejścia — sekcja 4 podaje oba odczyty; (3) protokół nr 2 podaje 30/30 zweryfikowanych odkodowań w teście 2D — plik wyników zawiera ich 28, a suma sesji 48/48 jest poprawna; (4) protokół nr 2 podaje, że predyktor większościowy był gorszy od lewego sąsiada na 6 z 7 obrazów — przy tym samym koderze był gorszy na 7 z 7, a od obu predyktorów trywialnych na 6 z 7; (5) protokół nr 2 opisuje `jim___ah` jako skan odręczny — 56 % czarnych pikseli i brak pustych wierszy wskazują raczej na obraz gęsty lub rastrowany, dlatego opisujemy go neutralnie.

### Korekty wpisu z 2026-07-14
- **Kierunek 1, kryterium (b).** Wpis: „zmierzone ratio na zdefiniowanych korpusach względem gzip/zstd — program pomiarowy nie został jeszcze uruchomiony". Teraz: zmierzone 2026-07-15 i obalone (sekcja 1); oś zamknięta, także w wariancie 2D (sekcja 6).
- **Kierunek 1, konstrukcja.** Wpis: „Konstrukcja (reguła zależności, budowa warstw, format kontenera): wstrzymana". Reguła i budowa warstw są teraz ujawnione — XOR sąsiednich bitów, iterowany warstwa po warstwie — bo oś jest sfalsyfikowana i zamknięta.
- **Kierunek 1, odwracalność.** Twierdzenie podtrzymane (100/100 i 4/4), z doprecyzowaniem: jedna wcześniejsza implementacja referencyjna zapisywała tylko pierwszy bit brzegowy i nie mogła odkodować więcej niż jednej warstwy; pomiary wykonano na poprawionej reimplementacji.
- **Kierunek 2, kryterium odtworzenia.** Wpis: „Rekonstrukcja co do bitu (spełnione, weryfikacja skrótem)". Dla kodeka PGA było to nieprawdziwe: weryfikacja dotyczyła wcześniejszego dowodu koncepcji. Pierwsza pełna walidacja kodeka PGA (2026-07-15) dała 0/10; odwracalność ma dopiero wariant naprawiony (sekcje 3–4).
- **Wynik negatywny +56 %.** Rozmiar zmierzono poprawnie (14,2 → 22,2 MB), ale artefakt nie jest w pełni odtwarzalny, więc nie był poprawnym zapisem bezstratnym. Wniosek — ekspansja na danych wysokoentropijnych — zostaje potwierdzony przez wariant naprawiony z weryfikacją SHA-256: na pierwszym 1 MiB tego samego pliku +52 % (odmiana wierna formatowi) i +12 % (mapa bez nadmiarowości), na danych losowych +55 % i +18 %.
- **Kierunek 2, stan techniki.** Wpis porównywał z gzip/zstd. Właściwym stanem techniki dla obrazów dwupoziomowych są CCITT G4 i JBIG2; G4 zmierzono w benchmarku 2, JBIG2 jeszcze nie.
- **Kierunek 2, status.** Wpis: „POC zaimplementowany / w walidacji". Teraz: zwalidowany — wersja historyczna nie spełnia kryterium odtworzenia, wariant naprawiony je spełnia, pozytyw nie przeniósł się poza syntetyczny raster (sekcje 3–5).
- **Wniosek zakresowy.** Wpis zawężał program do korpusów strukturalnych (dokumenty, skany, telemetria). Pomiary zawężają dalej: na skanach, renderowanych stronach tekstu, telemetrii i tekście wygrywają klasyczne kodery; zostaje nisza rastrów generowanych cyfrowo ze sztywną powtarzalnością (sekcja 7).

**Status programu:** aktywny. Oś kodowania zależnościowego — zamknięta jako sfalsyfikowana, w 1D i 2D. Kodek PGA — odwracalny w wariancie naprawionym, z przewagą zmierzoną wyłącznie na jednym syntetycznym rastrze generowanym cyfrowo; najbliższy krok: deduplikacja między dokumentami na seriach stron z jednego szablonu.

---
---

# 🇩🇪 DEUTSCHE FASSUNG

## Koch Laboratory — strukturbasierte Kompression, gemessen: Abhängigkeitsachse geschlossen, Nische eingegrenzt

Der Eintrag vom 2026-07-14 endete mit der Ankündigung eines definierten Benchmarks mit offengelegtem Korpus und Referenzverfahren, weil das Kompressionsverhältnis für keine der beiden Richtungen gemessen worden war. Dieser Eintrag berichtet über die Ergebnisse zweier Messsitzungen: Benchmark 1 (2026-07-15; 10 Dateien, 154 Messungen) und Benchmark 2 (2026-07-17; realistische Bilder, 51 Messungen). Die falsifizierbaren Kriterien wurden vor jeder Sitzung festgelegt, und jede Dekodierung wurde per SHA-256 gegen das Original geprüft. Aussagen des Eintrags vom 2026-07-14, die durch die Messungen überholt sind, stehen gesammelt im Abschnitt „Korrekturen zum Eintrag vom 2026-07-14".

Die Bilanz in drei Sätzen. Die Achse der Abhängigkeitskodierung ist in 1D und in 2D widerlegt und wird geschlossen. Die historische Version des PGA-Codecs konnte die Daten, die sie selbst kodiert hatte, nicht wiederherstellen. Nach der Reparatur schlägt der Codec die klassischen Referenzverfahren nur auf einem digital erzeugten Raster, dessen Wiederholungen starr und bitgenau sind — auf echten Scans und gerenderten Seiten unterliegt er.

**Methodik.** Problem → Stand der Technik → vor der Messung festgelegtes Kriterium → Methode → Ergebnis mit Randbedingungen → nächster Schritt. Kompressionsverhältnis = Ausgabegröße / Eingabegröße: kleiner ist besser, über 1 bedeutet Expansion. Gezählt wird der gesamte Container — Header, Metadaten, Randbits, Prüfsumme — nicht nur die Nutzlast. Die Kriterienbezeichnungen (F1–F7) entsprechen den privaten Messprotokollen, deren Hashwerte Abschnitt 9 angibt.

> **Publikationshinweis.** Für noch offene Richtungen veröffentlichen wir Problem, Stand der Technik, Kriterium und Ergebnis — nicht die Konstruktion. Wo „Konstruktion: zurückgehalten." steht, wird das technische Detail als Anmeldematerial vorgehalten. Die Achse der Abhängigkeitskodierung ist als falsifiziert geschlossen; deshalb wird ihre Regel hier offengelegt.

### 1. Richtung 1 — Abhängigkeitskodierung in 1D: Regel offengelegt, Hypothese widerlegt
**Forschungsfrage.** Liefert eine umkehrbare Darstellung des Stroms als Schichten von Relationen zwischen Nachbarbits nach einem klassischen Endkodierer ein kleineres Ergebnis als derselbe Kodierer ohne Transformation sowie als zlib-9 und zstd-19?
**Regel (offengelegt, da die Achse geschlossen ist).** Die erste Schicht ist das XOR jedes Paares benachbarter Bits, `b[i] ⊕ b[i+1]`; jede weitere Schicht entsteht nach derselben Regel aus der vorigen. Die Umkehrung von L Schichten erfordert L Randbits.
**Analytische Anmerkung.** Für die gemessenen Tiefen L = 1, 2, 4, 8 ergibt die Iteration genau `b[i] ⊕ b[i+L]` (Binomialkoeffizienten modulo 2): Schicht L ist eine XOR-Differenz im Abstand L, acht Schichten sind die XOR-Differenz benachbarter Bytes. Bei den gemessenen Tiefen liegt die Transformation damit innerhalb der klassischen Differenzkodierung.
**Stand der Technik (veröffentlicht).** Differenz- und prädiktive Kodierung (einschließlich XOR-Differenz), RLE, DEFLATE (zlib, gzip), zstd.
**Kriterien (vor der Messung festgelegt).** F1: SHA-256 der Rekonstruktion gleich SHA-256 des Originals bei jeder Dekodierung. F2: Schicht L mit Endkodierer ergibt ein kleineres Ergebnis als derselbe Container ohne Transformation (L = 0), für L ∈ {1, 2, 4, 8}. F3: Ergebnis kleiner als zlib-9 und als zstd-19.
**Methode.** Korpus aus 10 Dateien (Abschnitt 8); L ∈ {0, 1, 2, 4, 8}; Endkodierer RLE oder zlib-9; insgesamt 100 Messungen. Gemessen wurde eine kanonische Neuimplementierung mit vollständigen Randbits — die frühere Referenzimplementierung speicherte nur das erste davon und konnte daher nicht mehr als eine Schicht dekodieren.
**Ergebnis.** Kompressionsverhältnis mit Endkodierer zlib-9 (gesamter Container) sowie die Referenzverfahren auf den Rohdaten:

| Datei | L = 0 (Kontrolle) | L = 1 | L = 2 | L = 4 | L = 8 | zlib-9 | zstd-19 |
|---|---|---|---|---|---|---|---|
| technischer Text | 0,2929 | 0,2972 | 0,3083 | 0,3261 | 0,3541 | 0,2903 | 0,2778 |
| synthetisches Formularraster | 0,1411 | 0,1439 | 0,1581 | 0,1714 | 0,1720 | 0,1411 | 0,1069 |
| Telemetrie (CSV) | 0,1623 | 0,1624 | 0,1648 | 0,1694 | 0,2050 | 0,1622 | 0,0740 |
| dünn besetzte Datei | 0,0028 | 0,0029 | 0,0030 | 0,0033 | 0,0034 | 0,0026 | 0,0014 |
| lange Bitläufe | 0,0067 | 0,0065 | 0,0065 | 0,0065 | 0,0065 | 0,0065 | 0,0094 |
| Zufallsdaten | 1,0005 | 1,0005 | 1,0005 | 1,0005 | 1,0005 | 1,0003 | 1,0001 |
| Ausschnitt der ausführbaren Datei | 0,8641 | 0,8701 | 0,8762 | 0,8838 | 0,8937 | 0,8640 | 0,8509 |
| nur Nullen | 0,0012 | 0,0012 | 0,0012 | 0,0012 | 0,0012 | 0,0011 | 0,0001 |
| alternierende Bits | 0,0012 | 0,0012 | 0,0012 | 0,0012 | 0,0012 | 0,0011 | 0,0001 |
| periodische Datei | 0,0034 | 0,0035 | 0,0036 | 0,0034 | 0,0028 | 0,0033 | 0,0004 |

Mit zlib-9 verschlechtert sich das Ergebnis auf fünf Dateien monoton mit der Schichtzahl: technischer Text, synthetisches Formularraster, Telemetrie, dünn besetzte Datei und Ausschnitt der ausführbaren Datei. Die Ausnahmen sind ausschließlich synthetisch. Auf der Datei mit langen Läufen spart eine Schicht 48 B gegenüber demselben Container ohne Transformation (1707 statt 1755 B) und erreicht damit genau das Ergebnis von zlib-9 allein (1707 B). Auf der periodischen Datei sparen acht Schichten 175 B (725 statt 900 B), während zstd-19 dieselbe Datei auf 97 B bringt. Auf den Zufallsdaten und in den beiden entarteten Fällen (nur Nullen, alternierende Bits) bleibt das Ergebnis bis auf 1 B konstant. Mit RLE verwandelte die Transformation eine Expansion nur auf einer Datei in Kompression — der Datei mit alternierenden Bits (16,00 → 0,0629 bei einer Schicht), die zlib-9 ganz ohne Transformation auf 0,0011 bringt. Wo die Schichten das RLE-Ergebnis auf anderen Dateien verkleinerten, blieb es größer als die Eingabe (z. B. Formularraster 1,51, Telemetrie 5,74, technischer Text 6,33).
**Auswertung der Kriterien.** F1: erfüllt, 100/100. F2: mit zlib-9 nur auf zwei synthetischen Dateien erfüllt (lange Läufe, periodische Datei), mit RLE nur dort, wo das Ergebnis weiterhin größer als die Eingabe ist, sowie auf den alternierenden Bits. F3: auf 10 von 10 Dateien nicht erfüllt — kein Ergebnis mit Transformation war zugleich kleiner als zlib-9 und als zstd-19. **Hypothese widerlegt.**
**Mechanismus (gemessen).** Auf Text und Telemetrie treiben die ersten Schichten die Dichte der Einsen gegen 0,5: technischer Text 0,3932 → 0,4607 (L = 1) → 0,4805 (L = 2), Telemetrie 0,4261 → 0,5119 (L = 1). Die Transformation verwischt also die Struktur, die der Endkodierer sieht, statt sie freizulegen. Wo die Dichte sinkt, verbessert sich das Ergebnis trotzdem nicht: Auf dem synthetischen Formularraster senkt eine Schicht die Dichte von 0,1918 auf 0,1365, das Kompressionsverhältnis mit zlib-9 steigt jedoch von 0,1411 auf 0,1439. Dasselbe Muster wiederholt sich in 2D (Abschnitt 6).
**Nächster Schritt.** Keiner — Achse geschlossen (Abschnitt 7).
**Status:** durch Messung falsifiziert (F3 auf 10/10 Dateien nicht erfüllt; F1 erfüllt, 100/100); Achse geschlossen.

### 2. Richtung 1 — die vollständige Schichtpyramide: umkehrbar, aber quadratisch
**Forschungsfrage.** Ist die historische Implementierung, die die vollständige Schichtpyramide bis hinunter zu einem einzigen Bit aufbaut, durchführbar?
**Kriterien.** Wie in Abschnitt 1 (F1–F3).
**Methode.** Historische Implementierung ohne Änderungen, angewandt auf Präfixe des technischen Textes von 128 bis 1024 B.
**Ergebnis.**

| Eingabe | Ausgabe | Ausgabe / Eingabe | Kodierzeit |
|---|---|---|---|
| 128 B | 515.415 B | ×4027 | 0,086 s |
| 256 B | 2.055.749 B | ×8030 | 0,257 s |
| 512 B | 8.285.223 B | ×16.182 | 0,969 s |
| 1024 B | 33.310.203 B | ×32.529 | 3,844 s |

Die Pyramide für n Eingabebits enthält n(n−1)/2 Abhängigkeitsbits; die gemessene Ausgabe entspricht n(n−1)/2 Bytes auf 2 % genau (98,1–99,3 %), also etwa einem Byte pro Abhängigkeitsbit. Größe und Zeit wachsen quadratisch. Hochgerechnet auf die 14,2 MB große ausführbare Datei aus dem Eintrag vom 2026-07-14: etwa 6,5 PB Ausgabe und etwa 23 Jahre Rechenzeit auf der Messmaschine. F1 erfüllt (4/4); F2 und F3 widerlegt — das Ergebnis ist etwa 4000- bis 32.500-mal so groß wie die Eingabe. Beobachtung zur Konstruktion: Der Dekodierer dieser Implementierung liest ausschließlich die erste Schicht (mit einem Randbit genügt sie zur Wiederherstellung); das Speichern der Schichten ab der zweiten ist damit von vornherein redundant.
**Nächster Schritt.** Keiner — Implementierung zurückgezogen.
**Status:** umkehrbar (4/4), jenseits von Spielzeuggrößen undurchführbar; zurückgezogen.

### 3. Richtung 2 — der PGA-Codec in der historischen Version: Wiederherstellung auf 10 von 10 Dateien gescheitert
**Forschungsfrage.** Stellt die Codec-Version, an der die im Eintrag vom 2026-07-14 beschriebene Expansion um 56 % gemessen wurde, die Daten bitgenau wieder her?
**Kriterien (vor der Messung festgelegt).** F1 für den PGA-Codec sowie F4: Größe des gesamten Containers (alle Abschnitte und Metadaten) kleiner als die Eingabe auf dem strukturierten Korpus, bei erfülltem F1.
**Methode.** Historische Version ohne Änderungen; Kodierung und Dekodierung jeder der 10 Korpusdateien; SHA-256 des Ergebnisses gegen das Original.
**Ergebnis: Wiederherstellungskriterium widerlegt.** Die Kodierung lief auf 10 von 10 Dateien durch, die Dekodierung auf 0 von 10: Auf 4 Dateien brach sie mit einem Fehler ab, auf 6 lieferte sie Daten, deren SHA-256 nicht mit dem Original übereinstimmte. Die Code-Analyse ergab zwei Formatfehler: Die Restschicht trug nicht die Daten, die sie wiederherstellen sollte, und der Kartenabschnitt war nicht eindeutig parsbar. Die Folge reicht über den Benchmark hinaus: Auch das historische 22,2-MB-Artefakt aus dem Eintrag vom 2026-07-14 ist nicht vollständig rekonstruierbar. Seine Größe wurde korrekt gemessen (14,2 → 22,2 MB), eine gültige verlustfreie Kodierung war es jedoch nicht. F4 lässt sich nicht bewerten, solange F1 nicht erfüllt ist.
**Nächster Schritt.** Behebung beider Fehler in einer Messvariante (Abschnitt 4).
**Status:** negatives Ergebnis — die historische Version erfüllt das Wiederherstellungskriterium nicht (0/10); Korrektur des Eintrags vom 2026-07-14 siehe unten.

### 4. Richtung 2 — nach der Reparatur: ein bedingter Positivbefund, der sich auf realistische Daten nicht übertrug
**Forschungsfrage.** Stellt der Codec nach der Formatreparatur die Daten bitgenau wieder her, und liefert er eine positive Größenbilanz — zuerst auf dem Korpus von Benchmark 1, dann auf realistischen Bildern?
**Stand der Technik (veröffentlicht).** Für Binärbilder: CCITT G4 (ITU-T T.6, zweidimensionale Lauflängenkodierung relativ zur Referenzzeile) und JBIG2 (ITU-T T.88), das Text selbst über Symbolwörterbücher mit Mustervergleich kodiert. Diese, nicht gzip und zstd, sind die eigentliche Konkurrenz für diese Datenklasse. G4 wurde in Benchmark 2 gemessen; JBIG2 nicht — in der Messumgebung stand kein Kodierer dafür zur Verfügung.
**Kriterien (vor der Messung festgelegt).** F4 (Benchmark 1): Container kleiner als die Eingabe auf dem strukturierten Korpus, bei erfülltem F1. F5 (Benchmark 2): Kompressionsverhältnis kleiner als zstd-19 **und** als CCITT G4 auf dem realistischen Korpus.
**Methode.** Messvariante mit Behebung beider Fehler, in zwei Ausprägungen: formattreu zur historischen Version sowie mit einem von Redundanz befreiten Kartenabschnitt (Konstruktion zurückgehalten). Benchmark 2 verwendet die zweite Ausprägung auf sieben Binärbildern (Abschnitt 8), mit der tatsächlichen Rastergeometrie.
**Ergebnis — Benchmark 1.** F1 wiederhergestellt: 10/10 in beiden Ausprägungen. F4 erwies sich in der Formulierung des Protokolls als schwach: Einen Container kleiner als die Eingabe hat die formattreue Ausprägung auf 4 von 5 strukturierten Dateien (auf dem technischen Text 1,098), die Ausprägung mit redundanzfreier Karte auf 5 von 5. Entscheidend ist der Vergleich mit den Referenzverfahren. Die formattreue Ausprägung war auf jeder der 10 Dateien größer als zstd-19; auf dem synthetischen Formularraster belegte die Karte 97 % ihres Containers. Die Ausprägung mit redundanzfreier Karte lieferte den ersten gemessenen Positivbefund des Programms — auf einer Datei: synthetisches Formularraster 0,0495 gegenüber 0,1069 (zstd-19) und 0,1411 (zlib-9), also 2,2-mal besser als das beste Referenzverfahren. Auf den übrigen 9 Dateien gewannen die Referenzverfahren (z. B. Telemetrie 0,2981 gegenüber 0,0740, technischer Text 0,7949 gegenüber 0,2778). Auf hochentropischen Daten expandiert der Codec: Zufallsdaten 1,176 (formattreue Ausprägung: 1,547), Ausschnitt der ausführbaren Datei 1,115 (1,519). Stand nach Benchmark 1: bedingter Positivbefund — synthetisches Raster, Glyphen am Kachelraster ausgerichtet, ohne Rauschen, ohne Vergleich mit G4 und JBIG2.
**Ergebnis — Benchmark 2: F5 widerlegt.**

| Bild | PGA | zlib-9 | zstd-19 | CCITT G4 | bestes Verfahren |
|---|---|---|---|---|---|
| `fax2d` — CCITT-Fax, 1728×1082 | 0,1888 | 0,1374 | 0,1241 | **0,1209** | G4 |
| `g3test` — CCITT-Fax, 1728×1103 | 0,2629 | 0,1806 | **0,1624** | 0,1756 | zstd-19 |
| `jim___ah` — dichtes Testbild, 664×813 | 0,2443 | 0,2083 | **0,1910** | 0,6301 | zstd-19 |
| Textseite, serifenlose Schrift | 0,0489 | 0,0440 | **0,0364** | 0,0376 | zstd-19 |
| Textseite, Serifenschrift | 0,0348 | 0,0376 | **0,0302** | 0,0346 | zstd-19 |
| Formularseite | 0,0351 | 0,0263 | **0,0133** | 0,0372 | zstd-19 |
| synthetisches Formularraster (Kontrolle aus Benchmark 1) | **0,0492** | 0,1411 | 0,1069 | 0,4030 | PGA |

Der Codec unterliegt zstd-19 auf 6 von 6 realistischen Bildern und CCITT G4 auf 4 von 6. Das einzige Bild, auf dem irgendeine eigene Methode des Programms am besten abschneidet, bleibt das synthetische, am Kachelraster ausgerichtete Raster. F1: 7/7. Methodischer Hinweis: In Benchmark 1 rasterte der Codec die Eingabe ohne Kenntnis ihrer tatsächlichen Geometrie; dass die Rasterbreite mit der tatsächlichen Breite des synthetischen Formularrasters übereinstimmte, war Zufall. Benchmark 2 verwendet in allen Messungen die tatsächliche Geometrie.
**Nächster Schritt.** Klären, wovon der einzige Positivbefund abhängt (Abschnitt 5).
**Status:** Umkehrbarkeit wiederhergestellt (10/10, 7/7); Generalisierung widerlegt (F5); Positivbefund auf das synthetische Raster beschränkt.

### 5. Richtung 2 — Mechanismus des Positivbefunds: starre, bitgenaue Wiederholung
**Forschungsfrage.** Wovon hängt der einzige Positivbefund ab — von der Ausrichtung der Muster am Kachelraster, von ihrer Bitgenauigkeit oder von beidem?
**Kriterien.** Ein in Protokoll 1 angekündigter Empfindlichkeitstest ohne Erfolgsschwelle (gemessen wird die Degradationskurve gegenüber G4 und zstd-19 auf denselben Daten); F7 (vor der Messung festgelegt): Die Wahl des Rasterversatzes stellt das Kompressionsverhältnis des ausgerichteten Rasters wieder her.
**Methode.** Das synthetische Formularraster wurde horizontal um weniger als eine Kachelbreite verschoben und durch zufälliges Kippen von Bits mit einer Wahrscheinlichkeit von 0,1 bis 2 % verrauscht; jedes Ergebnis mit SHA-256-Prüfung (13/13).
**Ergebnis — Verschiebung.** Je nach Verschiebung ein Kompressionsverhältnis von 0,0492 bis 0,0880, das Wörterbuch mit bis zu 46 statt 16 Einträgen; G4 unverändert (0,403).
**Ergebnis — Rauschen.**

| gekippte Bits | PGA | Wörterbucheinträge | zstd-19 | CCITT G4 |
|---|---|---|---|---|
| 0 | 0,0492 | 16 | 0,1069 | 0,4030 |
| 0,1 % | 0,0713 | 20 | 0,1327 | 0,4158 |
| 0,5 % | 0,1290 | 78 | 0,1969 | 0,4652 |
| 1 % | 0,1991 | 561 | 0,2555 | 0,5195 |
| 2 % | 0,3851 | 2656 | 0,3434 | 0,6180 |

Der Vorsprung vor zstd-19 schrumpft von 2,2-fach auf sauberen Daten auf 1,9-, 1,5- und 1,3-fach und kehrt sich bei 2 % gekippten Bits um (0,3851 gegenüber 0,3434). Das Kompressionsverhältnis des Codecs verschlechtert sich um den Faktor 7,8, das von G4 um den Faktor 1,5. Schluss: Der Positivbefund verlangt eine starre, bitgenaue Wiederholung der Muster — wie sie digital erzeugte Raster (born-digital) aufweisen, Scans dagegen nicht.
**Ergebnis — F7: technisch bestätigt.** Absichtlich gegen das Kachelraster verschobenes Raster: 0,1416; nach Wahl des Rasterversatzes: 0,0492, also genau der Wert des ausgerichteten Rasters; Kosten: wenige Bits; F1 erfüllt. Das gilt für eine Verschiebung des ganzen Rasters; auf gerenderten Seiten und Scans, wo sich die Position der Zeichen einzeln ändert, wurde es nicht gemessen.
**Konstruktion: zurückgehalten.**
**Nächster Schritt.** Die Hypothese der dokumentübergreifenden Deduplikation (Abschnitt 7).
**Status:** Randbedingungen gemessen; F7 technisch bestätigt.

### 6. Richtung 3 — räumliche Abhängigkeit in 2D: Hypothese widerlegt, mit Ablation der Prädiktoren
**Forschungsfrage.** Trägt die Abhängigkeitsregel, von 1D auf räumliche Nachbarn im Raster übertragen, ein Signal, das in 1D fehlte?
**Stand der Technik (veröffentlicht).** Kontextmodellierung von Binärbildern in JBIG und JBIG2: Kontextschablone aus 10–16 Nachbarpixeln und adaptiver arithmetischer Kodierer. Hier wurde bewusst eine Minimalvariante gemessen, um zu entscheiden, ob die Achse überhaupt ein Signal trägt.
**Kriterien (vor der Messung festgelegt).** F6(a): Der Mehrheitsprädiktor (Mehrheit aus drei Nachbarn: links, oben, oben links) ist besser als die trivialen Prädiktoren (nur linker oder nur oberer Nachbar). F6(b): Das Ergebnis ist kleiner als das beste Referenzverfahren für das jeweilige Bild.
**Methode.** Residuum = Pixel ⊕ Vorhersage; kausaler Dekodierer; Endkodierer zlib-9 oder zstd-19; 7 Bilder × 4 Varianten = 28 Messungen; F1: 28/28.
**Ergebnis.**

| Bild | links + zlib-9 | oben + zlib-9 | Mehrheit + zlib-9 | Mehrheit + zstd-19 | bestes Referenzverfahren |
|---|---|---|---|---|---|
| `fax2d` | 0,1359 | 0,1646 | 0,1691 | 0,1567 | 0,1209 (G4) |
| `g3test` | 0,1800 | 0,2178 | 0,2355 | 0,2154 | 0,1624 (zstd-19) |
| `jim___ah` | 0,2085 | 0,3058 | 0,2849 | 0,2642 | 0,1910 (zstd-19) |
| Textseite, serifenlose Schrift | 0,0450 | 0,0456 | 0,0554 | 0,0469 | 0,0364 (zstd-19) |
| Textseite, Serifenschrift | 0,0387 | 0,0409 | 0,0465 | 0,0376 | 0,0302 (zstd-19) |
| Formularseite | 0,0272 | 0,0264 | 0,0316 | 0,0165 | 0,0133 (zstd-19) |
| synthetisches Formularraster | 0,1412 | 0,1610 | 0,2447 | 0,2189 | 0,1069 (zstd-19) |

F6(a) widerlegt: Mit demselben Kodierer (zlib-9) war der Mehrheitsprädiktor auf 7 von 7 Bildern schlechter als der linke Nachbar und auf 6 von 7 schlechter als beide trivialen Prädiktoren (Ausnahme: Auf `jim___ah` schlug er den oberen Nachbarn). F6(b) widerlegt: Keine Variante war auf irgendeinem Bild besser als das beste Referenzverfahren (0/7). Die jeweils beste Variante war auf 3 von 7 Bildern besser als G4 allein (`jim___ah`, Formularseite, synthetisches Raster), auf keinem besser als zstd-19 auf den Rohdaten.
**Mechanismus.** Das Residuum ist dünner besetzt als das Original und komprimiert dennoch nicht besser. Auf `jim___ah` hat es 0,165 Einsen pro Pixel gegenüber 0,557 im Original und ergibt mit zlib-9 praktisch dasselbe Ergebnis (0,2085 gegenüber 0,2083). Mit demselben Kodierer war das Residuum des linken Nachbarn höchstens geringfügig kleiner als das Original (`fax2d`: 0,1359 gegenüber 0,1374). Eine mit beiden Messungen vereinbare Deutung: Ein Kodierer der LZ-Familie gewinnt aus Wiederholungen — identischen Glyphen und Zeilen — und nicht aus der Bitdichte; die Abhängigkeitstransformation verändert die Dichte (in 1D in Richtung 0,5, in 2D nach unten), fügt aber keine Wiederholungen hinzu und zerstört einen Teil der vorhandenen. Der Gewinn aus der dünneren Besetzung gleicht den Verlust an Übereinstimmungen nicht aus.
**Nächster Schritt.** Keiner in dieser Form: Weiterarbeit hieße Kontextmodellierung mit arithmetischem Kodierer, also Konvergenz zum Stand der Technik (JBIG/JBIG2), keine neue Achse.
**Status:** falsifiziert (F6a, F6b); Achse geschlossen.

### 7. Schluss des Programms: Abhängigkeitsachse geschlossen, Nische eingegrenzt
- **Die Achse der Abhängigkeitskodierung ist als falsifiziert geschlossen**, in 1D (Abschnitte 1–2) und in 2D (Abschnitt 6), mit Endkodierern der Klassen RLE und LZ. Festgehaltener Mechanismus: Die Transformation verändert die Bitdichte, fügt aber keine der Wiederholungen hinzu, auf denen der Kodierer aufbaut.
- **PGA-Codec — Nische eingegrenzt.** Das Programm ging vom „besseren Allzweck-Kompressor" (in den Iterationen 2021–2025 widerlegt) über „strukturierte Korpora" (Dokumente, Scans, Telemetrie — für Scans, gerenderte Seiten, Text und Telemetrie durch Messung widerlegt) zu einer engen Datenklasse: digital erzeugte Raster mit starrer, bitgenauer Wiederholung (programmgenerierte Formulare, Druckdaten aus dem Spooler, aus Vorlagen befüllte Dokumente). Dort war der Codec etwa 2,2-mal besser als das beste klassische Referenzverfahren — gemessen mit vollständiger Wiederherstellung — und etwa 2,6-mal nach einer weiteren Verbesserung des Containers; das ist eine Schätzung aus Abschnittsgrößen, noch nicht durchgängig gemessen. Konstruktion: zurückgehalten.
- **Randbedingungen der Nische, offen benannt.** Der Vorsprung wurde auf einem einzigen synthetischen Raster gemessen; verglichen wurde mit G4, zlib-9 und zstd-19, nicht mit JBIG2, das Symbolwörterbücher verwendet und den nächstliegenden Stand der Technik bildet; die Robustheit gegenüber Rauschen ist schwach (Abschnitt 5). Das ist eine Abgrenzungshypothese, kein nachgewiesener Vorsprung auf realen Daten.
- **Nächste Hypothese: dokumentübergreifende Deduplikation.** Ein gemeinsames Wörterbuch für eine Serie von Seiten aus derselben Vorlage (z. B. Serien von Rechnungen oder Protokollen); Erwartung: Die Kosten des Wörterbuchs amortisieren sich über die Serie, und das Kompressionsverhältnis der Serie liegt deutlich unter dem einer einzelnen Seite. JBIG2 sieht Symbolwörterbücher vor, die mehreren Seiten gemeinsam sind; ohne Vergleich damit ist das Ergebnis dieser Hypothese nicht aussagekräftig. Das Kriterium wird, wie in beiden Sitzungen, vor der Messung festgelegt.

### 8. Korpus, Referenzverfahren, Reproduzierbarkeit
**Benchmark 1 (2026-07-15), 10 Dateien.** Fünf strukturierte: technischer Text (Projektdokumentation, 18.309 B), synthetisches Formularraster mit 1 Bit/Pixel (512 KiB; Tabellenlinien und ein kleiner fester Satz von Glyphen auf dem Kachelraster, ohne Rauschen), synthetische Telemetrie im CSV-Format (418.929 B), dünn besetzte Datei (256 KiB; Nullen und ein 16-Byte-Datensatzkopf alle 4096 B), Datei mit langen Läufen von 0x00- und 0xFF-Bytes (256 KiB). Zwei hochentropische Kontrolldateien: pseudozufällige Daten (256 KiB) und das erste MiB der kompilierten ausführbaren Windows-Datei aus dem Eintrag vom 2026-07-14. Drei Grenzfälle zu je 256 KiB: nur Nullen, alternierende Bits (Byte 0xAA), periodische Folge (Bytes 0–63). Die synthetischen Dateien sind deterministisch (fester Startwert des Zufallsgenerators).
**Benchmark 2 (2026-07-17), 7 Binärbilder.** Drei echte Bilder aus dem öffentlichen Testbildarchiv von libtiff `pics-3.8.0` (download.osgeo.org/libtiff/pics-3.8.0.tar.gz): `fax2d` (1728×1082) und `g3test` (1728×1103) — klassische CCITT-Fax-Testbilder — sowie `jim___ah` (664×813; dicht: 56 % schwarze Pixel, keine einzige leere Zeile). Drei mit TrueType-Systemschriften bei etwa 200 dpi gerenderte Seiten (1656×2336): Text in serifenloser Schrift, Text in Serifenschrift sowie ein Formular mit Linien und Beschriftungen in nichtproportionaler Schrift; Zeilen- und Zeichenpositionen absichtlich nicht am Kachelraster ausgerichtet. Kontrolle: das synthetische Formularraster aus Benchmark 1. Die Raster sind als rohe Bits gespeichert, 1 Bit/Pixel, zeilenweise, Schwarz = 1.
**Referenzverfahren.** zlib-9 (DEFLATE, wie in gzip), zstd-19, CCITT G4 (in Benchmark 2; gemessen als vollständige TIFF-Datei). JBIG2 — nicht gemessen: kein Kodierer in der Messumgebung.
**Umgebung.** Windows 11 (AMD64), Python 3.12.10, numpy 2.5.0, Pillow 10.4.0, zstandard 0.25.0.
**Prüfung.** Benchmark 1: 154 Messungen; bei 144 Dekodierungen stimmte der SHA-256-Hashwert mit dem des Originals überein, die 10 Abweichungen betreffen die historische Version des PGA-Codecs (Abschnitt 3). Benchmark 2: 51 Messungen, darunter 48 Dekodierungen — alle 48 übereinstimmend; die übrigen 3 Messungen betrafen ein Konstruktionsdetail des Containers, ihr Ergebnis ist zurückgehalten.
**Reproduzierbarkeit.** Die Ergebnisse der Referenzverfahren auf den libtiff-Bildern kann jeder wiederholen: Das Archiv ist öffentlich, Geometrie und Rasterformat sind oben angegeben. Die Ergebnisse der eigenen Methoden erfordern die zurückgehaltene Konstruktion; mit den Protokollen verbinden sie die Hashwerte in Abschnitt 9.

### 9. Nachweise: SHA-256-Hashwerte und Zeitstempel
Die Messprotokolle bleiben privat: Sie enthalten die zurückgehaltene Konstruktion. Ihre hier veröffentlichten SHA-256-Hashwerte erlauben später die Prüfung, dass ein künftig vorgelegtes Dokument dasselbe ist, das im Juli 2026 gestempelt wurde. Die Nachweisverzeichnisse enthalten die SHA-256-Hashwerte der Rohergebnisdateien und der Korpusmanifeste sowie die Commit-Kennungen des Messcodes (das erste zusätzlich den Hashwert des Whitepapers von 2025).

| Dokument | SHA-256 | OpenTimestamps-Zeitstempel |
|---|---|---|
| Messprotokoll Nr. 1 (Benchmark 1) | `0830cc94c042c05103e45a47de2bdae0db2411f683fcd309e6630c5700552375` | 2026-07-15, Bitcoin-Attestierung |
| Nachweisverzeichnis zu Benchmark 1 | `8b85e99708febb32682a100f5b2336bd0ffc8cf0e0ef4c7671aa91576de6dc56` | 2026-07-15, Bitcoin-Attestierung |
| Messprotokoll Nr. 2 (Benchmark 2) | `c478b509fda7419a60b0ad2149d28f4dec83aae5fb074ffbeab0cd3bb3aba1de` | 2026-07-17 |
| Nachweisverzeichnis zu Benchmark 2 | `7b2a7a3fa8a49c3adda2e6604cdf462106f1ab4fc11a26c4246ba8e955324b09` | 2026-07-17 |

Die Hashwerte beider Protokolle stimmen mit den Einträgen in den Verzeichnissen überein. Wo ein Protokoll eine Zahl oder Beschreibung enthält, die nicht zu den Rohergebnisdateien passt, gibt dieser Eintrag den Wert aus den Ergebnisdateien an: (1) Protokoll Nr. 1 sagt, dass sich das Ergebnis mit zlib-9 auf jeder Datei monoton mit der Schichtzahl verschlechtert — laut Daten gilt das für 5 von 10 Dateien (Abschnitt 1); (2) Protokoll Nr. 1 bewertet F4 gegenüber zstd-19, obwohl F4 gegenüber der Eingabegröße formuliert ist — Abschnitt 4 gibt beide Lesarten an; (3) Protokoll Nr. 2 nennt 30/30 geprüfte Dekodierungen im 2D-Test — die Ergebnisdatei enthält 28, die Sitzungssumme 48/48 ist richtig; (4) Protokoll Nr. 2 gibt an, der Mehrheitsprädiktor sei auf 6 von 7 Bildern schlechter als der linke Nachbar gewesen — mit demselben Kodierer war er es auf 7 von 7, schlechter als beide trivialen Prädiktoren auf 6 von 7; (5) Protokoll Nr. 2 beschreibt `jim___ah` als handschriftlichen Scan — 56 % schwarze Pixel und keine leere Zeile sprechen eher für ein dichtes oder gerastertes Bild; wir beschreiben es daher neutral.

### Korrekturen zum Eintrag vom 2026-07-14
- **Richtung 1, Kriterium (b).** Eintrag: „gemessene Ratio auf definierten Korpora gegenüber gzip/zstd — das Messprogramm wurde noch nicht durchgeführt". Jetzt: am 2026-07-15 gemessen und widerlegt (Abschnitt 1); die Achse ist geschlossen, auch in der 2D-Variante (Abschnitt 6).
- **Richtung 1, Konstruktion.** Eintrag: „Konstruktion (Abhängigkeitsregel, Schichtaufbau, Containerformat): zurückgehalten". Regel und Schichtaufbau sind jetzt offengelegt — XOR benachbarter Bits, Schicht für Schicht iteriert —, weil die Achse falsifiziert und geschlossen ist.
- **Richtung 1, Umkehrbarkeit.** Die Aussage bleibt bestehen (100/100 und 4/4), mit einer Präzisierung: Eine frühere Referenzimplementierung speicherte nur das erste Randbit und konnte nicht mehr als eine Schicht dekodieren; gemessen wurde an einer korrigierten Neuimplementierung.
- **Richtung 2, Wiederherstellungskriterium.** Eintrag: „Bitgenaue Rekonstruktion (erfüllt, hashverifiziert)". Für den PGA-Codec traf das nicht zu: Die Prüfung betraf einen früheren Machbarkeitsnachweis. Die erste vollständige Validierung des PGA-Codecs (2026-07-15) ergab 0/10; umkehrbar ist erst die reparierte Variante (Abschnitte 3–4).
- **Negatives Ergebnis +56 %.** Die Größe wurde korrekt gemessen (14,2 → 22,2 MB), das Artefakt ist aber nicht vollständig rekonstruierbar und war damit keine gültige verlustfreie Kodierung. Der Schluss — Expansion auf hochentropischen Daten — ist durch die reparierte Variante mit SHA-256-Prüfung bestätigt: auf dem ersten MiB derselben Datei +52 % (formattreue Ausprägung) und +12 % (redundanzfreie Karte), auf Zufallsdaten +55 % und +18 %.
- **Richtung 2, Stand der Technik.** Der Eintrag verglich mit gzip/zstd. Der eigentliche Stand der Technik für Binärbilder sind CCITT G4 und JBIG2; G4 wurde in Benchmark 2 gemessen, JBIG2 noch nicht.
- **Richtung 2, Status.** Eintrag: „POC implementiert / in Validierung". Jetzt: validiert — die historische Version erfüllt das Wiederherstellungskriterium nicht, die reparierte Variante erfüllt es, der Positivbefund überträgt sich nicht über das synthetische Raster hinaus (Abschnitte 3–5).
- **Abgrenzungsschluss.** Der Eintrag verengte das Programm auf strukturierte Korpora (Dokumente, Scans, Telemetrie). Die Messungen verengen weiter: Auf Scans, gerenderten Textseiten, Telemetrie und Text gewinnen die klassischen Kodierer; es bleibt die Nische digital erzeugter Raster mit starrer Wiederholung (Abschnitt 7).

**Programmstatus:** aktiv. Achse der Abhängigkeitskodierung — als falsifiziert geschlossen, in 1D und 2D. PGA-Codec — in der reparierten Variante umkehrbar, mit einem Vorsprung, der nur auf einem einzigen synthetischen, digital erzeugten Raster gemessen ist; nächster Schritt: dokumentübergreifende Deduplikation auf Seitenserien aus einer Vorlage.

---
---

# 🇬🇧 ENGLISH VERSION

## Koch Laboratory — structure-first compression, measured: the dependency axis closed, the niche narrowed

The entry of 2026-07-14 ended by announcing a defined benchmark with a disclosed corpus and reference baselines, because the compression ratio of neither direction had been measured. This entry reports two measurement sessions: Benchmark 1 (2026-07-15; 10 files, 154 measurements) and Benchmark 2 (2026-07-17; realistic images, 51 measurements). The falsifiable criteria were set before each session, and every decode was checked against the original with SHA-256. Statements of the entry of 2026-07-14 that the measurements overturned are collected under "Corrections to the entry of 2026-07-14".

The balance in three sentences. The dependency-coding axis is refuted in 1D and in 2D and is closed. The historical version of the PGA codec could not restore the data it had itself encoded. After repair, the codec beats the classical baselines only on a born-digital raster whose repetitions are rigid and bit-exact — on real scans and rendered pages it loses.

**Method.** Problem → state of the art → criterion set before measurement → method → result with boundary conditions → next step. Compression ratio = output size / input size: lower is better, above 1 means expansion. The whole container counts — headers, metadata, boundary bits, checksum — not the payload alone. Criterion labels (F1–F7) follow the private measurement protocols, whose hashes are given in section 9.

> **Publication note.** For directions still open we publish the problem, the state of the art, the criterion and the result — not the construction. Where "Construction: withheld." appears, the technical detail is retained as filing material. The dependency-coding axis is closed as falsified, so its rule is disclosed here.

### 1. Direction 1 — dependency coding in 1D: rule disclosed, hypothesis refuted
**Research question.** Does a reversible representation of a stream as layers of relations between neighbouring bits yield, after a classical final coder, an output smaller than the same coder without the transform, and smaller than zlib-9 and zstd-19?
**Rule (disclosed, since the axis is closed).** The first layer is the XOR of every pair of neighbouring bits, `b[i] ⊕ b[i+1]`; each further layer is built from the previous one by the same rule. Inverting L layers requires L boundary bits.
**Analytic note.** For the measured depths L = 1, 2, 4, 8 the iteration yields exactly `b[i] ⊕ b[i+L]` (binomial coefficients modulo 2): layer L is an XOR difference at distance L, and eight layers are the XOR difference of neighbouring bytes. At the measured depths the transform therefore lies within classical delta coding.
**State of the art (published).** Delta and predictive coding (including XOR delta), RLE, DEFLATE (zlib, gzip), zstd.
**Criteria (set before measurement).** F1: the SHA-256 of the reconstruction equals that of the original on every decode. F2: layer L plus the final coder yields a smaller output than the same container without the transform (L = 0), for L ∈ {1, 2, 4, 8}. F3: the output is smaller than zlib-9 and than zstd-19.
**Method.** A corpus of 10 files (section 8); L ∈ {0, 1, 2, 4, 8}; final coder RLE or zlib-9; 100 measurements in total. The measurements used a canonical re-implementation that stores all boundary bits — the earlier reference implementation stored only the first of them and so could not decode more than one layer.
**Result.** Compression ratio with zlib-9 as the final coder (whole container), and the baselines on the raw data:

| file | L = 0 (control) | L = 1 | L = 2 | L = 4 | L = 8 | zlib-9 | zstd-19 |
|---|---|---|---|---|---|---|---|
| technical text | 0.2929 | 0.2972 | 0.3083 | 0.3261 | 0.3541 | 0.2903 | 0.2778 |
| synthetic form raster | 0.1411 | 0.1439 | 0.1581 | 0.1714 | 0.1720 | 0.1411 | 0.1069 |
| telemetry (CSV) | 0.1623 | 0.1624 | 0.1648 | 0.1694 | 0.2050 | 0.1622 | 0.0740 |
| sparse file | 0.0028 | 0.0029 | 0.0030 | 0.0033 | 0.0034 | 0.0026 | 0.0014 |
| long bit runs | 0.0067 | 0.0065 | 0.0065 | 0.0065 | 0.0065 | 0.0065 | 0.0094 |
| random data | 1.0005 | 1.0005 | 1.0005 | 1.0005 | 1.0005 | 1.0003 | 1.0001 |
| executable slice | 0.8641 | 0.8701 | 0.8762 | 0.8838 | 0.8937 | 0.8640 | 0.8509 |
| all zeros | 0.0012 | 0.0012 | 0.0012 | 0.0012 | 0.0012 | 0.0011 | 0.0001 |
| alternating bits | 0.0012 | 0.0012 | 0.0012 | 0.0012 | 0.0012 | 0.0011 | 0.0001 |
| periodic file | 0.0034 | 0.0035 | 0.0036 | 0.0034 | 0.0028 | 0.0033 | 0.0004 |

With zlib-9 the result worsens monotonically with the number of layers on five files: the technical text, the synthetic form raster, the telemetry, the sparse file and the executable slice. The exceptions are all synthetic. On the long-run file one layer saves 48 B against the same container without the transform (1,707 vs 1,755 B), which merely equals zlib-9 alone (1,707 B). On the periodic file eight layers save 175 B (725 vs 900 B), while zstd-19 codes the same file to 97 B. On random data and in the two degenerate cases (all zeros, alternating bits) the output is constant to within 1 B. With RLE, the transform turned expansion into compression on one file only — the alternating-bit file (16.00 → 0.0629 at one layer), which zlib-9 codes to 0.0011 with no transform at all. Where layers reduced the RLE output on other files, it stayed larger than the input (e.g. form raster 1.51, telemetry 5.74, technical text 6.33).
**Criteria outcome.** F1: met, 100/100. F2: with zlib-9 met only on two synthetic files (long runs, periodic file); with RLE only where the output is still larger than the input, and on the alternating bits. F3: unmet on 10 of 10 files — no output with the transform was smaller than both zlib-9 and zstd-19. **Hypothesis refuted.**
**Mechanism (measured).** On the text and the telemetry the first layers push the density of ones towards 0.5: technical text 0.3932 → 0.4607 (L = 1) → 0.4805 (L = 2), telemetry 0.4261 → 0.5119 (L = 1). The transform thus blurs the structure the final coder sees instead of exposing it. Where the density falls, the output still does not improve: on the synthetic form raster one layer lowers the density from 0.1918 to 0.1365, yet the zlib-9 ratio rises from 0.1411 to 0.1439. The same pattern recurs in 2D (section 6).
**Next step.** None — the axis is closed (section 7).
**Status:** falsified by measurement (F3 unmet on 10/10 files; F1 met, 100/100); axis closed.

### 2. Direction 1 — the full layer pyramid: reversible, but quadratic
**Research question.** Is the historical implementation, which builds the full pyramid of layers down to a single bit, feasible?
**Criteria.** As in section 1 (F1–F3).
**Method.** The historical implementation, unmodified, run on prefixes of the technical text from 128 to 1,024 B.
**Result.**

| input | output | output / input | encoding time |
|---|---|---|---|
| 128 B | 515,415 B | ×4,027 | 0.086 s |
| 256 B | 2,055,749 B | ×8,030 | 0.257 s |
| 512 B | 8,285,223 B | ×16,182 | 0.969 s |
| 1,024 B | 33,310,203 B | ×32,529 | 3.844 s |

The pyramid for n input bits holds n(n−1)/2 dependency bits; the measured output tracks n(n−1)/2 bytes to within 2% (98.1–99.3%), i.e. about one byte per dependency bit. Size and time grow quadratically. Extrapolated to the 14.2 MB executable from the entry of 2026-07-14: about 6.5 PB of output and about 23 years of computation on the measurement machine. F1 met (4/4); F2 and F3 refuted — the output is roughly 4,000 to 32,500 times the size of the input. An observation on the construction: the decoder of this implementation reads the first layer only (together with one boundary bit it suffices for reconstruction), so storing the layers from the second upwards is redundant by design.
**Next step.** None — implementation withdrawn.
**Status:** reversible (4/4), infeasible beyond toy sizes; withdrawn.

### 3. Direction 2 — the PGA codec in its historical version: reconstruction failed on 10 of 10 files
**Research question.** Does the codec version on which the 56% expansion reported on 2026-07-14 was measured restore data bit-exactly?
**Criteria (set before measurement).** F1 for the PGA codec, and F4: the size of the whole container (all sections and metadata) is smaller than the input on the structured corpus, with F1 met.
**Method.** The historical version, unmodified; encoding and decoding of each of the 10 corpus files; SHA-256 of the output against the original.
**Result: reconstruction criterion refuted.** Encoding succeeded on 10 of 10 files, decoding on 0 of 10: on 4 files it aborted with an error, on 6 it produced data whose SHA-256 did not match the original. Code analysis found two format defects: the residual layer did not carry the data it was meant to restore, and the map section was not unambiguously parseable. The consequence reaches beyond the benchmark: the historical 22.2 MB artefact from the entry of 2026-07-14 is not fully reproducible either. Its size was measured correctly (14.2 → 22.2 MB), but it was not a valid lossless encoding. F4 cannot be assessed while F1 is unmet.
**Next step.** Repair of both defects in a measurement variant (section 4).
**Status:** negative result — the historical version fails the reconstruction criterion (0/10); the correction of the entry of 2026-07-14 is given below.

### 4. Direction 2 — after repair: a conditional positive that did not generalise to realistic data
**Research question.** After the format repair, does the codec restore data bit-exactly and achieve a positive size balance — first on the Benchmark 1 corpus, then on realistic images?
**State of the art (published).** For bilevel images: CCITT G4 (ITU-T T.6, two-dimensional run coding relative to a reference line) and JBIG2 (ITU-T T.88), which itself codes text through symbol dictionaries with pattern matching. These, not gzip and zstd, are the real competition for this class of data. G4 was measured in Benchmark 2; JBIG2 was not — no encoder was available in the measurement environment.
**Criteria (set before measurement).** F4 (Benchmark 1): container smaller than the input on the structured corpus, with F1 met. F5 (Benchmark 2): compression ratio smaller than zstd-19 **and** than CCITT G4 on the realistic corpus.
**Method.** A measurement variant with both defects repaired, in two forms: faithful to the historical format, and with a map section stripped of redundancy (construction withheld). Benchmark 2 uses the second form on seven bilevel images (section 8), with the true raster geometry.
**Result — Benchmark 1.** F1 restored: 10/10 in both forms. F4 as worded in the protocol proved weak: the format-faithful form has a container smaller than the input on 4 of 5 structured files (on the technical text 1.098), the form with the redundancy-free map on 5 of 5. The comparison with the baselines decides. The format-faithful form was larger than zstd-19 on every one of the 10 files; on the synthetic form raster the map took 97% of its container. The form with the redundancy-free map produced the program's first measured positive — on one file: the synthetic form raster at 0.0495 against 0.1069 (zstd-19) and 0.1411 (zlib-9), i.e. 2.2× better than the best baseline. On the other 9 files the baselines won (e.g. telemetry 0.2981 against 0.0740, technical text 0.7949 against 0.2778). On high-entropy data the codec expands: random data 1.176 (format-faithful form: 1.547), executable slice 1.115 (1.519). Status after Benchmark 1: a conditional positive — synthetic raster, glyphs aligned to the block grid, no noise, no comparison with G4 or JBIG2.
**Result — Benchmark 2: F5 refuted.**

| image | PGA | zlib-9 | zstd-19 | CCITT G4 | best |
|---|---|---|---|---|---|
| `fax2d` — CCITT fax, 1728×1082 | 0.1888 | 0.1374 | 0.1241 | **0.1209** | G4 |
| `g3test` — CCITT fax, 1728×1103 | 0.2629 | 0.1806 | **0.1624** | 0.1756 | zstd-19 |
| `jim___ah` — dense test image, 664×813 | 0.2443 | 0.2083 | **0.1910** | 0.6301 | zstd-19 |
| text page, sans-serif font | 0.0489 | 0.0440 | **0.0364** | 0.0376 | zstd-19 |
| text page, serif font | 0.0348 | 0.0376 | **0.0302** | 0.0346 | zstd-19 |
| form page | 0.0351 | 0.0263 | **0.0133** | 0.0372 | zstd-19 |
| synthetic form raster (control from Benchmark 1) | **0.0492** | 0.1411 | 0.1069 | 0.4030 | PGA |

The codec loses to zstd-19 on 6 of 6 realistic images and to CCITT G4 on 4 of 6. The only image on which any of the program's own methods comes out best remains the synthetic raster aligned to the block grid. F1: 7/7. A methodological note: in Benchmark 1 the codec rasterised its input without knowing its true geometry, and the raster width matched the true width of the synthetic form raster by coincidence; Benchmark 2 uses the true geometry in every measurement.
**Next step.** Establish what the only positive depends on (section 5).
**Status:** reversibility restored (10/10, 7/7); generalisation refuted (F5); the positive is confined to the synthetic raster.

### 5. Direction 2 — the mechanism of the positive: rigid, bit-exact repetition
**Research question.** What does the only positive depend on — the alignment of patterns to the block grid, their bit-exactness, or both?
**Criteria.** A sensitivity test announced in protocol 1, without a success threshold (what is measured is the degradation curve against G4 and zstd-19 on the same data); F7 (set before measurement): choosing the grid offset restores the compression ratio of the aligned raster.
**Method.** The synthetic form raster shifted horizontally by less than one block width, and made noisy by flipping random bits with a probability of 0.1 to 2%; every result verified with SHA-256 (13/13).
**Result — shift.** Depending on the size of the shift, a compression ratio of 0.0492 to 0.0880, with a dictionary of up to 46 entries instead of 16; G4 unchanged (0.403).
**Result — noise.**

| flipped bits | PGA | dictionary entries | zstd-19 | CCITT G4 |
|---|---|---|---|---|
| 0 | 0.0492 | 16 | 0.1069 | 0.4030 |
| 0.1% | 0.0713 | 20 | 0.1327 | 0.4158 |
| 0.5% | 0.1290 | 78 | 0.1969 | 0.4652 |
| 1% | 0.1991 | 561 | 0.2555 | 0.5195 |
| 2% | 0.3851 | 2,656 | 0.3434 | 0.6180 |

The lead over zstd-19 shrinks from 2.2× on clean data to 1.9×, 1.5× and 1.3×, and reverses at 2% flipped bits (0.3851 against 0.3434). The codec's ratio worsens by a factor of 7.8, G4's by a factor of 1.5. Conclusion: the positive requires rigid, bit-exact repetition of patterns — the kind that born-digital rasters have and scans do not.
**Result — F7: confirmed technically.** A raster deliberately shifted against the block grid: 0.1416; after choosing the grid offset: 0.0492, exactly the value of the aligned raster; cost: a few bits; F1 met. This holds for a shift of the whole raster; it was not measured on rendered pages or scans, where the positions of characters vary individually.
**Construction: withheld.**
**Next step.** The cross-document deduplication hypothesis (section 7).
**Status:** boundary conditions measured; F7 confirmed technically.

### 6. Direction 3 — spatial dependency in 2D: hypothesis refuted, with a predictor ablation
**Research question.** Does the dependency rule, carried over from 1D to spatial neighbours in a raster, carry a signal that 1D lacked?
**State of the art (published).** Context modelling of bilevel images in JBIG and JBIG2: a context template of 10–16 neighbouring pixels and an adaptive arithmetic coder. Here a minimal variant was measured on purpose, to decide whether the axis carries any signal at all.
**Criteria (set before measurement).** F6(a): the majority predictor (majority of three neighbours: left, upper, upper-left) beats the trivial predictors (left neighbour alone, upper neighbour alone). F6(b): the output is smaller than the best baseline for the given image.
**Method.** Residual = pixel ⊕ prediction; causal decoder; final coder zlib-9 or zstd-19; 7 images × 4 variants = 28 measurements; F1: 28/28.
**Result.**

| image | left + zlib-9 | upper + zlib-9 | majority + zlib-9 | majority + zstd-19 | best baseline |
|---|---|---|---|---|---|
| `fax2d` | 0.1359 | 0.1646 | 0.1691 | 0.1567 | 0.1209 (G4) |
| `g3test` | 0.1800 | 0.2178 | 0.2355 | 0.2154 | 0.1624 (zstd-19) |
| `jim___ah` | 0.2085 | 0.3058 | 0.2849 | 0.2642 | 0.1910 (zstd-19) |
| text page, sans-serif font | 0.0450 | 0.0456 | 0.0554 | 0.0469 | 0.0364 (zstd-19) |
| text page, serif font | 0.0387 | 0.0409 | 0.0465 | 0.0376 | 0.0302 (zstd-19) |
| form page | 0.0272 | 0.0264 | 0.0316 | 0.0165 | 0.0133 (zstd-19) |
| synthetic form raster | 0.1412 | 0.1610 | 0.2447 | 0.2189 | 0.1069 (zstd-19) |

F6(a) refuted: with the same coder (zlib-9) the majority predictor was worse than the left neighbour on 7 of 7 images, and worse than both trivial predictors on 6 of 7 (exception: on `jim___ah` it beat the upper neighbour). F6(b) refuted: no variant beat the best baseline on any image (0/7). The best variant beat G4 alone on 3 of 7 images (`jim___ah`, the form page, the synthetic raster) and zstd-19 on the raw data on none.
**Mechanism.** The residual is sparser than the original and still does not compress better. On `jim___ah` it has 0.165 ones per pixel against 0.557 in the original, and with zlib-9 yields practically the same output (0.2085 against 0.2083). With the same coder the left-neighbour residual was at most marginally smaller than the original (`fax2d`: 0.1359 against 0.1374). An interpretation consistent with both measurements: a coder of the LZ family gains from repetitions — identical glyphs and rows — not from bit density; the dependency transform changes the density (towards 0.5 in 1D, downwards in 2D) but adds no repetitions and breaks some of those present. The gain from sparsity does not make up for the lost matches.
**Next step.** None in this form: further work would mean context modelling with an arithmetic coder, i.e. convergence on the state of the art (JBIG/JBIG2), not a new axis.
**Status:** falsified (F6a, F6b); axis closed.

### 7. Program conclusion: the dependency axis closed, the niche narrowed
- **The dependency-coding axis is closed as falsified**, in 1D (sections 1–2) and in 2D (section 6), with final coders of the RLE and LZ classes. Recorded mechanism: the transform changes bit density but adds none of the repetitions the coder relies on.
- **The PGA codec — the niche narrowed.** The program moved from "a better general-purpose compressor" (refuted in the 2021–2025 iterations), through "structured corpora" (documents, scans, telemetry — refuted by measurement for scans, rendered pages, text and telemetry), to a narrow class of data: born-digital rasters with rigid, bit-exact repetition (programmatically generated forms, print spool output, documents filled from a template). There the codec was about 2.2× better than the best classical baseline — measured with full reconstruction — and about 2.6× after a further improvement of the container; the latter is an estimate from section sizes, not yet measured end to end. Construction: withheld.
- **Boundary conditions of the niche, stated plainly.** The lead was measured on a single synthetic raster; the comparison covered G4, zlib-9 and zstd-19, not JBIG2, which uses symbol dictionaries and is the closest state of the art; robustness to noise is weak (section 5). This is a scoping hypothesis, not a demonstrated lead on real data.
- **Next hypothesis: cross-document deduplication.** A shared dictionary for a series of pages from the same template (e.g. series of invoices or reports); expectation: the dictionary cost amortises over the series, and the compression ratio of the series falls well below that of a single page. JBIG2 provides symbol dictionaries shared across pages, so without a comparison with it the result of this hypothesis will not be conclusive. The criterion will be set before measurement, as in both sessions.

### 8. Corpus, baselines, reproducibility
**Benchmark 1 (2026-07-15), 10 files.** Five structured: technical text (project documentation, 18,309 B), a synthetic form raster at 1 bit/pixel (512 KiB; table lines and a small fixed set of glyphs placed on the block grid, no noise), synthetic telemetry in CSV format (418,929 B), a sparse file (256 KiB; zeros and a 16-byte record header every 4,096 B), a file of long runs of 0x00 and 0xFF bytes (256 KiB). Two high-entropy control files: pseudo-random data (256 KiB) and the first 1 MiB of the compiled Windows executable from the entry of 2026-07-14. Three edge cases of 256 KiB each: all zeros, alternating bits (byte 0xAA), a periodic sequence (bytes 0–63). The synthetic files are deterministic (fixed generator seed).
**Benchmark 2 (2026-07-17), 7 bilevel images.** Three real images from the public libtiff test-image archive `pics-3.8.0` (download.osgeo.org/libtiff/pics-3.8.0.tar.gz): `fax2d` (1728×1082) and `g3test` (1728×1103) — classic CCITT fax test images — and `jim___ah` (664×813; dense: 56% black pixels, not a single blank row). Three pages rendered with TrueType system fonts at about 200 dpi (1656×2336): text in a sans-serif font, text in a serif font, and a form with ruled lines and labels in a monospaced font; line and character positions deliberately not aligned to the block grid. Control: the synthetic form raster from Benchmark 1. The rasters are stored as raw bits, 1 bit/pixel, row by row, black = 1.
**Baselines.** zlib-9 (DEFLATE, as in gzip), zstd-19, CCITT G4 (in Benchmark 2; measured as a complete TIFF file). JBIG2 — not measured: no encoder in the measurement environment.
**Environment.** Windows 11 (AMD64), Python 3.12.10, numpy 2.5.0, Pillow 10.4.0, zstandard 0.25.0.
**Verification.** Benchmark 1: 154 measurements, 144 decodes matching the SHA-256 of the original; the 10 mismatches are the historical version of the PGA codec (section 3). Benchmark 2: 51 measurements, including 48 decodes — 48/48 matching; the remaining 3 measurements concerned a construction detail of the container and their result is withheld.
**Reproducibility.** Anyone can repeat the baseline results on the libtiff images: the archive is public, and the geometry and raster format are given above. The results of the program's own methods require the withheld construction; the hashes in section 9 bind them to the protocols.

### 9. Evidence: SHA-256 hashes and timestamps
The measurement protocols remain private: they contain the withheld construction. Their SHA-256 hashes, published here, make it possible to check later that a document produced in the future is the same one that was stamped in July 2026. The evidence lists contain the SHA-256 hashes of the raw result files and the corpus manifests, and the commit identifiers of the measurement code (the first list also the hash of the 2025 whitepaper).

| document | SHA-256 | OpenTimestamps stamp |
|---|---|---|
| measurement protocol no. 1 (Benchmark 1) | `0830cc94c042c05103e45a47de2bdae0db2411f683fcd309e6630c5700552375` | 2026-07-15, Bitcoin attestation |
| evidence list for Benchmark 1 | `8b85e99708febb32682a100f5b2336bd0ffc8cf0e0ef4c7671aa91576de6dc56` | 2026-07-15, Bitcoin attestation |
| measurement protocol no. 2 (Benchmark 2) | `c478b509fda7419a60b0ad2149d28f4dec83aae5fb074ffbeab0cd3bb3aba1de` | 2026-07-17 |
| evidence list for Benchmark 2 | `7b2a7a3fa8a49c3adda2e6604cdf462106f1ab4fc11a26c4246ba8e955324b09` | 2026-07-17 |

The hashes of both protocols match the entries in the evidence lists. Where a protocol states a number or a description that does not match the raw result files, this entry gives the value from the result files: (1) protocol no. 1 states that with zlib-9 the output worsens monotonically with the number of layers on every file — according to the data this holds for 5 of 10 files (section 1); (2) protocol no. 1 assesses F4 against zstd-19, although F4 is worded against the input size — section 4 gives both readings; (3) protocol no. 2 reports 30/30 verified decodes in the 2D test — the result file contains 28, and the session total of 48/48 is correct; (4) protocol no. 2 states that the majority predictor was worse than the left neighbour on 6 of 7 images — with the same coder it was worse on 7 of 7, and worse than both trivial predictors on 6 of 7; (5) protocol no. 2 describes `jim___ah` as a handwritten scan — 56% black pixels and no blank rows point rather to a dense or halftoned image, so we describe it neutrally.

### Corrections to the entry of 2026-07-14
- **Direction 1, criterion (b).** Entry: "measured ratio on defined corpora against gzip/zstd — the measurement program has not yet been run". Now: measured on 2026-07-15 and refuted (section 1); the axis is closed, in its 2D variant too (section 6).
- **Direction 1, construction.** Entry: "Construction (dependency rule, layer construction, container format): withheld". The rule and the layer construction are now disclosed — XOR of neighbouring bits, iterated layer by layer — because the axis is falsified and closed.
- **Direction 1, reversibility.** The claim stands (100/100 and 4/4), with one precision: an earlier reference implementation stored only the first boundary bit and could not decode more than one layer; the measurements used a corrected re-implementation.
- **Direction 2, reconstruction criterion.** Entry: "Bit-exact reconstruction (met, hash-verified)". For the PGA codec this was not true: the verification concerned an earlier proof of concept. The first full validation of the PGA codec (2026-07-15) gave 0/10; only the repaired variant is reversible (sections 3–4).
- **Negative result of +56%.** The size was measured correctly (14.2 → 22.2 MB), but the artefact is not fully reproducible and so was not a valid lossless encoding. The conclusion — expansion on high-entropy data — is confirmed by the repaired variant with SHA-256 verification: on the first 1 MiB of the same file +52% (format-faithful form) and +12% (redundancy-free map), on random data +55% and +18%.
- **Direction 2, state of the art.** The entry compared against gzip/zstd. The proper state of the art for bilevel images is CCITT G4 and JBIG2; G4 was measured in Benchmark 2, JBIG2 not yet.
- **Direction 2, status.** Entry: "POC implemented / in validation". Now: validated — the historical version fails the reconstruction criterion, the repaired variant meets it, and the positive does not carry over beyond the synthetic raster (sections 3–5).
- **Scoping conclusion.** The entry narrowed the program to structured corpora (documents, scans, telemetry). The measurements narrow it further: on scans, rendered text pages, telemetry and text the classical coders win; what remains is the niche of born-digital rasters with rigid repetition (section 7).

**Program status:** active. Dependency-coding axis — closed as falsified, in 1D and 2D. PGA codec — reversible in the repaired variant, with a lead measured only on a single synthetic born-digital raster; next step: cross-document deduplication on series of pages from one template.
