Ygoow: pre-registered measurement of our own security claims
Continues the entry of 2026-07-10, now naming the product. From July to September 2026 Ygoow’s security claims were turned into two-stage cards — criteria and predictions first, measurement after — on the target phone or in simulation: 37 cards (24 with the registration committed before the run), 199 ledger predictions of which 36 missed, eight published claims overturned and withdrawn, and corrections to statements of the July entry that no longer hold. Constructions withheld; traffic results come from a model, not from captured Tor traffic.
Koch Laboratory — Ygoow: pre-registered measurement of our own security claims
The entry of 10 July 2026 (“Metadata-resistant, account-less communication”) described seven research directions for a messenger with a deaf relay, without naming the product. The product is Ygoow (ygoow.com); the Android app has not been released yet. From 24 July to 29 September 2026 we turned its security claims into measurement cards. Below: the method, the counts, results already public on the product site, and corrections to the July entry. A summary of the method: Measured, not assumed.
Method: the card in two stages. Stage 1, written before the instrument runs: the question; the adversary — the strongest one we can build from the public protocol description; a population model with the chance level computed; criteria, numerical predictions, candidate fixes and a decision rule. Stage 2, the measurement: a null control (two samples of the same class) and a positive control (a known effect that must be detected in the same session); time and cost on the reference phone (ARM64, release build), traffic, storage and groups in simulation of the production code or of a model of it; a verdict on every criterion and every prediction, additions after the first numbers marked post hoc, deviations named. The practice tightened along the way: since 3 September every number has come from at least three runs or random seeds, and since 6 September stage 1 has gone into the repository in a commit of its own before the measurement.
Counts (as of 29 September 2026). A card is one protocol in the product’s repository.
- 37 cards (24 July – 29 September), all with stage 2; the results of 11 of them, from 20–29 September, are not yet published. 24, all from 6 September on, have stage 1 in a commit of its own before the measurement; the 13 earlier ones recorded criteria and results in the same commit, so there only the card’s text attests the order.
- Prediction ledger (from 3 September, 25 cards): 199 predictions — 113 held, 48 held in part or only in direction, 36 missed, 2 could not be assessed; 19 of these cards contain a miss.
- Claims overturned: 8 sentences that ygoow.com had published as a property of the product and that a card contradicted; all were withdrawn or rewritten (marked “claim overturned” below). Two more stood only in internal documents.
- By direction: 1 — 11, 2 — 4, 3 — 11, 4 — 3, 5 — 5, 6 — 3, 7 — 0.
Rules the misses forced.
- We measure timing only on the target phone, in a release build, and accept a negative only with a positive control that fired in the same session: ML-KEM was clean on the desktop and showed a 15.6 % timing difference on the phone; a debug run did not see even the known leak.
- The adversary must be at least as strong as the strongest one that can be built from the public protocol: a weaker one picks the wrong fix. Three predictions of its strength all missed in the same direction, by a factor of 6–24.
- Before registering a threshold: the chance level, whether the threshold is attainable at that level, and one case walked by hand through the mechanism — taken from the tail of the distribution.
- Traffic numbers come from a model of the public protocol and a declared population, not from captured Tor traffic.
1. Direction 1 — Post-compromise security without a stateful server: the primitives underneath
Question. Do encryption, signature, key agreement and key derivation leak a secret through their running time on the phone? State of the art. TVLA and dudect (Welch’s t-test), RFC 8032, NIST SP 800-38D, RFC 9106, zxcvbn. Criterion. |t| ≤ 4.5 in every run with a clean null control and a positive control that fired; output matching the test vectors. Method. 11 cards (2 from 20–29 September), phone and simulation; none measured the healing after compromise itself. Result. Refuted and fixed: Ed25519 (|t| = 60 and 71, natively 2.6 and 1.6), the quorum arithmetic (direction 7) and AES-GCM (an effect of about 1.4 %, visible only with a local clock; published before the fix, natively 24 of 24 readings at noise level). X25519 clean. Claim overturned: “an honest strength meter” — it overstated the strength of “word + digits + symbol” passwords by a median of 45 bits; 97 % of them passed the backup file’s check, today 0.3–0.6 %. The change of 28 August, made without a card (less memory in the backup’s key derivation), was overturned by a card two days later. Next. The Double Ratchet in the app, which also closes the forward-secrecy gap; an unexplained timing anomaly in the third-party SHA3 library. Construction: withheld. Status: primitives measured and in the app; post-compromise security built and verified against test vectors, outside the app.
2. Direction 2 — Post-quantum hybrid (X25519 + ML-KEM)
Question. Does the hybrid protect the first message against someone who records today and breaks X25519 later, and does it fit the key exchange between contacts? State of the art. FIPS 203, Signal PQXDH, X25519MLKEM768, KyberSlash. Criterion. The post-quantum half holds after X25519 is broken; an invitation fits one QR code at error-correction level M; |t| ≤ 4.5 on the phone. Method. 4 cards: analysis of the production flows, the app’s QR library, timing on the phone. Result. Claim overturned, published in July: the post-quantum key derived from the X25519 secret was meant to protect the first message, but two supported flows expose public keys from which that secret can be computed once X25519 is broken; since 23 August the post-quantum key no longer depends on X25519. Refuted: a single QR code is too dense — it takes a split QR code, NFC or a link. Refuted and fixed: a secret-dependent division in our own ML-KEM implementation — 15.6 % on the phone (|t| = 339), after the fix |t| = 3.9. Next. Wiring into live conversations and into the contact code; until then, “harvest now, decrypt later” is not covered in the app. Status: built and pinned to the published accumulated test vectors, outside the app.
3. Direction 3 — Metadata-resistance layer
Question. What does the relay read from the size, rhythm and number of frames without seeing content or the IP address? State of the art. Loopix, Tor’s netflow padding, the likelihood-ratio test, mutual information. Criterion. Size carries ≤ 0.5 bit about message length and depends on public parameters only; no frame is “certainly real” by its size; grouping a device’s addresses stays at ≤ 2× chance. Method. 11 cards (4 from 20–29 September): a model of the public protocol on the production code paths, a relay with a clock as the adversary. Result. Claims overturned and fixed: “never the exact length” (rooms unpadded; today 0.043 bit of length) and “indistinguishable decoy frames” (always the smallest class; today 0.0002 bit in size, 0.0018 bit in spacing). Claim overturned: “the relay cannot tell when you send” — cover traffic hides messages, not sessions, and is off by default; without it the relay keeps 98.5 % of the information about when someone talks. Claims overturned: “the relay cannot batch addresses into one device” and “personas do not reveal a shared device” — a likelihood-ratio adversary groups a device’s addresses 85–90 % of the time after one 15-minute epoch and 99.8 % within 30 minutes. Found: the unencrypted message counter in the header links a conversation across the address rotation 81–89 % of the time among 20 active conversations. Next. Two measured designs for closing the timing channel await a decision; header encryption in progress; a measurement on real traffic. Construction: withheld. Status: size padding and optional cover traffic in the app; the timing channel and the counter open.
4. Direction 4 — Address-less rendezvous with epoch rotation
Question. Does rotating the conversation address roughly every 15 minutes break the continuity visible to the relay without losing messages? State of the art. Tor onion services, store-and-forward delivery, linkability across identifier changes. Criterion. Loss of messages to an absent recipient ≤ 0.1 %; consecutive epochs of a conversation linked through the cursor alone ≤ 1 %. Method. 3 cards (2 from 20–29 September): simulation with and without background delivery. Result. Refuted and fixed: in the default configuration, messages to someone absent for more than 15–30 minutes could be lost without notice (57–58 % in the model; today 0 %), and the catch-up cursor itself linked a conversation’s addresses (100 %; today 0 %). Next. The results of two cards of 22 September — in a later entry. Construction: withheld. Status: address rotation and trail-free catch-up in the app; against the relay, rotation alone is not enough (direction 3).
5. Direction 5 — Coordinator-free healing of group epochs
Question. Does a membership change reach every member, including absent ones, and what does the relay learn about the group from control frames? State of the art. MLS (RFC 9420) with an ordering server, Sender Keys. Criterion. At least 99 % of members on the same key within 24 h with background delivery (95 % without); loss of room messages ≤ 0.5 % in every seed. Method. 5 cards (3 from 20–29 September): simulation on a model and on the production code. Result. Refuted: a change reached only members who opened the room while it was fresh — in the model, 3–9 % of members shared the key within a day, and 29–55 % of room messages were lost. The criterion was met only by a variant added after the first numbers (post hoc): 100 % and under 0.5 %. It was deployed anyway, and the deviation from our own rule is recorded on the card; the on-device test is pending. Found and fixed: when a room was opened, the relay read its size and membership; it still sees that the membership changes (100 % in the model) and, from the header counter, how many members are writing. Next. Protection against replayed key frames; forward secrecy for distributing sender keys; a removed member still sees when the room is active. Construction: withheld. Status: coordinator-free healing and background delivery in the app; the limits under “Next” remain open.
6. Direction 6 — Deniable store with a hidden volume
Question. Does the hidden profile stay unprovable after months of use — against someone who copies the phone’s storage and times the unlock? State of the art. VeraCrypt-class hidden volumes; the multi-snapshot adversary (Czeskis et al., 2008). Criterion. Without the decoy password, a detector on one or two copies performs at chance level; no data loss in at least 300 scripted lifetimes; no difference in unlock time. Method. 3 cards: the production storage code over months of simulated use, a stopwatch on the phone. Result. Claim overturned: “indistinguishable on disk” — after 30 days a single copy revealed the hidden profile on 69 % of simulated phones whose decoy was never opened, two copies on 98–100 %; unlocking took about 1.1 s longer, and using the decoy could overwrite the hidden profile. After the redesign: 0 data lost in 600 lifetimes, 0 size mismatches in 7,200 snapshots, unlock within −3 to +28 ms of a phone without a decoy. The prediction registered for the redesign missed — a password change had a signature of its own; the last of two further fixes was checked by a test, not on the population. Next. With the decoy password in someone else’s hands, two copies still show use of the hidden profile; re-randomisation without the key costs 14–40 times the write budget today. Construction: withheld. Status: in the app, measured over the lifecycle; the limits with a surrendered decoy password are named.
7. Direction 7 — Conditional access with quorum
Question. When the unlock conditions — password, file, link, K-of-N quorum, hardware key — are evaluated, do the secret and the location stay hidden? State of the art. Shamir secret sharing over GF(2⁸); conditions enforced on the client. Criterion. Quorum arithmetic with |t| ≤ 4.5 on the phone; evaluating a condition sends nothing outside Tor. Method. No card of its own: the quorum arithmetic was measured by a card of direction 1, the location condition was examined outside the cards. Result. Refuted and fixed: the table-based multiplication in GF(2⁸) leaked (|t| up to 1222); the branch-free, table-free version is exhaustively equivalent over all 65,536 input pairs. Negative result: the location condition was removed because evaluating it sent the surrounding Wi-Fi and cell networks to Google, outside Tor. Next. A measurement of the condition graph itself. Construction: withheld. Status: the lock with any combination of conditions in the app; the condition graph without a measurement card.
Corrections to the entry of 2026-07-10
- Method, “verification as a testable property”: vectors check correctness, not side channels and not behaviour over time — four primitives passed them and leaked on the phone, and the decoy store lost data.
- Direction 2, “Status: design”: built today, outside the app; a native library (FFI) was not needed to remove the measured leak.
- Direction 3, “a constant Poisson rate”: what runs is a fixed grid whose phase is drawn afresh at every address rotation. “Never the exact length” and “indistinguishable decoy frames”: false until 1 September for rooms and for large cover frames. “Each persona rides its own Tor circuit”: true, but a shared device clock linked the personas. “pcap diff”: not carried out; traffic numbers come from the model.
- Direction 4, “cannot batch addresses into ‘one device’” and “at most one 15-min window”: false against the relay, true only against an outside observer and a malicious contact. “The client carries the cursor” and the status “implemented and tested”: the cursor itself linked a conversation’s addresses, and messages to absent recipients could be lost; fixed on 17 September.
- Direction 5, “no loss of MLS security properties”: not met in full — the limits are listed under direction 5.
- Direction 6, “mechanism implemented”: it lost data and was detectable after 30 days; redesigned on 17 September.
Status of the direction. Built, but outside the app: post-compromise security and the post-quantum hybrid. Open towards the relay: the timing channel and the header counter. Outside this series: an independent review of protocol and implementation, and a measurement on real traffic. A later entry will describe the eleven cards from 20–29 September; the current state is kept on ygoow.com/progress.
Proof of time
This text was hashed and stamped on publication. Verifying it takes both files: the stamp proves that one exact byte string existed at that moment, and only the source below is that byte string. Any later edit changes the hash and breaks the match — which is the point.
SHA-256 of the stamped text: e7809aa121bebb53a42cf9597437a55ba01b50a48bcf2a28cf4c049345f99442
- Sigelith anchor portfolio-ygoow-measurements.public.md.beattime.json
- OpenTimestamps proof portfolio-ygoow-measurements.public.md.ots
- Stamped source text portfolio-ygoow-measurements.public.md
Verify with: ots verify <proof> --file <source>