Build log ·

Phase 2 — v1 wire format and the three hotfixes the staged rollout caught

Phase 2 introduced the v1 self-describing wire format on disk for new uploads, through a five-stage rollout with mandatory bake windows. The staged design caught a memory bug in code review, a base-branch misroute in PR review, and two v0-shape assumptions in cross-deployment testing. The artifact most worth keeping is a cross-layer audit method that generalizes to Phase 3.

Written by Cho García

Phase 2 was a wire-format migration. The v0 format ShieldFive has used since launch is a working AES-256-GCM implementation that pre-dates the design discipline now applied; v1 replaces it with a self-describing format whose chunks bind their own offsets and a context tag into AAD and which a reader can decode without out-of-band parameters. v1 is the foundation Phase 3 will sit on. v0 and v1 coexist on disk indefinitely; no file is ever rewritten and no user is ever migrated.

This post documents the staged rollout that surrounded the format change, the three hotfixes the staging caught, and the cross-layer audit practice that came out of one of them. It is the second build-log entry; the first covered the Phase 1 library extraction.

The staged rollout design

Phase 2 was designed as five sequential stages with mandatory bake windows before any code was written. The design lives in docs/phase2-design.md in the web repository and predates the first Phase 2 commit by about three weeks.

  • Stage 0 — refactor the existing decrypt path so v0 reads go through a dispatcher, without yet introducing v1.
  • Stage 1 — admit cipher_version = 2 at the database constraint level, with no code path yet writing it.
  • Stage 2 — implement v1: an upload proof v2 with a three-way cross-check (2a), a v1 writer worker (2b), and the API surface plus a feature flag for staged enablement (2c).
  • Stage 3 — cross-deployment test: upload a v1 file on a preview deployment, decrypt it on production using the production decryption path.
  • Stage 4 — flip the feature flag in production.

A 24-hour bake window separates each sub-stage; Phase 1 itself carried a seven-day bake before Stage 0 began. Each stage carries a stop-and-surface clause: if the previous stage's input assumption is not yet observable as true in production, the stage does not start.

The point of the staged design is not the calendar. A wire-format migration has four or five places it can go wrong — TypeScript validation, SQL validation, client dispatcher, server dispatcher, the writer itself — each of which a single end-to-end test of a single happy path will fail to catch. Five stages with five different "the previous stage's assumption is visibly true" checks give each layer its own gate.

Stage 0 and Stage 1 — the boring stages

Stage 0 (PR #107) refactored useDecryptWorker and sfCryptoDecrypt to route by cipher_version instead of branching inline. No new wire format. The diff was almost entirely test updates: the dispatcher had a v0 branch and a v1 branch that threw "v1 not yet implemented." 24h bake.

Stage 1 (PR #109) dropped the files_cipher_version_known constraint and re-created it as cipher_version IN (1, 2). v1 was now permissible at the database level; nothing yet wrote it. 24h bake.

These stages are boring on purpose. The interesting part of any migration is the stage where a real change finally happens; the boring stages exist to make that stage smaller.

Stage 2a — three-way cross-check on upload proof v2

Stage 2a (PR #110) introduced the v1 upload proof: a base64 envelope of [0x02][0x02][32-byte HMAC] where the HMAC binds the wire-format version into the proof itself. Verification runs in TypeScript on the server before the row is finalized; the SQL function does not need to know its shape. The three-way cross-check is between the proof the client sends, the proof the server reproduces from the ciphertext, and the proof the server re-derives from the metadata — three independent paths that have to agree before the upload is accepted.

The Stage 2a verification was structural: every place that ever emits or consumes an upload proof was audited against a written list of "what shape does this code path expect, and is that correct for both v0 and v1." The pattern is the precursor to the cross-layer audit that became necessary later in Stage 2c.

Stage 2b — the drain-pump pushback (commit 1ee2fdb2)

The original Stage 2b design for the v1 writer worker was a "drain pump": the worker would accept frames as they were produced and buffer them internally, flushing them to object storage as a separate pipeline stage. This worked in unit tests against small inputs. It also made the worker the owner of an unbounded buffer.

The pushback surfaced in code review and was specific: the worker's peak memory was a function of how fast object storage accepted bytes, which is a function of the user's network, which is not bounded. On a slow upload of a large file, the drain pump would hold gigabytes of ciphertext in memory while waiting for object storage. The library's whole point — proven by the §0 streaming-memory characterization run before any application code was written — is that peak memory is bounded and independent of input size. The drain pump threw that property away inside the worker.

The refactor at commit 1ee2fdb2 (PR #111) replaces the drain pump with a single-frame-in-flight design: the worker holds at most one chunk at a time, and the next frame is requested from the library only once object storage has accepted the previous one. Back-pressure propagates from object storage through the worker to the library, which propagates it to the input stream. Peak memory is bounded.

This is the form of bug the staged design exists to catch. 1ee2fdb2 landed before Stage 2b shipped to production; the design caught the bug in code review because Stage 2b had a written "what assumption from Stage 0 must hold here" clause, and the drain pump violated it on inspection.

Stage 2c — base branch misroute and the cross-layer audit

Stage 2c's first attempt was PR #112, which introduced the API surface and the feature flag. The diff was wrong — not because the code was wrong, but because the branch was cut from the Stage 2b feature branch rather than from main. Half the diff was Stage 2b code already merged through a different PR; reviewers could not tell what was new and what was rebase noise.

The recovery was a clean re-branch from main as PR #115. PR #112 was closed with a pointer; no commits were lost. The fix took twenty minutes; the time worth recording is the half hour spent checking that the rebased branch's git diff main...HEAD actually matched the intended Stage 2c scope before opening #115.

The substantive output of Stage 2c was the cross-layer audit document at docs/phase2-stage-2c-audit.md in the web repository. The audit re-checks every TypeScript and SQL site that ever touches a proof, hash, or cipher parameter, classifies each site as already correct, already broken, or new code from Stage 2c, and confirms that the dispatcher routes around v0-shape assumptions on the v1 path. The doc is structured as two tables (TypeScript layer, SQL layer) with a stop-condition checkpoint. It is the artifact a future reader can verify Phase 2's correctness against without reading the migration's PR history.

Stage 3 — the two hotfixes

Stage 3 was the cross-deployment test: upload a v1 file on the preview deployment with the v1 feature flag set, then decrypt it on production using the production decryption code (which had shipped its v0/v1 dispatcher in Stage 0 and admitted v1 at the constraint level in Stage 1). The test was designed to validate the full v1 path against the production reader, isolated from any v1-aware writer.

It failed twice.

The first failure was that complete-upload's row-level NULL check rejected v1 rows because v1 writes cipher_chunk_size = NULL and cipher_nonce_prefix = NULL — v1 doesn't have a fixed chunk size or a nonce prefix, since each frame is self-describing. The check had not been gated on cipher_version = 1. PR #116 re-gated it and, for hygiene, added an explicit 400 for v1 multipart-via- proxy uploads, which are not implemented in Phase 2 because the web client uses direct-to-object-storage multipart.

The second failure surfaced on the next pass, after PR #116 had baked: the finalize_file_upload_with_quota SQL function rejected the v1 proof because it carried a regex '^[0-9a-f]{64}$' matching v0's 64-hex HMAC. v1 proofs are base64, about 48 characters. PR #117 dropped the regex outright. Cryptographic verification of the proof happens in TypeScript before the RPC is called, so the SQL-layer shape check added no security; it added only a coupling between the database and a specific wire-format string shape that every future revision would have to re-migrate.

Both failures were genuine: real bugs that would have been user-visible if v1 had been enabled in production without the cross-deployment test. The pattern in both is the same — a v0-shape assumption baked into a code path that the v1 dispatcher routes through. The cross-layer audit was written after PR #117 partly to verify there were no remaining instances of the pattern. It found none.

The stop-and-surface clause was exercised both times: Stage 3 did not advance to Stage 4 until each hotfix had its own 24-hour bake window. The bake windows were not negotiable.

Stage 4 — the flag flip

Stage 4 was a single environment-variable change on the production deployment, taking about sixty seconds. v1 is now the default for new uploads. Existing v0 files continue to be decryptable; the v1-rejecting code paths that Stage 3 surfaced no longer exist.

Stage 4 was anti-climactic by design. The interesting part of a staged migration is everything that happens before Stage 4. If the flag flip is interesting, the rollout failed.

What the cross-layer audit method is, and why it generalizes

The audit doc is structured around one premise: a wire-format migration breaks when a v0-shape assumption is baked into a code path that both v0 and v1 traverse. The shapes in question are narrow — a regex, a NULL/non-NULL gate, an integer-vs-string check — and so are the code paths. A fresh-eyes grep of every site that mentions a proof, hash, or cipher parameter is feasible; reviewing each site against "is this v0-shape, and is that correct for the paths v1 routes through it" is feasible. The output is a table.

Phase 3's post-quantum hybrid wire format will be a third entry on the same axis. It will introduce a new proof shape and possibly a new chunk layout. The cross-layer audit method as written applies to that migration with the inputs swapped: grep for v0-and-v1-shape assumptions, classify, fix, re-grep. The document is a template, not just a record of one phase.

What's next

v1 has been the default writer for new uploads since the Stage 4 flag flip. v0 files continue to be readable indefinitely through the bridge that landed in Phase 1. No file on disk has been rewritten and no user has been migrated.

The next build-log post is the internal security review summary for Phase 2 — written separately because security review reads differently than build-log narrative.

Phase 3, sometime after Phase 2's tail of unknown-unknowns is empty, makes the post-quantum hybrid suite (ML-KEM-1024 alongside the classical secretbox share, HKDF-combined, with XChaCha20-Poly1305 for chunk AEAD) the default for new uploads. The same staged-rollout design and the same cross-layer audit method apply, with the v2 wire format as the new entrant.