Phase 3: Refracting Light v2 hash tests

Run: 6 October 2026, Apple M1 (8 GB), Python 3.14, seeded and reproducible Spec under test: REFRACTING-LIGHT-SPEC.md Bottom line: No statistical weakness found. That is necessary for a hash but nowhere near sufficient: SHA-1 and MD5 also passed tests like these years before they were broken. Refracting Light stays experimental. Use SHA-256 or BLAKE3 for anything consensus-critical.

Results

Test

Script / evidence

What was measured

Result

Verdict

A. Avalanche

phase3_hash_tests.py → phase3-runs/phase3-20261006-124709.json

Flip each input bit (4-, 16- and 64-byte messages) and count changed folded-output bits

Mean 64.04 / 63.97 / 64.05 of 128 (ideal 64). Minimum seen: 42

PASS

A2. Outlier re-check

phase3_bit_recheck.py → bit-recheck-20261006-124601.json

Bit 190 (64 B) looked weak (61.88). Re-tested with a fresh seed, 2,000 samples

63.88, z = −0.96. Bits 49, 350, 0, 100, 300 and 511 all within ±2.2

False alarm

B. Output bit bias

same as A

200,000 random 8-byte messages, ones-count per output bit

χ² = 104.3 on 128 dof (expected 128 ± 16), max |z| = 2.24

PASS

C. Single-bit differentials

same as A

For 7 input-bit positions in 16-byte messages, 4,000 random messages each: does one flip give a repeatable output XOR?

4,000 distinct output differences out of 4,000 every time. No differential repeated even twice

PASS

D. Truncated birthday

same as A

400,000 counter messages, collisions on 16/24/32-bit slices (low and high)

All within ±2.2σ of the ideal count. No full 128-bit collision

PASS

E. Reduced-step diffusion

phase3_diffusion_steps.py → diffusion-steps-20261006-124340.json

One flipped digit from a random mid-message state, measured after k further steps

k=0: 91/512 unfolded, 51/128 folded. k=1: 245/512. k≥2: ~256/512 and ~64/128 (full)

See note

F. Linear masks

phase3_linear_masks.py → linear-masks-20261006-124450.json

6,128 GF(2) masks, 24,000 samples: random 2–4-bit output masks, the 128 four-bit “fold position” masks, input⊕output masks

max |z| = 3.51 (random expectation ≈ 4.34). None over 4. Fold-position masks max 2.31

PASS

Note on test E (no finalization stage)

A difference that enters in the very last step only reaches 91 of 512 state bits (51 of 128 folded). Full mixing takes 2 steps. Because of the record layout, every message bit first enters at least 16 steps before the end (the P and Q check shards follow the data), so the statistical safety margin is about 8×.

Two cautions:

  1. A message flip is re-injected through P and Q, possibly in the last step. A designed (non-random) attack could try to use that last poorly mixed step to cancel an earlier difference.

  2. Adding a finalization stage of a few blank steps, or a final mix after the fold, is cheap. It would remove the question. That is the Phase 3 item 2 recommendation, and it would be a v3 (new test vectors).

Not done in this pass, and why

Planned item

Status

MILP

Not run. SAT (below) covered the same ground first

Targeted search for message pairs that land in the 384-dimensional fold kernel

Stage 1 done (section H): no foothold found. A full dedicated attack (meet-in-the-middle, hand-built trails) remains open for outside reviewers

Independent C/Rust port, byte-identical

Not started. The Python spec check is independent of the reference code, but in the same language

Weak-variant sweep with a finalization stage

Waits on the v3 decision above

G. SAT attacks (z3 5.1.0, installed into .venv-qtl on 6 Oct 2026)

Model check: phase3_sat_attacks.py builds RL-w, a width-scaled copy of v2 in which every 64-bit word becomes w bits. At w = 64 it reproduces all recorded v2 vectors. Its z3 encoding matches the concrete code at w = 6, 8 and 12. So the solver attacks the real design. At w ≤ 16 the constant 65537 reduces to 1, so small variants are, if anything, weaker than v2.

G1. Full hash, small widths (phase3-runs/sat-20261006-125649.json, 60 s limit)

Variant

Attack

z3

Plain brute force (est.)

RL-4 (8-bit output)

preimage

found in 14.3 s, verified

256 tries, 0.09 s

RL-4

collision

timed out (60 s)

~20 tries, 0.007 s

RL-6 (12-bit output)

preimage

timed out (60 s)

4,096 tries, 1.5 s

G2. Reduced steps, raw step function (phase3_sat_reduced.py → sat-reduced-20261006-130308.json, 90 s limit). z3 had to find k input bits that reach a random folded target from a random mid-message state.

Width

Largest k solved

Time at that k

First timeout

w = 8 (16-bit out)

12 steps

30.3 s

16 steps

w = 16 (32-bit out)

12 steps

34.3 s

16 steps

w = 64 (real v2, 128-bit out)

8 steps

13.1 s

12 steps

Reading: In every case, z3 was thousands of times slower than simply trying all 2^k inputs, and its time grew at least as fast as brute force. The solver found no algebraic shortcut through the step function, even at 8–12 steps. Every real message passes through at least 112 steps, and at least 16 after any message bit first enters. This is the expected result for a well-mixing design. It rules out easy algebraic weaknesses only; it is not a proof against expert differential or meet-in-the-middle attacks.

H. Fold-cancellation search, stage 1 (fold_cancellation_search.py → fold-cancellation-20261006-131515.json)

A folded collision happens exactly when the final 512-bit state difference Δ lies in the 384-dimensional kernel of the fold. Generic search costs about 2^64, so this stage looks for a foothold: a class of message differences that pushes Δ toward the kernel. It also checks whether a class confines Δ to a subspace. A GF(2) rank below full would cut the collision cost for that class from about 2^64 to about 2^((128−deficiency)/2).

The difference classes target the record’s P/Q re-injection and the weak last step. Each class used 20,000 pairs of 16-byte messages. Positions are message byte indices; D_k[j] = m[4k+j].

Class

Message difference

Folded Δ weight mean (z)

Min (expected ≈41)

Rank folded / unfolded

random-pair (baseline)

independent messages

63.98 (−0.48)

42

128 / 512

single-early

m[0] ^= 0x80

64.06 (+1.52)

43

128 / 512

late-reinject (Q copy hits the very last step)

m[3] ^= 0x01

63.97 (−0.87)

41

128 / 512

last-data-byte

m[15] ^= 0x01

64.05 (+1.35)

41

128 / 512

p-silent (P unchanged)

m[3], m[7] ^= 0x01

64.00 (+0.08)

44

128 / 512

q-silent (Q unchanged)

m[3]^=1, m[11]^=3, m[15]^=1

64.03 (+0.79)

43

128 / 512

pq-silent (no re-injection at all)

m[3]^=1, m[7]^=2, m[11]^=3

64.01 (+0.21)

43

128 / 512

No folded difference ever repeated (each class: 20,000 distinct). Every class has full rank, so no subspace foothold exists in these classes.

The weak last step, measured directly. Across 200,000 random states, a difference that enters only in the final step gives a folded Δ weight with mean 51.1 and minimum 21 (a random 128-bit value would give a minimum of about 39). That is real structure. It is not reachable by message pairs, because two messages would need identical 512-bit states before the last step, which is itself a full state collision. But it is the clearest argument for a finalization stage: with even 2 blank steps, the last-step structure disappears (section E: full mixing after 2 steps).

Conclusion: No fold-cancellation foothold in v2 from the targeted classes. Finalization is still recommended. It costs a few steps, removes the one measurable weak spot, and is easier to argue to reviewers than “not reachable”.

I. Avalanche, v2 vs v3 (avalanche_test.py → avalanche-v2-v3-20261006-133741.json)

Definition. Flip one input bit and count how many output bits change. Ideal: each output bit behaves like a coin toss, so the count is Binomial(128, ½), with mean 64 and sd 5.66. No input position may sit noticeably low. Strict Avalanche Criterion (SAC): for every (input bit, output bit) pair, the output flips with probability ½. Checked on a 128 × 128 grid for 16-byte messages.

Messages × flips

Mean (64)

SD (5.66)

Weakest position z (random max ≈)

SAC max |z| (≈4.56)

Unfolded (256)

v2 4 B

4,000 × 32

63.970

5.658

−2.01 (2.88)

—

256.02

v2 16 B

2,000 × 128

64.003

5.650

−2.21 (3.33)

4.20

256.02

v2 64 B

120 × 512

64.023

5.660

−3.24 (3.72)

—

255.88

v3 4 B

4,000 × 32

64.020

5.656

−1.63 (2.88)

—

256.00

v3 16 B

2,000 × 128

64.010

5.666

−2.09 (3.33)

3.94

255.99

v3 64 B

120 × 512

64.060

5.660

−2.68 (3.72)

—

255.95

Both pass. No position and no SAC cell was outside random expectation. SAC flip probabilities ranged 0.456–0.547 (v2) and 0.456–0.544 (v3). Message-level avalanche can’t tell v2 and v3 apart, because v2’s weak last step only shows when the difference enters at the very end (section H and the finalization sizing). Run time: 12.3 minutes, 4 workers, load ≈ 8.5–9.6.

J. v3 retest (RL_VERSION=v3, 6 Oct 2026, evidence phase3-runs/*-v3-*.json, log v3-retest.log)

Test

v2

v3

Verdict

Bias χ² (128 ± 16), max |z|

104.3, 2.24

148.8, 2.32

both pass

Single-bit differentials (7 positions × 4,000)

all distinct

all distinct

both pass

Truncated collisions, worst |z|

2.17

1.94

both pass

Linear masks max |z| (≈4.34) / fold-position masks

3.51 / 2.31

3.64 / 2.45

both pass

Fold cancellation: 7 classes, rank

full

full; mins 40–44

both pass

Avalanche + SAC (section I)

pass

pass

both pass

SAT preimage RL-4 (8-bit out)

14.3 s

50.1 s

v3 harder for z3

SAT collision RL-4 / preimage RL-6

timeout

timeout

—

Last-step difference, folded min (random ≈ 39)

24

41 (finalization sizing, k = 8)

fixed in v3

The last-step-only line in the v3 fold-cancellation log still reads 21. That’s expected: it exercises the raw step function, which v3 doesn’t change. v3’s protection is the 8 steps that follow, measured by finalization_sizing.py. The step-level tests (E and G2) also apply unchanged to v3.

Conclusion: v3 matches v2 on every message-level test and removes the one measured weakness. v3 is the current version.

K. Real truncated-collision attempt on v3: 40 to 64 bits (6 Oct 2026)

Method. rl_search.c is an independent C implementation of the spec. It matches all v2 and v3 vectors, and 1,000 random messages (500 per version) match the Python reference. It runs a van Oorschot–Wiener parallel collision search (distinguished points, 4 threads, 1.14 M v3 hashes/s) on top n bits of v3(PREFIX ‖ BE64(x)), with a fresh random 4-byte prefix per search. collision_attempt.py drives it and re-verifies every pair in Python. Control: the identical search on SHA-256 (--control). The yardstick is the birthday bound for an ideal n-bit output, sqrt(π/2 · 2^n) ≈ 1.25 · 2^(n/2) hashes. Evidence: collision-attempt-20261006-200559.json, -200824.json, -201926.json.

Bits

v3 searches

v3 mean ratio

SHA-256 searches

SHA-256 mean ratio

40

16

0.85 ± 0.12

—

—

48

32

0.84 ± 0.11

24

0.94 ± 0.09

56

4

1.08 ± 0.24

4

0.69 ± 0.19

64

1

1.68 (9.06 G hashes, 2 h 13 min)

1

1.25 (6.72 G hashes, 5 min)

All

53

0.88 ± 0.08

29

0.92 ± 0.08

Every one of the 82 pairs was verified. Individual searches range from about 0.1× to 3× the mean, as expected for a single birthday search. v3 and SHA-256 are indistinguishable at every width. Pooled, they are 0.04 apart with a standard error of about 0.11. Both pooled means sit slightly below 1.0. That’s a property of the method and the sample, and the control shows it equally, so it isn’t specific to v3.

64-bit collision found (top 64 of 128 bits equal, full digests differ):

a = 0040df25fe25cd3ac96eec65   v3 = 5916199fd7b6b8d0 665cb3a0f4fed711
b = 0040df25cac7c07fd028985d   v3 = 5916199fd7b6b8d0 f805a08c4d2bd9a0

Reading. Up to 64 bits, finding truncated collisions in v3 costs what it costs for SHA-256: the ideal birthday amount. No shortcut was found. A full 128-bit collision at this cost would take about 2^64 hashes, roughly 500,000 years on this Mac. This bounds generic attacks. It does not rule out a structural attack that a cryptanalyst might design.

CPU reads during the runs

Time

1/5/15-min load

Notes

12:46:24

45.3 / 25.0 / 13.3

Full Phase 3 run (8 workers) overlapped briefly with the linear-mask and re-check tests

12:47:15

41.7 / 27.3 / 14.9

Full run finishing

12:47:40

34.8 / 26.5 / 14.8

After the run. kernel_task at ~158% CPU (macOS thermal management). WebKit pages ~180%. Art-v4.py ~11%

Fold-cancellation search: 4 workers, 1 min 44 s, load peaked about 5.7. SAT runs are single-threaded (≈100% of one core, about 8 minutes in total). pmset -g therm recorded no thermal or performance warning level throughout. The raw log is in phase3-runs/cpu-reads.log. All Phase 3 scripts now use 4 workers, not 8, to keep the fans down.

L. Grover test

L.1 Scope, evidence and reproduction

MEASURED — 7 October 2026. This experiment ran on a classical CPU, using NumPy 2.5.2 and Python 3.14.7, with one worker and numerical-library threads limited to one. No quantum hardware, Qiskit or new packages were used. Final run time: 25.97 seconds. The existing six folded/unfolded v3 vectors and nonce vector passed; digest_w(m,64) agreed with refracting_light_v3.digest(m) on those inputs.

MEASURED — primary evidence: final raw JSON, generated by grover_rl_test.py. The JSON includes the seed, source fingerprints using SHAKE256, per-iteration probabilities, output histograms, boundary controls, exhaustive lane checks, and counts for all 600 full-width step counters. This section’s measured claims refer to that file unless stated otherwise.

MEASURED — retained development evidence: grover-20261007T110320631961Z.json and grover-20261007T110505905806Z.json are earlier runs in phase3-runs/. Their resource bookkeeping is superseded by the final JSON: they incorrectly recorded a 43-byte address input; the second also recorded an 88-bit fixed tag. Their small-width simulation and circuit checks passed. These files are retained without modification. The final run uses the correct 12-byte tag, 44-byte input and 600 steps. My initial conversational correction of the handoff’s byte count was wrong; the handoff was correct.

MEASURED — reproduction command (creates a new timestamped JSON; does not append to this report):

cd '/Users/premise/Documents/ChatGPT/Ciphers and Chains'
python3 -B grover_rl_test.py --seed 20261007 --input-qubits 12

MEASURED — mutation scope. The only existing file edited for this task is this report, by appending section L. QTL, ART, the chain specification, reference hash code, PEER_REVIEW/, and prior experiment files were not modified.

L.2 A–B: equality oracle and Grover statevector simulation

MEASURED. For each word width w = 2, 3, 4, the input register has 12 qubits: integers 0 through 4095 encoded as exactly two big-endian bytes, with the upper four bits zero. The output widths are 4, 6 and 8 bits. Five distinct targets per width were sampled uniformly from the output range using seed 20261007; targets were not chosen by hashing a sampled input. Every input was hashed into a table, then rehashed against the reference. Every selected equality predicate and its phase-flip operation was checked over all 4096 inputs; applying the oracle twice restored the test vector.

MEASURED. This is a table oracle, not a gate-level circuit. The simulator explicitly evolves a normalized real statevector through phase inversion and diffusion. It records every iteration from zero through 2*ceil((pi/4)*sqrt(N/M)), covering twice the first-peak optimum. It also tests artificial marked sets with zero, one, half and all inputs marked. All assertions passed.

PROVEN — exact prediction for this ideal Grover procedure:

N = 2^n_in
M = number of marked inputs
θ = asin(sqrt(M/N))
P(success after k iterations) = sin²((2k+1)θ)
k_first_peak_real = π/(4θ) − 1/2
k_opt = nearest nonnegative integer to k_first_peak_real

PROVEN. Integer rounding may produce tied optima. The often-quoted (pi/4)*sqrt(N/M) is an asymptotic approximation, not the exact integer optimum; it omits both the arcsine and the half-iteration correction. For M=0 there is no successful search. These formulas assume a perfect oracle and no noise. See Boyer, Brassard, Høyer and Tapp, Tight bounds on quantum searching.

MEASURED — results, N = 4096 in every row. Targets are hexadecimal; iteration counts are decimal. Exact theoretical and measured peak probabilities agree within 2.23 × 10⁻¹⁵ over all recorded iterations.

w

Output bits

Target

M

Asymptotic k

Exact integer k

Measured best k

Measured peak

2

4

07

307

2.8688

2

2

96.644075%

2

4

0d

257

3.1355

3

3

95.994731%

2

4

0e

262

3.1054

3

3

95.278755%

2

4

05

244

3.2179

3

3

97.612715%

2

4

01

298

2.9118

2

2

95.846600%

3

6

25

51

7.0386

7

7

98.870713%

3

6

0c

73

5.8831

5

5

99.044650%

3

6

23

55

6.7778

6

6

99.628493%

3

6

10

60

6.4892

6

6

99.995814%

3

6

0d

68

6.0956

6

6

98.819085%

4

8

34

16

12.5664

12

12

99.994704%

4

8

f3

14

13.4340

13

13

99.992577%

4

8

02

15

12.9785

12

12

99.675596%

4

8

61

16

12.5664

12

12

99.994704%

4

8

3d

18

11.8477

11

11

99.797831%

PROVEN — interpretation limit. Given N and M, ordinary Grover with a truth-table phase oracle has this same curve for every marked set, including highly structured functions. Agreement therefore cannot establish that RL behaves as a generic cryptographic function, or rule out a different quantum algorithm exploiting its arithmetic. Zalka’s optimality result is a black-box query result, not a proof that a specific RL circuit has no shortcut. See Grover’s quantum searching algorithm is optimal.

MEASURED. No trial used fewer iterations than its exact first-peak prediction. This passes the oracle/simulator consistency criterion. It does not pass a cryptographic “no structural quantum speedup” criterion; this experiment cannot test that proposition. Counts M vary across targets, and more preimages legitimately require fewer iterations. Full output histograms are retained for inspection; no random-function distribution claim is inferred from five targets.

MEASURED — scaling caveats from the source. At toy widths, constants change effective behavior: 65537*v becomes v mod 2^w for w <= 16. At w=2, the reference’s initial velocity values 5 and 7 exceed a two-bit word before the first step. This experiment deliberately preserves that reference behavior. Neither the toy digests nor their target multiplicities establish full-width security.

L.3 C: reversible step construction and verification

MEASURED. The new script constructs explicit X, CNOT and Toffoli gates for the step arithmetic. It uses fresh zero ancillas, ripple-carry additions, shifts plus additions for multiplication by 257, 17 and 65537, a signed z <= 0 condition, the outward clamp, quotient/remainder extraction, rotations as wire permutations, and XORs. The signed intermediate width is W=w+16. It computes outputs, copies them to a new output register and reverses the arithmetic gates, restoring scratch to zero and preserving inputs.

PROVEN — construction scope. This embeds the function as (state, bit, 0) → (state, bit, step(state,bit)); it does not assume the original state transition is invertible. For 0 <= i < 600, w >= 2, the magnitude bound 275*(2^w−1) + 16*(i+4) + 16 < 2^(W−1) prevents signed overflow in the clamp intermediates. Taking the low w remainder bits and the next w sign-extended bits supplies the floor-divmod quantities needed by the modular update, including negative values.

MEASURED. At w=4, each output lane depends on four four-bit words and one input bit. For every lane at counters i=0,64,599, all 2^17 = 131,072 local assignments were compared against phase3_sat_attacks.step: 1,572,864 lane comparisons, with zero mismatches. The reversible gate simulation also checked input preservation and zero scratch on every assignment. These were bit-parallel classical basis-state simulations, not quantum statevectors of thousands of qubits.

PROVEN — compositional coverage. Every one of the 2^33 possible full-step state/bit inputs projects into a checked local assignment for each lane. Since all lanes read the preserved old state and share the same bit, the local exhaustive results cover every full-state input at those three fixed counters. This is not enumeration or formal verification of every counter.

MEASURED. Integrated full-step wiring was additionally checked against the reference on 32 seeded random states and 18 boundary states for each of the six (w,i) combinations below, including the full width w=64. The boundary cases exercise zero/small/maximal words and both input bits. Full-width correctness is sampled, not exhaustive. All 300 integrated checks passed.

MEASURED / ESTIMATE — circuit counts. Qubit, Toffoli and logical gate-depth counts are obtained from the generated clean-step circuits. T-counts use the explicit 7 T/T† per exact Toffoli accounting convention; no Clifford+T compilation was executed. Depth counts X/CNOT/Toffoli layers with each Toffoli treated as one logical operation, using a conservative wire-dependency schedule. They are not physical clock cycles. The standard seven-T decomposition is described in Selinger, Quantum circuits of T-depth one.

w

Fixed counter i

Logical qubits

Toffoli gates

T-count estimate

Logical gate depth

C

2

4

07

307

2.8688

2

4

0d

257

3.1355

3

2

4

0e

262

3.1054

3

2

4

05

244

3.2179

3

2

4

01

298

2.9118

2

3

6

25

51

7.0386

7

3

6

0c

73

5.8831

5

3

6

23

55

6.7778

6

3

6

10

60

6.4892

6

3

6

0d

68

6.0956

6

4

8

34

16

12.5664

12

4

8

f3

14

13.4340

13

4

8

02

15

12.9785

12

4

8

61

16

12.5664

12

4

8

3d

18

11.8477

11

ESTIMATE. The JSON also records the loose bound T-depth <= 7 * logical_gate_depth. This deliberately unoptimized construction uses substantial scratch and history; its cost is not a lower bound on an attack. Circuit optimization, reversible pebbling and exploiting fixed input bits can change the time/space tradeoff.

L.4 D: full RL v3-128 preimage resource estimate

PROVEN — message length and step count. The literal b'E46-address\0' has 12 bytes. With a 32-byte candidate public key the message has 44 bytes. Its 74-byte record plus one finalization byte gives:

steps = 8*(8 + 6*ceil(message_bytes/4)) + 8
      = 8*(8 + 6*11) + 8 = 600

MEASURED. The estimator generated and counted the same reversible construction at w=64 for every counter 0 through 599, rather than multiplying the four-bit cost by sixteen. Arithmetic circuits were counted for all counters; basis-state correctness was tested at the three counters described above. One forward traversal of the 600 clean steps totals 47,813,472 T gates under the stated convention.

ESTIMATE — RL-only oracle architecture. Retain 601 512-bit state registers, reuse the largest arithmetic scratch region between steps, compute the fold, apply a target-equality phase, then uncompute the fold and entire state history. Thus one phase-oracle query has two hash traversals, compute and uncompute. One Grover iteration has one such phase oracle plus a diffuser; “two traversals” must not be counted again as two Grover iterations.

ESTIMATE — surrounding work. Fixed-length record encoding is affine over GF(2), so it can use X/CNOT gates with no T gates. The envelope budgets at most 600*(352+1) = 211,800 sequential X/CNOT layers per encoding traversal, 512 CNOTs per fold, 254 Toffolis for the 128-bit equality phase, and 510 Toffolis for a diffuser on the 256 variable candidate bits. It reserves 352 message qubits (96 fixed tag bits), 600 record bits, 128 fold bits and 255 predicate ancillas in addition to state history and arithmetic scratch. These surrounding stages have not been integrated into a full gate-level oracle or simulated; they are explicit conservative construction budgets.

ESTIMATE — unoptimized RL-only logical resources:

Quantity

Estimate / construction bound

Logical qubits, including surrounding work

<= 324,076

T-count per phase-oracle query

95,628,722

T-count per Grover iteration, including diffuser

95,632,292

X/CNOT/Toffoli logical depth per phase oracle

<= 5,475,571

Logical depth per iteration, including diffuser

<= 5,476,086

T-depth per iteration, loose bound

<= 35,360,164

Ideal RL-128 preimage iterations

1.4488038916 × 10^19

T-count for that idealized search

1.3855243681 × 10^27

Logical depth for that idealized search

<= 7.9337747076 × 10^25

HYPOTHESIS. Applying the ideal 128-bit preimage model to RL assumes an approximately 2^-128 marked fraction in a sufficiently large candidate domain and no better structural attack. Under that assumption, near-peak success takes (pi/4)*2^64 Grover iterations. The experiment does not establish that assumption. The quantum cost here concerns matching a chosen digest, not finding any collision.

PROVEN / ESTIMATE. In the black-box search model, splitting work across P quantum machines improves elapsed query depth by approximately sqrt(P), not P, until small-domain limits matter. Thus the RL idealization becomes roughly (pi/4)*2^64/sqrt(P) iterations per machine, while aggregate work grows. No conversion to seconds or years is justified without a physical gate schedule, error model, error-correction overhead, connectivity and magic-state factory budget. The logical figures above include none of these physical costs.

L.5 Address conclusion and limits

MEASURED — current design. CHAIN-DESIGN.md sections 2a, S12 and S15 select the single-owner payload RL_v3_128(tag || public_key) || SHAKE256_128(tag || public_key).

HYPOTHESIS / ESTIMATE. If this joint function behaves like an ideal 256-bit function over the relevant candidate domain, matching both halves on the same input has density approximately 2^-256, giving about (pi/4)*2^128 = 2.6726 × 10^38 Grover iterations. The RL half alone has the idealized 2^64 iteration scale. Requiring a simultaneous match is the intended reason for concatenation; independently finding a match for each half on different inputs does not produce a valid address preimage.

PROVEN — limitation of the inference. Concatenating two 128-bit outputs does not itself prove 128-bit quantum preimage security. For example, if the RL half were constant, the combined function would impose only the remaining 128-bit condition. Joint structure and attack algorithms must be reviewed. The toy Grover run cannot justify marking the whole address construction “quantum safe.”

ESTIMATE — correction for peer review. Under the ideal joint-function model, generic classical preimage work is approximately 2^256, and generic quantum preimage work approximately 2^128. The current chain document’s “about 2^128 classically and about 2^128 with Grover” is not the correct generic preimage comparison for an ideal 256-bit payload. This is an annotation only: the chain file was not edited.

HYPOTHESIS / ESTIMATE — wallet attack scope. An arbitrary byte-string preimage is not necessarily a usable public key with a known signing secret. A spend-capable attack may need coherent key generation and additional validity checks, or may attack the signature/key-generation scheme directly. Overall wallet security can therefore be limited by other components. These tests neither recover keys nor certify funds safe.

ESTIMATE — incomplete resource work. The resource table above is for the RL-only condition. A complete concatenated-address oracle must also implement SHAKE256 and, for a spend-capable search, any necessary key-generation/validity circuit. Those costs were not built or measured; their complete resource fields in the JSON are explicitly null. Multiplying the RL-only per-iteration cost by 2^128 would not be a complete address attack estimate.

L.6 E: mining and difficulty, in plain language

PROVEN — ideal search scaling. If a fraction 2^-b of candidate headers qualifies, independent classical attempts need 2^b hashes on average. An ideal coherent Grover search reaches high success in roughly (pi/4)*2^(b/2) iterations, provided the search space is large enough. A submitted valid header can still be checked with one ordinary hash. This does not make a Grover iteration as cheap as a classical hash.

ESTIMATE — protocol effect. Applying the ASERT rule in CHAIN-DESIGN.md, sustained faster block production makes the target harder, pushing aggregate block timing back toward its schedule. This is the sense in which difficulty adjustment can absorb additional quantum mining throughput. It does not remove the quantum miner’s relative advantage or guarantee decentralization. A lone capable quantum miner could dominate rewards or reorganizations despite restored block timing. The ASERT mechanism is documented by Bitcoin Cash Node’s upgrade specification; the quantum-mining consequence here is an inference, not a tested ASERT result.

ESTIMATE — operational limits. New blocks can invalidate the current search, and noise, coherent-oracle cost, restarts and competing miners affect achievable advantage. No mining network or ASERT response simulation was run. In particular, the handoff’s 600-step address example is not the mining header: the proposed 120-byte header requires 1,512 RL steps. Section 2b specifies comparing the 128-bit digest against the upper 128 bits of a 256-bit target; section 2’s shorter wording should be reconciled during consensus review. No consensus behavior was changed here.

L.7 Review disposition

MEASURED — passed: published v3 vectors; exhaustive selected table-oracle predicates; all 15 small-statevector Grover curves; boundary controls; compositional exhaustive four-bit step checks at three counters; integrated sampled/boundary four- and 64-bit step checks.

HYPOTHESIS — still unestablished: RL’s full-width preimage strength, resistance to structural quantum attacks, and the combined address’s claimed quantum security level. These are not “passed” by this test.

ESTIMATE — remaining work: independently review the reversible arithmetic and joint-hash assumptions; compile and verify a complete oracle including surrounding logic and SHAKE256; include key-generation costs for spend-capable attacks; optimize and evaluate fault-tolerant physical resources before quoting real attack times. No full-size quantum search, physical quantum run, complete address-oracle simulation or security proof was performed.

M. WO-14: Rad and commitment-stamp Grover re-run

M.1 Scope and reproducibility

MEASURED — 8 October 2026. The run used the release-built C++ librl through build/release/rl_cli, Python 3.14.7, NumPy 2.5.2, one worker, and seed 20261007. No packages were installed. Reproduce it with:

python3 -B grover_rl_test.py --seed 20261007 --input-qubits 12 --rl-cli build/release/rl_cli

MEASURED — primary evidence: WO-14 raw JSON. It records source fingerprints, the per-iteration toy curves, library cross-check counts, resource estimates, and runtime. The run completed in 36.16 seconds.

MEASURED — C++ library cross-check. At the production word width w=64, e46::rl::rl_v3_128 matched both phase3_sat_attacks.digest_w(w=64) and refracting_light_v3.digest for 4,133 messages: all 4,096 two-byte encodings of a 12-bit input, the published v3 and nonce vectors, selected lengths and seeded messages, and 16 recorded WO-08 commitment-stamp inputs. There were zero mismatches. All 16 recorded stamp hashes also retained their required 18 leading zero bits. The message set is a differential check, not an exhaustive check of all 64-bit or arbitrary-length inputs.

MEASURED — toy oracle scope. As resolved by owner decision 14, w=2,3,4 use Python-only width-scaled table oracles. The script does not claim the production C++ API supports those widths. For each toy width it exhaustively built the 4,096-entry digest table, exercised five equality targets, and checked the corresponding Grover probability at every recorded iteration against the exact formula in section L. It also ran a given-Rad-prefix predicate and a toy zero-prefix stamp predicate. The actual M counts are recorded in the JSON. These table-oracle checks validate Grover mechanics for the selected marked sets; they do not measure function-independent security or provide a gate-level implementation of the predicates.

M.2 Quantum query-cost estimates at the WO-08 base tier

The WO-08 base threshold is Z_base = 41. The next trigger tier is Z=48. The estimates below use ideal random-function oracle models; they are query scales, not quantum executions or physical resource counts.

Task

Z=41

Z=48

Interpretation

Find any collision (BHT scale)

2^(41/3) ≈ 13,004 queries (2^13.67)

2^16 = 65,536 queries

BHT is the generic quantum collision-search scale over a sufficiently large ideal-function domain.

Match a chosen Rad’s Z-bit prefix (Grover scale)

2^(41/2) ≈ 1,482,910 queries; first-peak approximation (π/4)2^(41/2) ≈ 1,164,675

2^24 = 16,777,216; first-peak approximation ≈ 13,176,795

Searches for any input whose prefix equals the given Rad’s prefix, at ideal marked density 2^-Z. This is not the cost of checking or searching only the already-known finite WO-08 candidate set.

Find an 18-bit commitment stamp (Grover scale)

2^9 = 512 queries; first-peak approximation ≈ 402

Same 18-bit predicate

Ideal zero-prefix oracle; excludes oracle construction and fault-tolerant overhead.

ESTIMATE — bounded-puzzle caveat. WO-08 samples S=2^20 candidate indices. Under an ideal random-function model, the expected number of pairs at Z=41 is S(S−1)/2^42 ≈ 0.25, with Poisson approximation about 22.1% probability of any pair. At Z=48, the expected number is about 0.00195 pairs, with about 0.195% probability of any pair. A finite puzzle may contain no collision, so the BHT scale above must not be read as a promise that each bounded puzzle has a solution. Once a particular Rad is already known, the generic Grover row instead describes searching a sufficiently large domain for its chosen prefix.

ESTIMATE — limitations. These expressions omit oracle synthesis, reversible SHAKE or message framing where applicable, error correction, machine parallelism, restarts, and wall-clock conversion. The classical table oracle is supplied by construction and cannot test whether RL has exploitable arithmetic structure. No full-size quantum search, physical quantum run, or security proof was performed.

M.3 Validation and disposition

MEASURED — passed: release and sanitizer CMake presets built; all 16 CTest cases passed in each configuration; the production-width differential check had zero mismatches; toy equality, Rad-prefix and stamp-prefix Grover probabilities matched the exact simulator formula; and the existing reversible-circuit checks and resource estimator completed in the same run. One development assertion incorrectly compared the first rounded peak with every later peak across two oscillation periods; it was corrected to retain per-iteration formula checks without imposing that false global-maximum condition. The final recorded run passed.

HYPOTHESIS — unestablished: RL’s full-width collision resistance, generic behavior under coherent quantum access, and resistance to structural quantum attacks. Grover/BHT query scaling is conditional on the black-box model and is not a cryptographic security result.