RedTail polynomials ternary and clamp research

Research record dated 5 October 2026. Participants: Michael Cohee and Codex. This document preserves the current conversation, source findings, a limited collision experiment, and Michael’s proposed research direction. It is an exploratory record, not a specification or a claim of cryptographic security. Existing project code and papers were not modified.

Latest steering: ABS(x) is deferred at Michael’s request. Current work focuses on collisions with signed symbols left intact. Earlier ABS discussion below is historical. See the active collision trials.

Research question

Can RedTail-X, polynomial operations, a bounded clamp, and ternary symbols support a new hash or cipher suitable for a Bitcoin-like system? Michael proposes investigating what happens when a construction escapes its bounds and changes its tail or transformation in response.

The initial analysis ruled out directly using the existing linear parity reduction as a secure hash. Michael’s subsequent proposal introduces a clamp-and-escape mechanism that the earlier experiment did not implement. Its behavior remains open until the mechanism is specified.

Michael’s direction in his own words

The following is preserved verbatim from the steering message, including provisional terminology and spelling:

One: Red Tail, Polynomials and Ternary ( -1, 0, and 1 ) as bits or nits. Two: When looking at the CLAMP or part of RedTail-X when applied to Reed Solomon 4+2, we do not get collisions. The Clamp avoids them. The formula would need a modification to throw, in a good way, a new cipher every time we escape the clamp. RS4+2 stays inside the clamp of RedTail-X. It is claming +INF and -INF to the bound limits of RS42. Without it we get binary collision. Look at Shannon Directly, Shannon Entropy. RedTail was made to stop one type of entropy. We can bond it to another that will break the formla requiring modfication to the tail for those collisions and hashes possibly. This is hard and I need to read more here, but thats my direction and navigation right now.

Working interpretation, subject to Michael’s refinement:

  1. Explore signed ternary symbols {-1, 0, +1} and polynomial representations.

  2. Identify a bounded region in which recovery or uniqueness holds.

  3. Define escape from that region precisely.

  4. On escape, modify the tail or select a new transformation while retaining a reproducible rule.

  5. Investigate whether coupling different measures of uncertainty or different processes creates useful cryptographic behavior.

The statements that the clamp prevents collisions and bounds infinities to RS42 are research hypotheses. The inspected code establishes a numerical entropy clamp, but does not establish that broader mechanism.

What the existing sources establish

The architectural-engineering and Layered-Eval manuscripts are presentations of substantially the same exact-file continuity case study. They distinguish preservation of ingested bytes, completion of share repair, and correctness of a repaired application view. BRMR reconstructs missing DXF owner handles from surviving references and stores repair views separately. Its Merkle commitments depend on a cryptographic hash.

RedTail-X v1.1 defines final-share sizing for systematic Reed–Solomon 4+2 encoding:

s = ceil(t / 4)
tail_bytes = 6s
c = Gd, with G = [I4; C] over GF(256)

Here t is the partial stripe length. The four data shares preserve input bytes with at most three padding bytes; two parity shares enable recovery from any two known share erasures. The manifest preserves original length and hashes. Reported emitted-byte savings were 17.93% against the prototype’s original full-padding baseline on 213 PDFs. That is a storage measurement, not a hash security measurement.

At fixed input length, the complete systematic codeword is injective because it includes the data. Across variable lengths, original length must also be retained: zero padding can otherwise make distinct strings such as a byte followed by different numbers of zero bytes indistinguishable. The full representation, including the length, is the relevant object.

The earlier response’s conclusion should therefore be read narrowly: the existing formula does not supply a secure SHA-256 replacement. It did not demonstrate collisions in the complete length-bearing RedTail-X representation or disprove every future nonlinear construction inspired by it.

The clamp found in the implementation

In representation.rs, the production entropy estimate computes:

p = count / sample_length
contribution = -p * log2(clamp(p, 1e-12, 1))

The original p remains the weight. An absent symbol therefore contributes zero while the logarithm receives a positive argument. For positive probabilities at least 1e-12 this agrees with the ordinary entropy term. Smaller positive probabilities would be numerically approximated. With the configured 8192-byte sample, that lower clamp affects only zero-count symbols.

This bounds the logarithm’s input; it does not clamp an RS codeword, preserve a secret key, or supply collision resistance. GF(256) arithmetic itself has no positive or negative infinity. Its byte symbols are field elements, not a saturated interval of real numbers.

A separate entropy_example.rs clamps the probability in both places. The report excludes that estimator approach because independently clamped probabilities need not sum to one. It must not be conflated with the production logarithm guard or Michael’s proposed escape rule.

If a proposed clamp means ordinary saturation, C(x) = min(b, max(a, x)), all x above b share one output and all x below a share another. Saturation alone loses distinctions. An escape marker plus retained residual could preserve them, but its complete encoding and size would need to be counted. A different meaning of clamp needs its own definition.

Shannon directly

Shannon’s original 1948 paper, A Mathematical Theory of Communication, gives the uncertainty measure H = -K sum(p log p). With base-two logarithms and K = 1, the units are bits. For an alphabet of m symbols, the maximum is log2(m), attained by the uniform distribution. Read Section 6 on uncertainty, Section 7 on source entropy, and Sections 12–13 on equivocation and noisy channels.

Applied here, a uniform ternary symbol carries log2(3), approximately 1.585 bits. Byte-symbol entropy has a maximum of eight bits per byte. Zero-probability terms use the limiting value zero. A large observed histogram entropy does not establish a hard inversion problem.

Analytical interpretation for this project: distinguish source uncertainty, uncertainty about missing data given survivors, numerical singularities, and an adversary’s uncertainty about a secret. These are different quantities. RedTail-X’s reported contribution removes padding, while the RS redundancy enables recovery from erasures. Saying it “stops entropy” needs a declared random variable, observer, and measurement.

Ternary and polynomial choices

The conventional name for a ternary digit is a trit. The proposed {-1, 0, +1} alphabet is balanced ternary when used with positional base-three arithmetic. “Nit” is retained above as Michael’s provisional terminology, not an adopted unit definition.

One possible representation to investigate is:

P(z) = sum(a_i z^i), with a_i in {-1, 0, +1}

This alone does not define a hash or cipher. Formal polynomials, integer evaluation, modular evaluation, and finite-field evaluation produce different mappings. Revision after Michael’s next message: investigate signed ternary independently of the byte-field arithmetic, with explicit encoding when passing through RS. Binary storage can represent negative integers using a signed encoding; a single bit has only two states. The specific identity -1 = +1 in GF(2^8) is a property of that field, not a prohibition on storing negative values in binary. Distinct byte labels can carry three trit states without performing signed arithmetic in the field. ABS folding is now a separate, explicitly conditional experiment, with and without sign retention.

If exponents encode position, sorting polynomial terms need not lose position. If position, length, zero terms, or small coefficients are discarded without recoverable metadata, distinct source representations can merge. The current symbolic canonicalizer intentionally merges equivalent terms and uses floating-point coefficients; it is not an exact byte identity primitive.

Collision experiment and its limits

The initial investigation reimplemented the custom parity equations from polynomial_code.rs in a small Python calculation. It did not execute the Rust encoder or implement the new clamp proposal.

For each byte position:

P = d0 XOR d1 XOR d2 XOR d3
Q = d0 XOR (2*d1) XOR (3*d2) XOR (4*d3)

Multiplication uses GF(256) with reduction polynomial 0x11d. Both input columns (0,0,0,0) and (1,2,3,0) produce P = Q = 0. Repeating the column construction across four 16-byte shares gives two distinct 64-byte messages with identical 32-byte parity output. Repeating across 20-byte shares also gives an 80-byte example with equal 40-byte parity. All 256 field multiples of the difference column were checked and preserved parity. SHA-256 distinguished the tested message pairs.

The actual custom Rust encoder uses fixed 256 KiB shares; the small calculation applies its byte-column algebra at smaller illustrative widths. The production library uses a different generator matrix. The particular vector above is not claimed to be its kernel vector.

The general conclusion concerns a proposed linear digest h(x) = Ax with more input dimensions than output dimensions: a nonzero kernel vector delta gives h(x + delta) = h(x). Such vectors can be found by linear algebra. Full codeword recovery, a parity-only digest, and a future nonlinear clamp construction are distinct cases.

Cipher hash and escape semantics

A cipher requires a key and a decryption rule. A hash requires a reproducible digest of an input. Bitcoin-style proof of work additionally requires public verification and resistance to shortcuts for reaching a target. A new output, key, nonce, state, and algorithm are different meanings of “a new cipher every time”; the next specification should identify which one is intended.

For an escape mechanism, record:

  • The state, its arithmetic, its bounds, and the exact escape predicate.

  • What an escape changes: tail bytes, polynomial coefficients, key, nonce, or transformation identifier.

  • Whether the change is deterministic or uses external randomness.

  • What the verifier or decryptor receives, including any residual and transition history.

  • What remains secret and what security property is being claimed.

  • How identical inputs are treated and whether output length stays fixed.

A finite fixed-length digest cannot be globally collision-free for an unbounded message domain. The cryptographic goal is computational difficulty of finding collisions. A bounded injective representation can avoid collisions within its domain, but that fact alone does not establish one-wayness. Adding fresh randomness can change the output, but does not establish either property without a complete construction.

Next research steps

  1. Write the clamp as an exact mapping and identify whether it is the existing logarithm guard, a domain restriction, saturation with residuals, or a new state transition.

  2. Select ternary arithmetic and give worked examples for -1, 0, +1, all boundaries, and the first two escapes.

  3. Define the entropy quantities to be coupled and the measurable effect expected from their coupling.

  4. Enumerate a small complete domain and test injectivity of the full representation, including length and escape metadata. Separately test any projected digest.

  5. If the goal is hashing, define all rounds and finalization before avalanche, collision, preimage, and proof-of-work shortcut analysis. If the goal is encryption, define keys, decryption, and authentication first.

At the time of the initial record these steps were proposals. The subsequent trial record now reports exhaustive small-domain experiments for signed ternary, conditional ABS folding, boundary escape, tail retention, and deliberate collisions. The exact intended clamp remains unspecified. No new cryptographic cipher, mining algorithm, or proof of security was produced.

Source register

Local papers reviewed in the initial investigation:

Additional implementation evidence:

External references consulted during the conversation:

  • Shannon’s original paper linked above; the entropy discussion was checked directly in this turn.

  • Bitcoin block header specification: 80-byte headers and target comparison.

  • NIST hash functions: cryptographic hash properties.

  • NIST FIPS 197: AES combines nonlinear substitution with linear mixing and rounds; this is background, not validation of RedTail-X.

  • NIST FIPS 202: standardized SHA-3 constructions.

Source code links describe the local files inspected on this date. They are not pinned archival revisions. The reported corpus benchmarks were read from the papers and were not rerun. The parity calculation was a limited mathematical reproduction, not a full cryptographic audit.