Author: Anthony Garces

  • The Security Check That Checked Nothing

    The Security Check That Checked Nothing

    TL;DR: A security check that treats a missing required file as nothing to check can let attacker-chosen policy decide the result. Test absence, not only tampering. The parent slice log keeps the failure on record.

    Your security check catches a changed file. What does it do when the file is not there at all? That question found a hole in Ranex after a completed control had already been marked done.

    In this note

    The check returned without checking

    The bug was not a scanner that missed a clever payload. It was a trust-root check that returned when the evaluated commit did not carry the path it was asked to inspect.

    Ranex evaluates evidence against a gate catalog and a producer keyring. Those files are policy: the catalog says which claims a verdict needs; the keyring says which producers can supply evidence. A working-tree edit to either committed file was caught. But refuse_uncommitted_trust_root had an early return for a path the evaluated ref did not carry. It had compared no bytes and allowed the run to continue.

    ADR-002 records the three reproduced routes: an uncommitted catalog named with --gate-catalog, a keyring hidden by a committed .gitignore and named with --producers, and a committed symlink whose reviewed name did not contain the bytes ultimately read. None needed key theft or a forged signature. Each reached the same return.

    The lesson is smaller than the incident and more useful. A control can work perfectly for the input you expected while doing nothing for the input an attacker gets to choose.

    Missing is a security state

    A missing required input must produce a decision. If it reaches a default branch, it has already become policy.

    It is tempting to think of absence as an operational nuisance: a file was not provisioned, a setting was omitted, a report was not generated. That is only true when the system owns the omission. When the party being measured can name the path, omit the field, or redirect the lookup, absence is part of the attack surface.

    Ranex already had the right general rule: absence blocks. No evidence is a FAIL, never a default and never a skip. The trust root was the exception, even though every later check depends on it. The function was written for an operator selecting a path. Its caller also allowed the measured party to select that path. The security boundary had changed; the branch had not.

    This is different from the SLICE-004 lesson in 59 Refusals, Zero Tests. That post is about measuring whether named refusal paths execute at all. This incident is narrower: one executed policy check treated an attacker-chosen missing trust root as permission to proceed.

    Look for this shape outside security tooling too. An authorization rule that skips when a role is absent. A deployment gate that reads a missing scan report as zero findings. A schema validator that makes a field optional without deciding whether the operation should be allowed without it. “Nothing to validate” is never a neutral answer until you have decided who benefits from it.

    Test your own security gates

    You can find this class of bug without adopting Ranex. Start at a check that decides whether work can proceed and ask what every escape route means.

    • List every early return, continue, fallback, and empty collection in the check.
    • For each required input, test absent, unreadable, malformed, redirected, and present-but-wrong states separately.
    • Ask whether the caller, the operator, or the measured party controls the file name or field. A safe default for one can be an exploit for another.
    • Ask the source of truth about the name supplied, before resolution can turn that name into a different object.
    • Keep one ordinary success test beside the refusal tests. A fix that refuses every path is not a security control; it is an outage.

    That last test matters. Security tests often prove only that bad inputs fail. A control that makes every input fail will look excellent until somebody tries to use it.

    The fix and its limit

    Ranex now refuses a trust-root path the evaluated commit does not carry. It compares the bytes stored under the supplied repository name with the bytes it reads, then passes those committed bytes to the loader rather than reopening a path. The same rule applies to the producer keyring used by run.

    The decision record says the ordinary committed catalog and keyring must still reach PASS. Its security test covers five refusal cases and one normal path; every refusal exits with no verdict. The record also reports that the check-and-reopen window was removed by parsing the committed bytes, and that the observed file opens fell from three and two to one each. Those are useful measurements, not a claim that review can judge whether a committed gate asks for enough.

    That limit remains. A reviewed catalog can require too little, and this control cannot tell. The point is not that files in git become wise. The point is that the bytes whose policy you reviewed are the bytes that decide the verdict.

    Ranex is pre-release. This is one repaired control in a kernel with a working verdict path, not a finished security product.

    Questions people actually ask

    Use these answers when a required security input is missing rather than merely incorrect.

    What is a trust root in a security system?

    A trust root is the committed file whose bytes decide whether later evidence counts. In Ranex, the gate catalog and producer keyring are trust roots because they decide every verdict.

    What does fail closed mean for a missing security input?

    Fail closed means Ranex refuses when a required trust-root path is absent instead of continuing with unchecked bytes. A missing input is a named refusal, not an empty success.

    How did Ranex find the trust-root bug?

    Ranex found the bug while auditing SLICE-003 after SLICE-002 had closed: a path absent from the evaluated commit made the trust-root check return without comparing bytes.

    Is Ranex ready to use today?

    No. Ranex is pre-release; its README describes a working verdict path with much of the surrounding product still unbuilt.

    Make absence loud

    Pick one security check you own. Delete or redirect the input it relies on in a disposable test. Then make the expected refusal specific: name the input, stop the operation, and preserve a normal success case.

    Do not settle for a check that catches the wrong value. Make it tell you what happens when there is no value to check.

    Try it. Break it. Tell me what broke.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • A Ranex Gate Verdict Asserts Four Things — and Refuses Four More

    A Ranex Gate Verdict Asserts Four Things — and Refuses Four More

    TL;DR: You cannot act on a green light until you can name what it asserted and what it refused. A Ranex gate verdict asserts four things and refuses four more. The accountability apparatus is built on claims that can survive inspection.

    You see green and want relief. That is human. But before you merge, deploy, or let an agent continue, ask the harder question: what did that verdict actually establish, and what did it stay silent about?

    This one is written for the person who has to build the answer. If your problem is the other one — somebody told you “it’s tested” and you need the questions to ask back — the non-technical version is over on Anito. Different reader, different job.

    In this note

    The four things a verdict asserts

    A verdict is a sentence with a fixed shape, not a feeling about the code. The README’s “What a passing build actually proves” section gives the exact shape:

    Every behavior on the graph the owner approved has at least one executable test. Every test ran. Every test passed. Here is the evidence, pinned to this exact code digest.

    Count the assertions. Approved-graph behavior has executable coverage. The tests ran. The tests passed. The evidence is bound to one exact code digest. That is the whole entitlement — not that the code is wonderful, not that the owner got every requirement right, not that the system is safe in every way you care about.

    That boundary is the value. A claim that is smaller than your hope can still be checked. A claim inflated until it means “everything is fine” cannot. When a check fails, a narrow claim also gives you somewhere real to investigate: the graph, the test, the execution, or the evidence binding. Four assertions, four places to look.

    The last of the four carries the most weight and gets the least attention. Evidence has to attach to a subject, and here the subject is a digest — which is why subject-bound evidence matters. The proof must refer to the code under judgment, not a nearby commit. And it is why no self-approval matters. The producer of evidence cannot be the approver who makes it count. Both make the passing sentence harder to fake.

    What the verdict refuses falls into four buckets. It does not grade the graph. It says nothing off the graph. It says nothing about non-functional properties without their own gates. And it never upgrades conformance into correctness. The rest of this note takes them in order.

    Refusal one: it does not grade the graph

    A graph can faithfully describe the wrong thing. The README leaves that judgment with the person who owns the target, who must use the thing and decide whether the graph was right. A gate can establish conformance to an approved specification. It cannot choose the specification for its owner.

    That is not a weakness hidden in the machinery. It is the correct boundary. A checkout flow can meet every approved behavior and still omit the behavior customers needed. A billing rule can match the written rule and still be a bad business decision. Code cannot settle a product decision merely because the code ran.

    So make the owner approval meaningful. Put the behavior in words or a graph that the owner can challenge before implementation begins. Then let the verdict speak about conformance, not wisdom. Mixing those claims gives nobody a place to stand when the product is wrong.

    Refusals two and three: off the graph, off the functional axis

    Anything off the approved graph is unconstrained. If a behavior was not specified, the gate has no requirement to test and no basis to promise it. Absence of a requirement is absence of a guarantee.

    Performance, accessibility, and security belong in the same category. They are not free properties attached to a functional pass. The README says they need separate checkers unless you add gates for them. If no gate measures them, a green functional verdict stays silent about them.

    This should change your release conversation. Do not ask whether the product is “green.” Ask which propositions have verdicts behind them. Do you have a performance gate? An accessibility gate? A security gate? A requirement outside the approved graph? Each missing answer is not automatically a failure of the code. It is a missing claim.

    That distinction saves you from a bad ritual: treating one successful command as permission to stop thinking. A gate should block or establish a defined proposition. It should not become a ceremonial stamp that inherits every concern nobody wrote down.

    Refusal four, and the verdict your pipeline can defend

    The fourth refusal is the quiet one: a verdict never upgrades conformance into correctness. Take a passing check from your own pipeline and finish this sentence: “This result establishes that…” Then keep writing until the artifact, the command, and the requirement are visible.

    • Which owner-approved behavior does this test cover?
    • Did the executable test run and pass?
    • Which exact code digest did it observe?
    • Who produced the evidence, and who approved it?
    • What behavior sits outside this claim?
    • Which non-functional properties have separate gates?

    If the strongest honest sentence is narrow, keep it narrow. “Conformant to an approved specification” is deliverable. “Correct” is not a claim anybody can make. The second word feels stronger only because it hides the work still unmeasured.

    The current product does not claim the full story

    Ranex is pre-release. Its README describes a working verdict path, including subject-bound evidence, absence blocks, no self-approval, and a run-to-evaluate loop. The flow graph and scenario compilation that the full passing-claim shape depends on are designed, not built.

    That means the quotation is the documented claim shape, not permission to pretend every part of the broader loop exists today. Approver identity is unauthenticated, same-UID key theft remains open, and the journal does not detect rollback or truncation. A product that asks you to interrogate verdicts should expose its own limits first.

    Your move is smaller than a platform decision. Pick one verdict your pipeline emits. State its strongest defensible claim. State what it cannot establish. Then add the next gate only for the property you actually need to know.

    Questions people actually ask

    What does a passing Ranex gate prove?

    A passing Ranex gate proves that approved graph behavior has executable tests, every test ran and passed, and evidence is pinned to the exact code digest.

    Does a passing Ranex gate prove software is correct?

    A passing Ranex gate proves conformance to an approved specification, not that the specification or software is correct.

    Does a passing Ranex gate prove security or performance?

    A passing Ranex gate does not prove non-functional properties unless separate gates check them.

    Try it. Break it. Tell me what broke. Read the MIT-licensed repository, then write the one sentence your next verdict can defend.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • What Ed25519 Signing Actually Proves About a Test Result

    What Ed25519 Signing Actually Proves About a Test Result

    TL;DR: An Ed25519 signature proves that a registered private-key holder signed an evidence record; it does not prove who approved it or that the reported result matches reality. Read how the kernel works.

    A signed test result arrives and everyone relaxes. Stop there for a second. What did the signature actually establish, and what did everyone quietly add to the story?

    A signature authenticates a signing act, not a whole software claim. Ranex signs evidence with Ed25519 and verifies it against a committed public keyring before admitting a record. That is worth having. It is also easy to over-read, so the useful part is drawing the line around it.

    In this note

    A signature proves a narrow fact

    Ed25519 evidence signing proves that the holder of a registered private key signed the record. The verifier has public keys only, so it can verify that signature but cannot forge one.

    That is the core claim. Before an evidence record is admitted, Ranex verifies its Ed25519 signature against the keyring committed in the repository. A record without a valid signature from a registered producer does not get to become evidence merely because it has a plausible producer label. The trust root is visible for review rather than supplied from whatever working directory happens to be present.

    Pair that with the other boundaries and the result is more useful. Evidence is bound to a subject digest, so the same command against a different commit does not prove the current subject. The committed catalog also binds a claim to the argv that can satisfy it, preventing a signed record of an unrelated command from satisfying a named claim. The kernel evaluates the resulting evidence against its gate rather than treating a signature as a PASS button.

    Still, say the claim precisely. The signature proves a registered key holder signed a record. Subject binding and claim-to-command binding constrain what the record can satisfy. They do not turn a cryptographic signature into a general proof that every reported observation was true. The difference matters when a report says “tests passed” and your next question should be “what did this evidence actually observe?”

    Signing does not authenticate approval

    Signing does not prove who approved a result. In the current kernel, --approver is a plain string, so a producer can name anyone as the approver.

    This is a known gap, not a footnote. No-self-approval compares the producer and approver strings, and the rule still does its job as a structural check. But the identities represented by those strings are unauthenticated. A valid evidence signature proves nothing about the person who reviewed, rejected, or approved the work. Do not tell your team that a signed record has solved human authorization when the repository says it has not.

    Signing also does not prove the test result is true in the broad sense people want. The README names a sharp boundary: an approved, hash-correct dependency can still choose its own exit code. Dependency approval reduces hidden change; it cannot make third-party code truthful. A signature can establish which registered key signed a record of that outcome. It cannot force the dependency to report reality.

    That is not an argument against signatures. It is an argument against asking one control to carry every trust decision. You need frozen tests, subject binding, claim-to-command binding, separate approval, and checks appropriate to the property you care about. Performance, accessibility, and security are not free consequences of a passing functional gate. They need their own gates if you want a claim about them.

    The most dangerous sentence in an incident is “it was signed, so it was safe.” Signed by whom? Covering which subject? For which claim and command? Approved by which authenticated identity? If you cannot answer those questions, the signature may be real while the conclusion is far too large.

    Handle keys like a boundary

    Private signing keys belong outside the repository. Ranex keygen refuses to write a private key anywhere inside it, while the committed governance/producers.yaml keyring contains public halves only.

    The README describes that file as the trust root. It is a mapping from producer names to public Ed25519 values, and review of that committed mapping is the control on it. This is a practical distinction. A repository can carry the information a verifier needs without carrying the material that lets an attacker forge a producer’s signature from the tree.

    export RANEX_SIGNING_KEY=~/.config/ranex/worker.key
    PYTHONPATH=src uv run python -m ranex.cli.main keygen --producer worker

    That command path is not ceremony. A private key in the repository would make the committed trust root and the signing capability travel together, which defeats the separation. The delegated worker loop is also described as starting in an environment that refuses to hold the signing key. A worker can do work; it is not automatically trusted to sign its own success.

    Use the same lens with your own signing setup. Where does the private half live? Who can read it? Which process holds it during a run? Which reviewed public keys are permitted to verify? A key system is not only an algorithm choice. It is the path from private authority to the process that gets to use it.

    Keep the open risk visible

    Same-UID key theft remains open. Ranex says its run path still reads the signing key before spawning, so anything running as the same user can take it; the stated advice is to use a scoped, spend-limited key.

    That risk belongs beside the strength of public-key verification, not beneath it. The verifier cannot forge because it has only public keys. A process that can steal the private key is a different threat. Strong verification cannot undo a compromised signing environment after the fact.

    The repository’s Status section marks signed evidence as working today, lists unauthenticated approver identity and same-user key theft among known gaps, and calls the project pre-release. Those are the claims to carry into a real review. The complete governed product is not all built, and the boundaries here are deliberately smaller than a promise of total safety.

    For longer-lived evidence, pair signing with an append-only record. The hash-chained journal retains an ordered record and detects out-of-band row edits, though it has its own rollback and truncation gap. Controls stack because their failure modes differ.

    Find one signed report your team relies on. Write down what its key proves, what its subject binding proves, and what it cannot prove about the approver or the observed system. That is not distrust. That is the beginning of a claim you can defend.

    Questions people actually ask

    What does Ed25519 evidence signing prove?

    Ranex evidence signing proves that the holder of a registered private key signed the record; the verifier holds public keys and cannot forge it.

    Does an evidence signature prove who approved a result?

    Ranex does not authenticate approver identity because the approver value is a plain string, so a signature proves nothing about who approved it.

    Where does Ranex keep evidence signing keys?

    Ranex keeps private signing keys outside the repository because keygen refuses repository paths, while the committed keyring contains public halves only.

    Try it. Break it. Tell me what broke. Read the Ranex repository, then trace one signing key and one conclusion your own pipeline makes from it.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Subject-Bound Evidence: Why “It Passed” Means Nothing Without a Commit Digest

    Subject-Bound Evidence: Why “It Passed” Means Nothing Without a Commit Digest

    TL;DR: “It passed yesterday” is not evidence about the code in front of you. A command result must be bound to the exact commit it observed. The accountability apparatus needs a record tied to the artifact, not a comforting memory.

    You have seen the familiar green check on a branch that moved after the job ran. The label still says tests passed. The code is different. What does that result prove now?

    In this note

    A passing command describes one subject

    A test command ran against some bytes, with some inputs, at a particular point in time. Its result can speak about that subject. It cannot automatically speak about the next commit, the next checkout, or the branch name after somebody moved it.

    This is not paperwork. A small change can alter the behavior a test was meant to establish. A test result that continues to vouch for code it never observed turns green into a reusable mood. You need a claim, an artifact, and a link between them.

    Ranex calls that link subject-bound evidence. The README’s “What makes a verdict trustworthy” and “Status” sections describe the invariant directly: the same command against a different tree proves nothing about this one. A changed subject is a changed question.

    Think about a report handed from one engineer to another. “The suite passed” sounds complete until you ask: passed against which commit? If nobody can answer, the report has lost the one fact that tells you whether it applies. The test result may be genuine. It is still not evidence for the code you are about to merge.

    The runner observes the committed tree

    Binding a digest in a record is not enough if the command ran somewhere else. Your working tree can contain edits. Tool inputs can drift. A checkout can be adjacent to the commit you intended rather than the commit itself. “Close enough” is where false proof enters.

    Ranex materialises the committed tree from verified blobs before it runs the bound command. Each blob is checked against the object identifier carried by the commit tree. The command runs in an environment built from empty, with a toolchain pinned to directories the observed party cannot write.

    That is a strong claim, and it has a cost. The README says trees that need installed dependencies, or carry a symlink or submodule, cannot be observed through this path. The system refuses them rather than pretending an altered observation was the same subject.

    Do not confuse a checkout with an observation. A checkout is convenient. An observation is the exact tree and execution context your record says it measured. When an agent has influence over the code and the environment, that distinction is not academic.

    The project reached this shape after recorded false-PASS paths shared a root problem: the observed tree was not the tree HEAD named, and inputs were chosen by the party being measured. The current README says those paths are closed. The useful lesson for you is simpler: verify the subject before you let a result describe it.

    Stale evidence must stop counting

    Once the tree moves past the digest attached to an evidence record, that record stops satisfying the claim. Ranex turns the missing applicable evidence into FAIL. It does not keep yesterday’s test result alive because a branch name stayed familiar.

    That can feel strict when you changed one line. Good. The claim is not “a related version passed a command.” The claim is about this subject. If the evidence no longer matches, the correct result is that you need fresh evidence.

    This works with absence blocks. A required claim without satisfying evidence is FAIL, never a default and never a skip. Evidence that is absent and evidence that is stale arrive by different paths, but neither can support the claim under judgment.

    The principle also carries beyond agents. A human-written patch deserves the same treatment. Better models do not change it. Faster CI does not change it. If you cannot bind the result to the artifact, you cannot tell whether the result applies.

    Bind the claim before you trust the result

    Use this before you accept “it passed” in a review, a release note, or an agent summary.

    • What exact claim is this command meant to support?
    • Which commit digest is the subject?
    • Did the command observe that committed tree rather than a working copy?
    • Were the command inputs chosen outside the party being measured?
    • Does the record stop applying when the subject changes?
    • Does missing applicable evidence refuse the claim?

    Run that checklist on the next green build. If the answer to the digest question is a branch name, you have work to do. If a result remains valid after the code changes, ask what it is really asserting. Do not let the label do more work than the evidence can carry.

    The working path is narrow

    Ranex is pre-release. Subject-bound evidence and the run-to-evaluate path are listed as working today. The flow graph and scenario compilation that would feed a broader build workflow are designed, not built.

    That limit matters. This post is not a claim that every engineering property is covered. It is a claim about an evidence boundary: the recorded command result is pinned to the subject it observed. Read the status, then judge the working path on that narrower promise.

    Questions people actually ask

    What is subject-bound evidence?

    Ranex subject-bound evidence is evidence pinned to the exact commit digest it describes.

    Why does a passing test stop counting after code changes?

    Evidence for an older subject does not support a claim about the changed subject, so Ranex refuses it.

    What tree does Ranex run a command against?

    Ranex materialises the committed subject tree from verified blobs before running the bound command.

    Try it. Break it. Tell me what broke. Read the MIT-licensed repository, then ask your next green build which bytes it actually observed.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • The Concurrency Test That Proved Nothing

    The Concurrency Test That Proved Nothing

    TL;DR: A race test proves nothing unless its operations overlap and the test fails against broken code; here, Effect.all was sequential by default. Before you trust another green concurrency test, start with the Field Notes slice log.

    You wrote the race test. You used the function named all. You watched it pass. Then the production bug survives because your “concurrent” work ran one operation after the other. Could this test have failed before the fix?

    In this note

    SLICE-013 found exactly that trap while repairing a stranded-tool defect in the Ranex harness fork. The test used Effect.all. In the Effect beta used by the harness, that API defaults to sequential execution. Without an explicit concurrency: 2, the test would have passed against the broken code and proven nothing.

    The function name was not the concurrency contract.

    This warning applies beyond Effect: parallel-looking code is not proof of parallel execution. A race test that cannot lose is not a race test.

    The bug lived before the runner decided there was work

    The failure was an empty-inbox crash path. The runner returned before reconciliation, leaving a tool projected as running forever.

    In the original path, the runner checked whether there was eligible input. No forced run, no steering input, no queued input: return. The reconciliation that marks interrupted tools sat after that guard. A crash with an empty inbox therefore skipped reconciliation entirely. Nothing else came along to look, and the tool stayed projected running.

    The immediate fix was a hoist. Reconciliation now runs before the eligible-input guard returns, so an empty-inbox call to run() can recover the stranded tool without scheduling a provider turn. The slice also guarded the normal short-circuit behavior for other input states. Fixing one path should not quietly make every empty inbox start doing unrelated work.

    But that only repairs sessions that receive a later run() call. A session nobody calls again remains stranded. The prototype had a reconcile capability, but no caller at startup. Capability without wiring is a polite way to leave the defect in place.

    So the completed slice added a startup sweep in the application graph. At process start, it reconciles stranded tools across sessions without scheduling a provider turn. That is the half that repairs the headline case after a crash: nobody calls run(), and recovery happens anyway.

    A repair path is not shipped until something calls it under the condition you claim it handles.

    Why the first race test was decoration

    The race was between reconcile() and run() for one session. Both could read a tool as running and both could publish an interruption, so the completed slice serialized that work per session.

    The test matters because duplicate recovery is worse than untidy. The slice requires exactly one durable Tool.Failed event; projected end state alone cannot catch the duplicate because the projector can show the same final status after two published events. The broken case reproduced two events. The fix used a per-session semaphore and the test then observed one.

    Calling Effect.all did not make the reproduction concurrent. In this Effect beta, its default is sequential. The source behavior is concurrency ?? 1. Unless the test declares concurrency: 2, one operation finishes before the next begins. There is no overlap. There is no race. There is only a test wearing a race-test hat.

    Concurrency is a behavior you must create and observe, not an intention expressed in a function name.

    The slice record says the unqualified test would have passed against the unfixed code. That is the standard worth borrowing. Do not only make your repaired code green. Put the test on a broken revision or temporarily remove the synchronization and require it to go red. If it stays green, you have a test-shaped story, not evidence.

    Shapes to hunt for before you trust the pass

    You can turn this into a short inspection pass. The point is to prove the system reached the dangerous overlap, not to add more concurrency theater.

    1. Read the framework documentation or installed source for the default execution mode. “All” and “join” do not tell you whether work overlaps.
    2. State the exact two operations that can act on the same state. In this slice it was reconcile() and run(), not two calls to run().
    3. Force both operations to pause after observing the vulnerable state and before publishing their result. Then release them together.
    4. Run the test against the unfixed code or remove the lock. Require the duplicate, stale write, or forbidden result to appear.
    5. Count durable events, writes, and external effects. Do not only inspect the final projection or UI state.
    6. Confirm the fix serializes only the reachable surface. A broad lock can hide a test defect while changing behavior you did not intend to change.
    7. Check the startup path. If recovery is supposed to happen after nobody calls the normal operation again, prove the application graph actually invokes it.

    The checklist gives you a way to reject a passing test that never created the condition it was named for.

    The fix also created a named hazard

    The startup sweep recovers abandoned sessions, but it is database-global. A second process booting can mark tools that a first live process is running as interrupted.

    That hazard was introduced by this slice, not discovered later and polished out of the story. The record says it is accepted scope because the harness normally runs one daemon. It also states the boundary plainly: before the harness runs as more than one process against one database, the fencing slice must gate the sweep on ownership.

    This is where “we fixed it” stops being useful engineering language. The hoist fixes the empty-inbox run() path. The startup sweep fixes sessions nobody calls run() on. The sweep also creates a cross-process ownership hazard. All three can be true at once.

    If you remove that last sentence from a release note because it makes the fix look less clean, you are not making the system safer. You are making the next operator discover the precondition by accident.

    A named hazard with an owner and a prerequisite is not a victory lap — it is a boundary your system can defend later.

    That posture is part of the accountability apparatus behind Ranex: a checkable claim, a record of what it does not cover, and a deterministic gate where authority matters. If that framing is useful, read the accountability apparatus. It is not a claim that Ranex is ready to run for you. Ranex is pre-release, and much of its broader picture remains designed, not built.

    What the record actually proves

    SLICE-013 closed on 2026-08-08 with all seven criteria met, landing as commit a8bc7bdf35 in anthonykewl20/ranex-harness. It is harness-fork work, as the README’s Current work section describes.

    The record proves the unsafe empty-inbox baseline was reproduced, the hoist repaired it, the startup sweep was wired into the application graph, repeated reconciliation produced one durable failure event, and the reachable reconcile-versus-run race was reproduced before serialization closed it. It also records green regression gates, but a green report is not the main point here. The meaningful proof is that the test could fail on the broken condition.

    You can read the slice records in docs/slices/done/ in the Ranex repository. The value is not that a repository has a concurrency test — it is that the test’s own execution semantics were checked before anyone trusted its pass.

    Questions people actually ask

    These answers explain how race tests can pass without creating the overlap they claim to cover.

    Why did the concurrency test pass against broken code?

    The Effect.all call defaults to sequential execution in the Effect beta used by the harness. Without explicit concurrency: 2, the two operations never raced.

    What did the reconciler fix?

    The reconciler moved interrupted-tool reconciliation before the eligible-input guard and added a startup sweep for sessions nobody calls run() on.

    Why is a startup sweep risky in multiple processes?

    The sweep is database-global, so another process starting can mark tools that a first live process is running as interrupted. Session-ID fencing is the recorded prerequisite before multi-process use.

    How do you know a race test proves something?

    A trustworthy race test runs against the unfixed code and requires failure. The test then starts genuinely concurrent operations using the framework’s documented execution semantics.

    Your next test is the one to distrust

    Pick the race test in your suite that makes you feel safest. Read the scheduler default. Put the test against the broken code. Count the event that must not happen twice. Then inspect the startup path for the recovery you assume exists.

    Leave the new hazard in the record with its prerequisite. Your future self does not need a cleaner story. Your future self needs the boundary.

    Try it. Break it. Tell me what broke. If you find a test that passes because nothing overlapped, give the repository a GitHub star and send an honest critique. A critique that finds the false pass is the useful contribution.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Inside the Append-Only, Hash-Chained Journal

    Inside the Append-Only, Hash-Chained Journal

    TL;DR: Ranex keeps verdict records in an append-only, hash-chained SQLite journal so ordinary rewrites are blocked and edited rows are detectable, while rollback and truncation remain an explicit open limit. Read how the kernel works.

    Someone asks why a change was allowed out the door. You open the audit record and find a clean story. What stops that story from being cleaned up after the fact?

    A record that can be quietly rewritten is not evidence of a decision. Ranex records verdict activity in an append-only hash-chained journal. The design does not turn a local database into an untouchable historical archive. It does give you a concrete way to detect a row edit and a concrete command to check the chain.

    In this note

    The record must resist a quiet rewrite

    An append-only journal preserves the sequence of admitted events rather than a polished final summary. That matters when your question is not “what does the database say now?” but “what evidence and verdict existed when this decision happened?”

    Ranex’s journal is a SQLite record. Ordinary updates and deletes are prohibited by SQLite triggers. Each journal row also carries a link into a hash chain, so a row has a relationship to the history before it. The operator command ranex journal verify recomputes that chain. Those facts provide a useful control: the normal application path cannot edit or delete a row, and an edit made around that normal path is not supposed to blend into the sequence unnoticed.

    Do not reduce this to “there is a database.” A database can retain the latest story while losing the important question. When did the story change, and what did it replace? Append-only changes the shape of what the journal is for. The journal is a record to replay alongside the deterministic verdict path, not a mutable dashboard state that happens to contain audit-looking fields.

    This is particularly useful when a worker produces a result you did not watch arrive. Ranex does not take the worker’s summary as proof; it reads the diff on disk and runs checks. The evidence, verdict, and journal give the operator something stronger than a completion message. They give the operator materials to inspect after the excitement has passed.

    Two controls cover different paths

    SQLite triggers and a hash chain are not duplicates. The triggers block ordinary rewrite attempts; the chain makes an out-of-band row edit visible when verification recomputes the links.

    Imagine a program, a stray maintenance script, or an agent with access to the database trying to issue an ordinary update or delete. The SQLite triggers prohibit that normal operation. That is a direct guard where application behavior meets the journal. It helps keep “fix the record” from becoming a casual recovery move.

    Now take a different route. An attacker edits the database file outside the normal path. A trigger cannot fire for a change that bypasses the SQL operation it protects. The hash chain is for that case. A changed row breaks the relationship with following rows, and ranex journal verify recomputes the chain for the operator. Verification is the important verb here. A chain nobody checks is only a claim about how rows were written.

    There is no magic in the word hash. The value of the design is that it gives you a defined test against a defined class of edits. You can run the verifier instead of accepting an application’s self-description. You can ask whether the record still links as it should. If it does not, the journal has evidence of a problem rather than a quietly amended history.

    The repository also says concurrent appenders are serialised before reading the previous link. That keeps append operations from racing into incompatible predecessor links. It is a consistency property for the chain. It is not a claim that the journal has solved every preservation problem, which brings us to the line that needs to stay visible.

    Check the journal you depend on

    You can apply the lesson before you use Ranex. Find the record your release process calls an audit trail, then ask whether it can tell an edit from an original event.

    • Does the normal application path prohibit updates and deletes of decision records?
    • Can an operator independently verify the relationship between successive records?
    • Does the record name the evidence, subject, and verdict it is preserving?
    • Would an edit outside the application leave a detectable break?
    • Who can copy, replace, or restore the record store?
    • Can a reviewer distinguish a current status from the sequence that produced it?

    These questions are the reason the journal exists. Generated work can leave you with more activity than a person can remember. A future reviewer needs to connect a requirement, evidence, the exact artifact evaluated, and the verdict that allowed a next step. The journal is one part of retaining that connection. It does not decide whether the requirement was good; the approved flow graph is intended as the root of trust in the broader designed loop. It retains what the kernel saw and decided.

    There is a natural boundary with signed evidence. Signatures establish something about who held a registered private key. The journal records the admitted sequence. Neither control replaces the other, and neither should be stretched into a claim it does not make.

    The limit is rollback and truncation

    The journal does not detect rollback or truncation. An internally consistent earlier prefix still verifies after later rows are removed.

    That is the boundary named in the README’s known gaps. If later rows disappear and the remaining database ends at an older valid point, recomputing the retained chain succeeds. The verifier detects an edited row in the retained history; it does not know, by itself, that history was made shorter. Serialising appenders does not close this gap. Trigger protection does not close it either.

    Keep the distinction sharp. “Tamper-evident for out-of-band row edits” is a useful claim. “Permanent history that detects every deletion” is not a claim Ranex makes. The repository’s Status section lists the append-only hash-chained journal as working today and names rollback or truncation as a known gap. The project remains pre-release.

    That clarity is better for your incident response. If verification fails, you have a reason to investigate an altered chain. If it passes, you know the retained rows are internally consistent, not that no later history was removed. You can decide what additional retention or external anchoring your own risk requires without pretending the local journal already supplied it.

    Questions people actually ask

    How does a hash-chained journal detect tampering?

    The Ranex journal links each row hash to the previous row, and journal verify recomputes the chain to expose an out-of-band row edit.

    Can a hash-chained journal detect deleted entries?

    The Ranex journal cannot detect rollback or truncation because an internally consistent earlier prefix still verifies after later rows are removed.

    What stops ordinary journal updates and deletes?

    Ranex uses SQLite triggers to prohibit ordinary updates and deletes, while the hash chain detects edits outside that normal path.

    Try it. Break it. Tell me what broke. Read the Ranex repository, then inspect whether your own audit record can show a rewrite.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • The Watchdog That Would Have Passed Every Test and Hung Anyway

    The Watchdog That Would Have Passed Every Test and Hung Anyway

    TL;DR: A watchdog fixes a stall only when timeout errors are terminal and non-retryable; idle and absolute budgets cover different waits. That is the shape to hunt for after reading the Field Notes slice log.

    You have a timeout. The test turns red when the provider goes silent. The build turns green when you add the watchdog. Then a real stalled turn sits there, your session stays busy, and somebody has to intervene by hand. A retryable timeout would have let the watchdog fire, let the criteria pass, and restarted the exact stall that caused the timeout.

    In this note

    That was the defect behind SLICE-012 in the Ranex harness fork. A stalled provider stream had no timeout, so its coordinator never settled and the active session stayed busy. The shipped change makes that stream reach a terminal state on its own. Good. The more useful part is the hole the tests could have missed.

    The session would still hang anyway.

    The timeout that fixes nothing

    A timeout is only a fix when its failure path reaches a terminal state. If your retry policy recreates the condition that timed out, retrying it is a loop, not recovery.

    Picture the order of events. A provider stream stops producing. The watchdog fires. The runner sees an error it is allowed to retry. It starts the stream again. The provider stalls again. The watchdog fires again. Your test can prove the timer fired. It can prove the retry happened. It can still miss the fact that the turn never ends.

    SLICE-012 made the classification explicit: both watchdog failures are typed, non-retryable errors. That detail is load-bearing. The earlier prototype had avoided the retryability problem by accident. A pre-implementation review found that no requirement had said the timeout must be non-retryable, so the completed slice made the decision and required a test that proves it is not retried.

    That is a small wording change with a large operational consequence. “Times out” is not enough. “Fails terminally, without retry” says what the system is allowed to do next.

    If retry recreates the failure condition, retry is a hang with better logging.

    If retry would only restart the stall, the failure class needs to say so. Your operator needs a terminal state, not a more energetic version of stuck.

    Why one timeout cannot do two jobs

    You need two thresholds when the waiting patterns differ by orders of magnitude. An idle deadline catches silence between chunks; an absolute budget limits the whole turn.

    The slice wraps the stream with two separate controls. The idle deadline resets every time a chunk arrives. It answers: “Did the provider go quiet in the middle of a response?” The absolute budget races against the whole consumer. It answers: “Has this turn taken too long, even though activity continues?” Each one has to be proven without the other.

    That independence matters. Feed chunks often enough and idle will never fire; only the absolute budget can stop an overlong turn. Disable idle and stall the stream; absolute must still cut it. One silent-stream fixture cannot prove both. It only proves whichever control fires first.

    The first pull is deliberately untimed by idle. The records explain why: inter-chunk gaps run around 10–100ms, while a reasoning model can take minutes to its first token. One number cannot serve both distributions. An idle threshold tight enough for the gaps would cut the first response on every call.

    There is a cost to that choice, and it is written down instead of tucked behind a pleasant name. A provider that accepts the connection and never sends a chunk is bounded only by the absolute budget. At the default, that can be up to 30 minutes. The record does not claim that setting is ideal. It names the limitation: a separate first-chunk budget, identified as the proper fix and left out of this slice.

    Do not let one comforting timeout setting pretend it covers three different waits. First response, inter-chunk silence, and whole-turn duration need their own evidence. If you cannot show that evidence, say what remains unbounded.

    What to inspect in your own pipeline

    You can find this class of failure without adopting Ranex. Start where a streaming call, a retry policy, and a session state meet. The job is to make the broken path fail before you trust the repaired one.

    • Find every timeout and write down its failure class: retryable, non-retryable, interrupt, or something else.
    • For each retry, ask whether the next attempt changes the condition that failed. If it does not, prove the system settles instead of cycling.
    • Separate time to first token, gaps between chunks, and total work time. Do not reuse a threshold merely because it is nearby.
    • Make a stream send one chunk and then stall. Confirm the session reaches a terminal state without a person stopping it.
    • Make a healthy slow stream complete below the idle threshold and below the whole-turn budget. A watchdog that cuts legitimate work is not a watchdog you can trust.
    • Run the timeout while tool work is in flight. Check the actual tool state and settlement, not a log line that says cleanup happened.
    • Use a non-default timeout and prove behavior changes. Reading a configuration value back is not proof that it controls the running system.

    There is no glamorous trick here. You are looking for the exit from failure, not the detection of failure. That distinction saves you from a green test whose only achievement is proving a timer owns a clock.

    The second hang hiding behind the first

    A watchdog failure is an error, not an interrupt. That difference left tool fibers dispatched during streaming uncleared, and the later settlement wait held the turn open forever.

    This is the defect that showed up while building the fix. Cleanup had been tied to interrupts. The watchdog produced a typed error instead, so the cleanup did not run. The runner then waited for tool fibers that had not been cleared. The visible timeout existed; the terminal state did not.

    The completed slice required a specific observable for this path: the tool fiber terminates and the tool is recorded as interrupted. That is stronger than a feeling that a tool was not stranded. It makes the question checkable after the fact.

    That is also why a green test suite deserves a hard question. What did it actually exercise? In the slice record, the unsafe baseline had to hang a real provider stream inside the runner. Testing a timeout helper in isolation would have been decoration. The runner, the failure class, the retry behavior, the tool cleanup, and the terminal session state all had to meet in the same proof.

    SLICE-012 closed on 2026-08-07 with all nine criteria met, landing as commit 23d6a5b4ee in anthonykewl20/ranex-harness. This is harness work, not a claim that Ranex is ready for use. Ranex is pre-release; much of the broader system is designed, not built. The watchdog is one shipped durability claim in the harness fork.

    If you want the surrounding model for why the verdict must sit outside an agent’s own loop, read how the kernel works. The slice record itself is in the Ranex repository under docs/slices/done/.

    Questions people actually ask

    These answers explain timeout failures, terminal states, and the waits each budget covers.

    Why can a watchdog timeout still leave a session hanging?

    A retryable timeout restarts the same stall when the condition that timed out will recur after retry. The watchdog can fire while the turn still never reaches a terminal state.

    Why use both idle and absolute timeouts?

    An idle deadline detects a quiet mid-stream provider, while an absolute budget bounds the whole turn even when chunks keep arriving.

    Does an idle timeout cover time to first token?

    No. The first pull is deliberately untimed because inter-chunk gaps and reasoning-model time to first token have very different latency distributions. A connection that never sends is bounded by the absolute budget.

    What happened to tool work when the watchdog failed?

    The watchdog produced an error rather than an interrupt, so streaming tool fibers were not cleared and settlement could keep the turn open. The slice added coverage for that path.

    Your next failure drill

    Take one timeout in your pipeline this week. Do not start by making the timer shorter. Stall the real operation after it begins. Then follow the failure all the way through: retry decision, cleanup, settlement, and the state your operator sees.

    Write the failure class down. Split thresholds that are serving different distributions. Name what remains unbounded.

    Try it. Break it. Tell me what broke. If this record helps you find a test that passes while its failure loops forever, give the repository a GitHub star and send an honest critique. The critique is more useful than applause.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Ranex vs. GitHub Rulesets and Copilot Hooks: What’s Actually Different

    Ranex vs. GitHub Rulesets and Copilot Hooks: What’s Actually Different

    TL;DR: Rulesets, editor hooks, and CI gates can control where work moves. Ranex asks what the evidence for that work must establish. Those boundaries can sit together. The accountability apparatus needs both a route and a record.

    You already have rules around a push, a merge, or an editor session. Good. The question is not whether you should throw them away. The question is whether a green result carries the information you need about the artifact that reached that channel.

    In this note

    Channel controls decide where work can move

    GitHub rulesets sit around repository actions such as push and merge. Editor hooks sit in an editor workflow. CI pipeline gates sit in the pipeline. Each can enforce the policy you configure at its own boundary.

    That is worth having. A branch should not advance when required conditions for that branch are absent. An editor action can be stopped before it runs. A pipeline can stop before it publishes its next result. Those controls make authority visible at the channel where the action happens.

    Keep the comparison boring because it is boring. These are different places to put a boundary. Your existing controls may be the right answer for your repository, your team, and the action you need to constrain.

    But a channel rule does not answer every evidence question by itself. A pipeline can report a result. A merge policy can require that result. An editor hook can permit or refuse an action. You still need to ask what command ran, what claim it was meant to support, and what exact artifact it observed.

    That is not a complaint about those controls. It is the normal limit of a boundary. A lock on a door decides who can enter; it does not describe what happened in the room. Your workflow needs the channel decision and a record that can answer the evidence question later.

    Evidence controls decide what a claim must show

    Ranex puts its boundary around the claim. Its kernel evaluates gate, evidence, subject, and approver as a pure function. The evidence is bound to a subject digest. A required claim without satisfying evidence is FAIL rather than a default or a skip.

    This is a narrower question than “can this merge happen?” It is: can this record support this claim about this code? The README’s “Status” section says the same command run against a different commit proves nothing about this one. That is the evidence boundary in plain language.

    Suppose a test job produces green output. The channel control can decide whether that output is required before merge. The evidence question remains open until you can state what the job ran, what it measured, and whether its record is bound to the code under judgment. A result with no stated proposition is hard to audit later.

    The evidence boundary also separates production from approval. Ranex lists no self-approval as working behavior: whoever produced the evidence cannot approve it. That does not make a channel control less useful. It supplies a different question for a workflow where an agent can produce both the change and the report about the change.

    Use the boundaries together

    You can adopt the evidence question without adopting Ranex. Start with the controls you already use, then make each required result answerable.

    • Which channel is being controlled: push, merge, editor action, or pipeline stage?
    • What exact claim does the required check support?
    • What command is authorized to support that claim?
    • Which code digest did the command observe?
    • Who produced the evidence, and who is allowed to approve it?
    • What happens when the evidence is missing or belongs to older code?

    That checklist does not ask you to abandon a familiar stack. It asks you to make the green light legible. Your ruleset can still protect the merge channel. Your hook can still protect the editor channel. Your pipeline can still coordinate work. The evidence record needs to carry its own meaning.

    There is no trophy for replacing working controls. Composition is the point. Let a channel control stop an unauthorized transition. Let an evidence control refuse a claim that lacks the right proof. The first answers where an action may go; the second answers what the action established.

    The status is not a replacement claim

    Ranex is pre-release. The README says it is a kernel with a working verdict path and very little else. Subject-bound evidence, absence blocks, no self-approval, the journal, and the run-to-evaluate path are listed as working today. The full flow graph and scenario compilation are designed, not built.

    Its limits are part of the comparison. Approver identity is unauthenticated: --approver is a plain string. Same-UID key theft remains open. The journal does not detect rollback or truncation. Those gaps mean you should read the status before treating any design goal as present capability.

    So do not read this as a replacement pitch. Use the channel controls you trust. Ask the evidence question wherever agent-generated work raises the stakes. If Ranex earns a place later, it earns it by making that question executable, not by declaring your current setup inadequate.

    Questions people actually ask

    How does Ranex differ from GitHub rulesets?

    GitHub rulesets govern repository channels such as push and merge, while Ranex evaluates evidence for a claim against a subject digest.

    Can Ranex work with editor hooks and CI gates?

    Ranex can sit beside editor hooks and CI gates because their channel controls and its evidence question address different boundaries.

    Is Ranex ready to replace an existing governance stack?

    Ranex is pre-release, so its README describes a working verdict path and known gaps rather than a replacement claim.

    Try it. Break it. Tell me what broke. Read the MIT-licensed repository, then write down what one required check in your stack actually proves.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • The First Real Delegation: An Agent, a Worktree, and a Judge With No Key

    The First Real Delegation: An Agent, a Worktree, and a Judge With No Key

    TL;DR: Run agent-written code without a signing key, measure it in a separate keyless process, and keep approval with another authority. This is the next entry after the Ranex slice log.

    You delegated work to an agent, then watched it run code it had just written. Would you let that same process hold the key that signs the result? A real delegation is built around the answer “no.”

    In this note

    The dangerous moment after the agent edits the tree

    The risk is not only that an agent writes the wrong code. The risk is that code it wrote gets to reach the authority that declares the work trustworthy.

    A headless agent in a worktree feels contained until you follow the process tree. It can edit the repository. It can trigger the suite. The suite is code it may have changed minutes earlier. If the execute process carries a signing key, the boundary exists only in the story you tell about it.

    SLICE-008 began with a blunt defect: there was no front door. A person had to connect dispatch, harness, emission file, and judgement by hand. The earlier gear-mesh proof used a model that did no work and hand-wrote its own evidence. That proved parts could turn. It did not prove that the kernel could judge work an agent actually produced.

    There was a sharper failure underneath. The existing ranex run path read the signing key before the bound command spawned. The slice records that RISK-06 reproduced that command taking the key from the process environment. If an agent-controlled test hook can reach that key, it can produce a signature while the rest of the system looks tidy.

    Do not file that concern away. Design it out of the execution phase.

    The lesson: separate execution, measurement, and judgement

    The practical split is direct: run the agent without a signing key, measure its committed work in a separate keyless invocation, then attest later. The judge produces a candidate with named missing claims, never a self-issued PASS.

    ranex task delegate dispatches a worktree and starts the harness headless. Its environment is built from empty: a pinned PATH, a scratch HOME outside the worktree, bridge variables, one model-credential variable, an auto-approve flag, and a wall-clock bound. When the signing-key variable is present, delegation refuses before it spawns anything.

    “Built from empty” means one concrete thing: the execution environment is deliberately assembled instead of inheriting whatever trust happened to be in the parent shell. That is a much better place to start asking what the worker can reach.

    When the agent finishes, the kernel parses its emission and matches the worktree and commit against its own dispatch record. A missing emission blocks. A mismatch blocks. An empty delegation where the commit equals the base is refused rather than judged.

    Next, the kernel runs the frozen suite against the emitted commit. The commit is materialised, sealed, and offline for that measurement. The invocation holds no key.

    Only after that does a separate, later invocation sign the result. Untrusted code is not running at that time.

    The worker can make a change. The measurement can describe the change. Neither action grants approval.

    The task judge result is a journalled CANDIDATE naming its missing claims. The stamp stays a human’s and out of band. That distinction is the whole design, not a label applied at the end.

    For the wider reason a separate check port matters, read the accountability apparatus. An accountability system needs an inspector outside the actor it inspects.

    A checklist for your own delegated loop

    You can audit this without adopting a kernel. Start with the process that executes agent-produced code, then trace every secret, timeout, artifact, and verdict out to its owner.

    • Start from an empty environment. Write down the exact variables and paths the worker receives. Treat inherited shell state as unreviewed input.
    • Refuse execution when a signing key is present. Do this before spawning the agent or its process tree, not after it has begun.
    • Inspect the live process tree. Do not accept a source-level assertion that a key was removed. Check what the child and parent can actually reach.
    • Bind a task to its worktree before work begins. Make the judge compare the emitted worktree and commit with a dispatch record it owns.
    • Measure a frozen suite against the emitted commit. Do not let a worker summary stand in for an independently observed result.
    • Bound the whole process group. A wall-clock limit must kill the group on a stall, not leave child processes behind.
    • Use a state short of approval. Candidate is an honest output when claims remain missing. A worker should not award itself PASS.

    One dad-joke-sized warning: a timeout that kills only the parent is not a timeout. It is a process group hug with the interesting parts still running.

    What this delegation proved — and what it did not

    SLICE-008 proved an end-to-end delegated run against a real free model. It ended in a journalled CANDIDATE naming missing claims, with no PASS, and a reviewable diff.

    It also tested the shapes that tend to become decoration. The execute phase refuses a signing key before spawn. The live process test checks that the delegated command cannot reach the key through environment, file path, or parent. A planted conftest.py may run and fail the suite, but it produces no signed record. Forged and missing emissions block before materialisation.

    The wall-clock bound terminates a stalled run’s whole process group, records the timeout, and journals no candidate. The fork’s operator-facing command presents as ranex, while opencode’s MIT attribution remains in the tree.

    Those are useful facts. They are not a claim that all delegation is safe.

    RISK-06 stayed open when this slice closed. The model credential sat in a network-open loop, where it could be posted elsewhere. The record said to use a scoped, spend-limited key, and that the existing ranex run path still read the key before spawning. Recorded is not mitigated — until it is: SLICE-046 later closed RISK-06 by binding ranex run’s command inside the qualified confinement session, so the worker can no longer take the signing key. The controller that runs that session is still same-uid trusted infrastructure — see the gap list for what that still leaves open.

    The slice also does not close merge collisions, verifiable separation, or gate quality. A weak gate can still accept plausible code. You own the target and the quality of the gate. The kernel cannot make either judgement disappear.

    The source record lives in docs/slices/done/SLICE-008-first-delegation.md in the Ranex repository. Read it if you are designing this boundary. The failures are part of the useful material.

    Questions people actually ask

    These answers cover keyless measurement and the remaining model-credential risk.

    Can a delegated agent expose its model API key?

    The model credential remains in a network-open loop — use a scoped, spend-limited key. RISK-06, the risk that the worker could take the signing key, was open when SLICE-008 closed; SLICE-046 closed it later by binding ranex run inside the confinement session.

    What does an environment built from empty protect against?

    The delegated harness receives a deliberate environment rather than inherited trust, and delegation refuses to start when the signing-key variable is present.

    Did SLICE-008 solve model credential exposure?

    No. The credential still sits in a network-open loop, so use a scoped, spend-limited key. SLICE-008 left RISK-06 open for ranex run; SLICE-046 closed it later by binding the command inside the confinement session.

    Your next move

    Take one agent job that runs tests or scripts. Identify the secret that can authorise its result. Then inspect the live child and parent processes while the job runs. If that secret is reachable, split execution from measurement before you trust another green status.

    Write down what remains open too. “Recorded, not mitigated” is not a failure of the notes. It is how the next person avoids mistaking a boundary for a finished fortress.

    Try it. Break it. Tell me what broke. If this was useful, star the Ranex repository and leave an honest critique. A hard question is more valuable than a polite nod.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Absence Blocks: Why No Evidence Is a Fail, Never a Skip

    Absence Blocks: Why No Evidence Is a Fail, Never a Skip

    TL;DR: A required claim without satisfying evidence must fail, because a skipped check is not proof; a fresh Ranex clone starts red for exactly that reason. Read how the kernel works.

    Your dashboard is green, yet a job did not run. Do you know which one? If the answer lives in a log nobody reads, your green result is carrying more confidence than the evidence earned.

    Ranex treats absence as a blocker. When a gate requires a claim and no satisfying evidence exists, the verdict is FAIL — not a default, not a skip, not a warning waiting for a tired person to scroll past.

    In this note

    A silent skip is not a pass

    A check that did not run has established nothing. Treating that empty space as success turns missing work into a green light.

    You know the shapes. A condition bypasses a job. A suite collects no tests and returns cleanly. A scanner stops early and leaves a warning below the fold. The summary says green because the system recorded an exit path, not because it recorded the evidence you needed. That is a failure mode, not an inconvenient edge case.

    Ranex makes the missing piece explicit. A gate carries required claims. Evidence must satisfy those claims for the subject under judgment. If the required evidence is absent, evaluation returns FAIL. The rule is an invariant of the kernel, alongside subject-bound evidence and no self-approval — though approver identity is unauthenticated today, so that check compares unverified strings. A gate that cannot block is refused at construction, because a gate that cannot stop anything is decoration.

    This is why absence and determinism belong together. A pure-function verdict gives the same answer for the same inputs. Absence blocks answers FAIL when an input cannot satisfy the rule. Without both, “same inputs” can still mean “missing evidence became permission.”

    There is a cost. A fail-closed system will make you produce evidence before it lets you call the work done. That can feel fussy when you know the command would have passed. But the entire point is that the system does not accept what you know without a record it can evaluate.

    A fresh clone starts red

    A fresh Ranex clone fails gate evaluation because it has no evidence yet. That failure is correct: no record exists for the required tests-executed claim.

    The README gives the operator path. Run the test suite with frozen dependencies, then ask the gate to evaluate the current subject. On a fresh clone, evaluation returns FAIL with a nonzero exit and names the missing claim. It does not create a cheerful initial state. It does not infer success from an empty evidence file. It tells you that nothing has been proven.

    PYTHONPATH=src uv run --frozen python -m ranex.cli.main gate evaluate HEAD \
        --approver reviewer_alice

    To produce evidence, the operator needs a signing identity whose private key lives outside the repository. The committed keyring contains public halves, not private keys. After deliberate dependency provisioning, ranex run records the observed command result and subject digest; evaluation then judges that evidence against the gate. The steps matter because the verdict is not an applause button for a command. It is a judgment about whether the bound evidence satisfies the required claim for this subject.

    Fresh-clone failure is a fast test of your own mental model. If a system can announce PASS before any relevant evidence exists, what did PASS mean? It meant the system had another rule, even if nobody wrote it down. Ranex writes its rule down in behavior: no satisfying evidence, no pass.

    Make absence visible in your pipeline

    You do not need to adopt a kernel to look for this fault line. Pick a recent green deployment and inspect the evidence path rather than the badge.

    • List every claim the release depends on, not only every job name.
    • For each claim, identify the record that satisfies it and the exact subject it covers.
    • Make a missing result fail rather than disappear into a default branch.
    • Make skipped, errored, and missing test outcomes visible as distinct states.
    • Change the tree and confirm earlier evidence no longer satisfies the new subject.
    • Ask whether an operator can tell the difference between unfinished work and a rejected record.

    That last distinction matters. A record that fails verification is not reported as “no evidence.” Ranex reports it as refused and includes a reason. Missing evidence means no satisfying record exists. Refused evidence means a record arrived but did not clear verification. An unfinished task and an attack are different events; hiding them under the same empty label destroys useful information.

    Staleness is absence in another form. Evidence is bound to a subject digest, so the same command on a different commit proves nothing about the current one. Once the tree moves past the digest that evidence covered, it stops satisfying the claim. That is not the system being difficult. It is the system refusing to let yesterday’s proof stand in for today’s code.

    Failure names the missing proof

    A useful FAIL tells you which required proof is missing or refused. It does not pretend that all failures are identical, and it does not claim a guarantee it lacks.

    The repository’s Status section lists absence blocks as working today and says the project is pre-release. The full flow-graph and scenario-compilation picture is designed rather than built. Keep that boundary in view. The working claim here is smaller: the verdict path fails when a required claim has no satisfying evidence.

    That small claim changes the conversation during a release. Instead of “the pipeline was green,” you can ask “which evidence satisfied this claim for this commit?” If there is no answer, you do not need a debate about whether the omission feels safe. The gate already has the right response.

    Go to the last build you trusted because it was quiet. Find one claim it was supposed to establish. Then remove its evidence in a disposable copy of the process. If the release still looks green, you found the work to do.

    Questions people actually ask

    What does absence blocks mean in CI?

    Ranex treats a required claim without satisfying evidence as FAIL, never as a default or a skip.

    Why does a fresh Ranex clone fail evaluation?

    A fresh Ranex clone has no evidence for the required tests-executed claim, so gate evaluation correctly returns FAIL with a nonzero exit.

    Is refused evidence the same as missing evidence?

    Ranex reports a record that fails verification as refused with a reason, while missing evidence means no satisfying record exists.

    Try it. Break it. Tell me what broke. Read the Ranex repository, then make one required check absent in a safe copy of your own pipeline.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records — the same fact gate the product enforces on code. It ships only after Anthony’s own review.