Blog

  • Measure Before Learning: Stop Calling One Run a Win

    Measure Before Learning: Stop Calling One Run a Win

    TL;DR: You cannot call a change better because you watched it work once; freeze the comparison and its rules before you look. The Ranex slice log records the same distinction between claims and evidence.

    You change a prompt, a handbook, a workflow, or a model setting. One run looks better. The temptation is immediate: ship the winner and tell yourself you learned something.

    You learned that one run happened. You did not learn that the change caused the result.

    In this note

    One run is an observation

    An observation records what happened on one run. A grade answers a stricter question: did a candidate beat a control under a comparison you defined before the result arrived?

    That difference is the center of ADR-016. It covers future evidence for revisions to handbooks, rules, processes, workflows, models, skills, and their harnesses. The ADR rejects treating every run as a grade because monitoring has no counterfactual and lets people change the treatment after they have seen the outcome. That’s painting the bullseye with extra steps — moving the treatment once you already know where the dart landed.

    A counterfactual is simply the missing comparison. If the candidate passed, what would the unchanged version have done on that same case, under the same conditions? Without that answer, your explanation has a hole right where causation should be.

    Ranex’s design prevents the shortcut structurally. An ExperimentObservation belongs in a separate experiment-record namespace and binds the experiment, pair, arm, subject, target, treatment, outcome digests, and fault classification. An ordinary observation cannot become a grade just because someone likes its number.

    Freeze the ruler first

    You should freeze the material and the rules before the candidate exists. If the candidate can shape the ruler, a favorable result tells you far less than it appears to.

    The design freezes a TargetCorpus and resolves a TreatmentManifest before dispatch. The treatment includes the relevant dependencies, model and provider metadata, harness, handbook, rules, process, workflow, prompt, skills, tools, and budgets. An alias without a provider-issued immutable version is ineligible unless the owner approves that concrete version.

    That sounds strict because reproducibility has a strict requirement. You cannot rerun “whatever model that name meant last week” and call it the same treatment. You also cannot let a proposed change alter the corpus that judges it. The candidate is untrusted input, never authorization.

    For the first scope, ADR-016 fixes one candidate and one binary PASS/FAIL primary endpoint. Before the trial, the design preregisters corpus size, alpha, practical-effect threshold, minimum discordant count, and fault policy. It uses a paired A/B trial on the same frozen corpus, with deterministic counterbalanced order and a fresh worktree or session.

    Inconclusive is a valid answer

    An inconclusive result means the comparison did not earn a stronger claim. It is not a soft pass waiting for a launch date.

    The first design uses exact McNemar on discordant pairs and reports an exact Clopper-Pearson interval for discordant-win probability. SUPERIOR requires more candidate wins than losses, the preregistered minimum number of disagreements, a two-sided result at or below alpha, and the registered practical margin. INFERIOR is symmetric. A grader or registered-fault-limit failure is FAULT. Every other complete result is INCONCLUSIVE.

    Those names are not branding. They stop a familiar move: seeing a secondary metric improve, ignoring a failed primary endpoint, and declaring progress anyway. ADR-016 makes secondary metrics descriptive and ineligible. It also states an important limit: exact power and sample-size design for this paired test is UNVERIFIED because the research did not find mature permissive exact McNemar power code.

    The design says what it can grade, what it cannot yet establish, and what happens when the evidence is thin.

    A checklist before you call it better

    You can take this process to your own change, whether or not your work involves AI. The point is to stop the result from choosing the rules that justify it.

    • Name the primary outcome. Choose one result that decides the comparison.
    • Freeze the cases. Keep the candidate from editing the target after it is proposed.
    • Resolve the treatment. Record the exact version, dependencies, tools, and settings you actually ran.
    • Run a paired control. Compare old and new on the same material rather than comparing two unrelated good days.
    • Set failure policy first. Decide which infrastructure faults void a pair and which outcomes count against the candidate.
    • Keep promotion separate. A grade is evidence; an accountable owner still decides whether to promote.
    • Allow inconclusive. If your process cannot return that answer, it is built to confirm a belief.

    Promotion is not self-approval

    A comparative grade does not authorize promotion. ADR-016 keeps promotion separate, requiring authenticated owner approval, stale-base checking, monotonic versioning, and a new version for rollback.

    This protects the boundary between a proposal and a decision. A candidate generator can recommend a revision. It cannot declare itself superior, approve itself, or convert a one-off run into authority. Unknown fault classification blocks and returns INCONCLUSIVE; absent evidence never promotes.

    That separation matters for people too. You can ask an assistant for a draft, a review, or a proposed fix. You still need a process that tells you whether the evidence supports changing the thing people depend on.

    Where Ranex stands

    Ranex measurement is not implemented. ADR-016 is an accepted source-backed design only; it does not claim a measurement feature in the working product.

    Its measurement-local M0 is a disposable two-week prototype that opens only after ADR-017 P0 permits it. M0 must produce a digest-bound green exit record before future F1 and later production work can proceed. A red record followed by dropping or superseding the design is an allowed result.

    Ranex overall is pre-release. The README says the product has a working verdict path and little else; this measurement path remains future work. Treat that status as a constraint on every claim in this note.

    Questions people actually ask

    What is the difference between an observation and a grade?

    ADR-016 defines an observation as a recorded run and a grade as the result of a preregistered controlled paired comparison; ordinary observations structurally cannot grade.

    What must be fixed before a paired trial starts?

    ADR-016 requires the frozen corpus, resolved treatment, primary endpoint, sample size, alpha, practical threshold, minimum discordant count, and fault policy to be fixed before the trial.

    Is Ranex measurement live today?

    Ranex measurement is not live today: ADR-016 is an accepted design, its M0 prototype is future work behind ADR-017 P0, and the product remains pre-release.

    Before you call a change a win, write down the old version, the cases, the threshold, and the answer you will accept if the result stays unclear. Try it. Break it. Tell me what broke.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • How a 12% Flaky Test Suite Got Approved Twice

    How a 12% Flaky Test Suite Got Approved Twice

    TL;DR: A green report is not evidence until its command is rerun against the artifact on disk; here, that exposed two failures in 16 runs. The parent slice log records the result.

    Your test report says stable. Is the suite stable, or did the report get there first? That question can save you from approving a defect you cannot see. A durability suite reported stable, received approval from two independent reviewers, and still failed when someone ran it again against the worktree on disk.

    In this note

    The failure rate was roughly 12%. It was not hidden behind a complicated attack. A fixture threw away the stderr that would have named the problem.

    If you run a CI pipeline, review agent work, or sign off on a release, this is your problem too. A green report is a statement. The files and command result are the thing that statement must answer to.

    The approval that did not survive contact with disk

    The suite was unstable even after it was reported stable and approved twice. The instability appeared only when the supervisor reran every gate against the actual worktree, instead of reading the session that created it.

    A test report approved by two checkmarks; re-run against the disk, three of the status dots blink red
    The report on the left was approved twice. The re-run on the right reads the worktree on disk, where three tiles flicker.

    SLICE-011 was a disposable prototype for five durability claims. It shipped nothing. Its job was to measure unsafe behavior red first, prove a proposed control green, preserve an exit record, and make later production work refuse to proceed without that digest-bound record.

    The fifth claim covered session ownership. First, the prototype had to show two processes could drain one session. Then it had to show the existing fence refused the second owner. That is the right shape: demonstrate the hole before celebrating the lock on the door.

    The reported result looked fine. The ownership check was marked 5/0, then 20/20 stable after a fixture fix. But before that correction, the reruns measured two failures in 16 full-suite runs.

    The fixture set PRAGMA busy_timeout after journal_mode. Two workers opening concurrently could hit database is locked outside the Effect catch. The failure signal existed. The fixture captured worker stderr and discarded it.

    The evidence was the problem, not the reviewers. Two independent reviewers approved what they were shown. The artifact on disk said something else.

    Neither failure was found by reading a report.

    There was a second hit in the same slice. Its compiled gate was built incorrectly for the case it existed to cover, following a mistaken supervisor instruction. A reviewer caught that, and it was reproduced on disk before the fix. The gate had derived the record from the open slice when it needed to resolve the prototype record in both the open and completed slice locations.

    That is the scar. The process that was supposed to keep a durability record honest contained a wrong gate and a flaky control. Neither yielded to polished prose.

    What your green report cannot tell you by itself

    A green report cannot establish that the command was rerun, that the fixture exposed its diagnostics, or that the gate tests its own refusal path. You need the artifact, the command, and an independent rerun.

    A session summary is useful for navigation. It is not evidence. The SLICE-011 record says this plainly: the session’s own summary is discarded self-report.

    This matters most when a test touches timing, concurrency, process teardown, retries, or shared state. Those are exactly the places where one clean run can convince you that a mechanism works while the next run proves the control was never stable enough to speak.

    Here is the deal: your negative control must be stable too.

    SLICE-011 had a missing written control: the negative control itself needed to be stable. A flaky control cannot tell you whether you found a race or changed the mechanism. It turns every result into an argument about noise.

    Do not replace that requirement with more confidence in the review. Review can catch a bad instruction, as it did here. It cannot transform an unrerun artifact into a measurement.

    Go hunt these shapes in your pipeline

    Start with the checks that block a deploy, approve an agent change, or certify a safety control. You are looking for ways the pipeline can describe success without preserving the result that earned it.

    • Find a check whose result is accepted from a worker, agent, or session summary. Rerun its exact command against the checked-out worktree.
    • Find fixtures that capture stdout or stderr. Make a deliberate failure and confirm the diagnostic reaches the person judging it.
    • Find concurrency or timing tests with a single successful run. Run the full suite repeatedly and record failures as failures, not as an inconvenience.
    • Find every negative control. Ask whether it is stable enough to distinguish the old mechanism from a race.
    • Find a gate that protects a record or workflow. Break the path it was created to refuse and verify that it refuses for the stated reason.
    • Find instructions that changed a control’s lookup or input. Treat the instruction as a hypothesis, then test the behavior on disk.

    Do this before you add another dashboard. You do not need a new reporting layer to learn whether a fixture swallowed the one line that mattered.

    What the slice actually proves, and what it does not

    The prototype proves its five durability claims red to green in a scratch harness worktree, with negative controls and a digest-bound record. It does not ship production durability.

    That boundary matters. The record exists to gate later durability production slices. It is not a license to claim that durable retry, durable blockers, or session-ID fencing are built in production. The README calls those three remaining production claims unbuilt and says Ranex is pre-release.

    The corrected ownership control reached 20/20 stable. That is a result for this prototype after the fixture fix, not proof that all concurrency checks in your repository are sound. I do not have evidence for that.

    The useful proof is narrower and stronger: when the supervisor reran the artifact on disk, the suite’s self-description lost to measurement. When a reviewer challenged the compiled gate, the gate was reproduced and corrected on disk. The process improved because it made claims answerable to an artifact outside the session that made them.

    That is also the point of the kernel: the verdict comes from executable checks and evidence, not from a worker’s confidence. Ranex is still pre-release, with a working verdict path and substantial designed work remaining. The slice records live in the repository under docs/slices/done/, including the parts that failed before they were fixed.

    Questions people actually ask

    These questions help you inspect flaky suites before you accept a green check.

    Why is a test report not enough evidence?

    A report can describe a run inaccurately or hide the artifact that produced it. Re-run the gate against the worktree on disk.

    How was the flaky test suite detected?

    A supervisor reran every gate against the worktree on disk and measured two failures in 16 full-suite runs.

    What caused the SLICE-011 flakiness?

    The fixture set PRAGMA busy_timeout after journal_mode, and it discarded worker stderr that contained the database-lock message.

    Your next approval needs an artifact

    Before you approve the next flaky-looking check, take the command out of the report and run it against disk. Force the failure path. Read stderr. Run the full suite again. Then ask whether the control that proves the old behavior is stable enough to mean anything.

    Try it. Break it. Tell me what broke. If this helped you catch a green report that could not survive a rerun, star Ranex on GitHub and send the honest critique. The useful critique is the one that finds the next hole.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Hashing Dependencies Does Not Make Them Honest

    Hashing Dependencies Does Not Make Them Honest

    TL;DR: A dependency hash proves which bytes arrived. It does not prove that those bytes will tell the truth after your test runner imports them. The parent slice log records the boundary.

    Your lock is pinned. Your wheels match their hashes. Can one of those approved wheels still make the test command say PASS? Yes. That is not a hash failure. It is the limit a hash never claimed to cross.

    In this note

    A hash proves identity, not behavior

    A SHA-256 digest answers one exact question: are these the bytes you expected? It cannot answer what those bytes do when executed.

    That distinction matters when dependencies participate in a measured command. Ranex materialises the committed subject tree. Ignored environments such as .venv are not part of that tree, so its own bound command, uv run pytest -q, needed a deliberate way to obtain dependencies without trusting an ambient writable environment.

    The tempting answer is “hash the wheels.” Hashing is necessary here. It is not independence. A wheel contains executable importable code. In this Python path, pytest loads installed pytest11 entry points before collection. The dependency can therefore influence the exit code the gate later evaluates.

    ADR-007 states the boundary directly: dependency code, the resolver, the interpreter, and the environment builder are trusted computing base for that run. Confinement can preserve the selected bytes and deny network access. It cannot turn dependency-selected behavior into an independent fact.

    This is not an argument for dropping hashes. It is an argument for refusing to call identity proof a behavior audit.

    Provisioning still earns its place

    Dependency provisioning reduces hidden change. It makes the run reproducible enough to inspect, rather than accepting whatever happens to be installed.

    Ranex separates preparation from measurement. deps fetch is the only networked phase. It resolves a clean copy of the committed manifest with no lock present, under an operator-pinned resolver, Python target, index set, and resolution epoch. The generated lock must byte-match the committed lock. Every selected wheel needs a SHA-256 address; source distributions, local paths, VCS sources, and missing hashes refuse.

    The rest follows the same discipline. Store entries are re-hashed when read. The dependency root is assembled from verified entries and made read-only. Approval presents package additions, removals, and version changes, not only opaque digests. The measured command runs offline, with the catalog-bound argv unchanged.

    Those controls answer real substitution and drift problems. An authored lock is not accepted as proof of the manifest. A corrupted stored wheel does not quietly load. A new package cannot arrive after approval. If you need to run a dependency-bearing suite against an exact subject, that is valuable work.

    For the full operator account that exposed nine defects around this path, read Nine Defects, Zero Unit Tests. This note stays on the narrower boundary: even a clean, approved, hash-correct dependency remains code that can affect the result.

    The approved wheel that can lie

    An approved, hash-correct wheel can force a passing verdict. Ranex has an executable limit for exactly that case.

    tests/security/test_slice006_approved_wheel_can_lie.py builds an approved wheel with a pytest11 plugin. The wheel passes every integrity control and then forces success. The test is expected to stay green because it proves a declared gap, not a defence.

    This is where security language gets expensive. “Verified dependency” can mean the hash matched. It must not be read as “safe dependency,” “trusted publisher,” or “honest test result.” The slice labels import-time execution not caught and preserves a control showing the suite genuinely ran. A green test for an open boundary is not a solved boundary.

    Direct imports keep the limit larger than one pytest mechanism. Disabling plugin auto-load would not make other imported dependency code harmless. The dependency is inside the run that produces the evidence.

    Questions for your dependency gate

    You can use this distinction in any ecosystem. Start by separating bytes, approval, and runtime behavior instead of making one word carry all three.

    • Does your resolver derive the lock from controlled inputs, or does it trust a lock the measured party could author?
    • Do missing hashes and unsupported sources refuse, or are they treated as convenient exceptions?
    • Are resolver, interpreter, indexes, and the artifact store outside the party being measured?
    • Can the measured run download, sync, or rewrite its dependency root after approval?
    • Does a person see the package delta they approve?
    • Which dependency code runs before your test command has collected a test?
    • Does your documentation say that a matching hash proves identity, not behavior?

    You will not remove the trusted computing base by naming it. You will stop hiding it behind a stronger claim than the evidence supports.

    Keep the claim small and useful

    Run the strongest dependency process you can justify: clean derivation, pinned inputs, content addressing, readable approval, and an offline measured run. Then state its result precisely.

    The result is not “the dependencies are honest.” It is “these approved bytes entered this controlled run.” That is a useful fact. It is also the fact you earned.

    Ranex is pre-release. Its dependency provisioning path is built and tested, but it is not a promise that third-party code cannot influence a verdict.

    Questions people actually ask

    These answers separate artifact identity from the behavior code can take after import.

    Does hashing a dependency prove it is safe to run?

    No. Ranex treats a SHA-256 hash as proof that approved bytes arrived, not proof that those bytes behave honestly after import.

    What does Ranex dependency provisioning check?

    Ranex cleanly derives the committed lock under pinned inputs, requires SHA-256-addressed wheels, records an approved package delta, and runs the measured command offline.

    Can an approved wheel force a passing test verdict?

    Yes. Ranex test_slice006_approved_wheel_can_lie.py demonstrates an approved, hash-correct wheel using a pytest11 plugin to force a passing verdict.

    Is Ranex dependency provisioning a finished product feature?

    No. The dependency path is built and tested, but Ranex remains pre-release and does not claim a finished product.

    Try it. Break it. Tell me what broke.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • If You Never Fetched It, You Never Researched It

    If You Never Fetched It, You Never Researched It

    TL;DR: A citation can look researched without proving anyone opened it. Fetch the source bytes, record them, and say what that check still cannot prove. The parent slice log explains why claims need evidence.

    You have written a sentence, attached a source, and felt sure of it without opening the link. A model can do the same thing faster and with more confidence. Neither reaction turns recall into research.

    In this note

    A citation is a shape you can fake

    A URL costs almost nothing to produce. A file with a recorded matching hash requires somebody to obtain bytes.

    Ranex found this problem in its own architecture-decision process. The original research rule required one URL under a prior-art heading. A blog post satisfied it. A link to the Ranex repository satisfied it. A dead link satisfied it. The rule looked like research without requiring anyone to show that they had read an implementation.

    ADR-003 records the corrected rule: a citation is a shape an agent can invent; a file with a matching hash is one it had to obtain. The purpose is to make a factual question answerable from files in the repository, not to measure whether a writer feels well read.

    This is why polished citations deserve the same suspicion as polished summaries. The title can sound right. The URL can point at a respected project. Neither tells you whether the cited code says what the writer claims.

    Fetched evidence turns “I saw this” into a file another reviewer can inspect without trusting the writer.

    For Ranex decisions, each prior-art entry must name a pinned source-file citation, its license, its weakness, and a vendored copy recorded with the blob hash Git would report. The decision requires two distinct sources. The copies live with the decision record and are tracked by git, so a fresh clone contains what the record says it relied on.

    The weakness requirement is doing real work. A source can be mature and still answer a different problem. Naming the weakness makes the reader ask where the borrowed design stops applying before it becomes a hidden assumption in a new system.

    Two sources are not an excuse to collect tabs until the work stalls. ADR-003 calls them a floor on rigor, not a reading quota. The goal is independent implementation evidence, not bibliography theater.

    Use a research checklist you can enforce

    You can apply this discipline to a design doc, an incident note, or a prompt handed to an agent. The machinery can be simpler; the questions should stay sharp.

    • Fetch, do not merely link. Save the actual source you rely on.
    • Pin the version. Cite a commit or release path, not a branch that can move after review.
    • Use independent sources. Do not count the same file twice under different links.
    • Record the license. Copying source is a licensing act, not casual note-taking.
    • Name the weakness. State what the source does not establish for your decision.
    • Make absence block. A cited file that was never fetched is a named failure, not a blank checkbox.

    Ranex tests these clauses against its documents. The record describes refusals for specifications cited instead of working code, moving branch links, duplicate sources, escaped evidence paths, untracked files, and more. Those checks are useful because they turn common shortcuts into visible failures.

    Say where the proof stops

    Fetching and hashing a file does not prove that the file came from the URL beside it. Keep that limit in the sentence.

    An offline repository check can verify the tracked bytes and their recorded hash. It cannot independently contact the cited host and establish provenance. ADR-003 explicitly records that an agent could vendor a file it wrote itself and give that file its true hash. Closing that gap requires a second, independent network fetch.

    That admission does not make fetched evidence pointless. It tells you exactly what it catches: citing from memory or inventing a source without obtaining its bytes. It does not catch a determined writer who controls both the citation and the vendored file. A smaller true claim is stronger than a broad claim you cannot test.

    Questions people actually ask

    These answers distinguish a citation someone supplied from source bytes a reviewer can inspect.

    What is the difference between a citation and fetched evidence?

    A citation is a URL or title someone can type from memory. In Ranex, fetched evidence is a vendored source file with a recorded hash that the writer had to obtain.

    Why does Ranex require two fetched sources for a decision?

    ADR-003 requires two distinct pinned source files because one implementation shows only one author choice; two sources are a rigor floor, not a reading quota.

    Does fetching a source prove where it came from?

    No. Ranex says a fetched hash proves recorded bytes exist, not that they came from the cited URL; an independent network fetch would be needed for that claim.

    What happens when no source is fetched?

    Ranex refuses the citation as a named problem. Absence is blocking rather than a default pass or a skipped research entry.

    Open the source before you repeat the claim

    Before you reuse a citation, open the pinned source file. Read the implementation, the license, and the limitation. Save what you relied on so the next reviewer can do the same.

    Ranex is pre-release, with a working verdict path and substantial surrounding work still unbuilt. This research rule is one documented control, not proof that every decision is right. It is a way to make invented authority more expensive.

    Try it. Break it. Tell me what broke.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Building Ranex in the Open: What Gating My Own Code Actually Caught

    Building Ranex in the Open: What Gating My Own Code Actually Caught

    Your build is green. Do you know what that green light proved?

    Not “did the tests pass”: did the thing that says PASS actually run anything, against the code you think it ran against? I couldn’t answer that about my own repository, so I pointed the verifier at itself. Ranex gates Ranex: the kernel evaluates this repository’s own test suite: 943 frozen test IDs, run provisioned, sealed and offline against the real current commit. It didn’t validate the design. It kept catching me. What it caught is below, so you can go looking for the same shapes in your own build.

    Why dogfooding a verifier is different

    Diagram: two dashboards, one showing 37 of 47 checks green while secrets were readable, the other 47 of 47 green while the gate ran nothing.
    Two real states from this build. Neither was caught by reading a report.

    Most dogfooding is a marketing exercise. “We use our own product!” Fine. But a verification tool pointed at itself has a sharper property: every time it catches something, it’s evidence the thing works. Every time it misses something, that’s a defect in the product itself, and in the code.

    Which makes the interesting output the misses, not the passes. I have 17 years in this industry. Four of those were at Pantheon, working platform-level problems on sites where downtime was measured in money. I’ve seen a lot of failures. None of that stopped the kernel from finding how many of them were mine.

    The green light that proves nothing

    Two states from this build are worth sitting with, because you can check for both of them yourself.

    In one, 37 of 47 checks were green while secrets remained readable.

    In another, all 47 were green while the gate had actually run none of them.

    Read that second one again. A perfect score, from a gate that executed nothing. Now ask what your own gate would report if it were in that state.

    Here’s the deal: the question was never “was the AI smart enough.” It was:

    Does this green light actually prove the proposition we think it proves?

    That problem exists with AI, humans, CI systems, security scanners, tests and auditors. It predates language models by decades, which means you had it before you had an agent. It’s the reason I think the durable idea here is meaningful evidence: not guardrails, not linting, not hallucination detection.

    Three that made it into the record

    Each of these is a closed slice in the repository, with the failure written down rather than smoothed over. Read them as three shapes to go hunting for in your own pipeline.

    A flaky suite that two reviewers approved

    A test suite running at roughly 12% instability was reported stable and approved by two independent reviewers. It was caught only because every gate is re-run against the worktree on disk rather than read from the session that produced it. The cause was a fixture swallowing the stderr that would have said so.

    Neither reviewer was careless. The report simply wasn’t the artifact.

    59 refusals that no test executed

    A slice closed on a cleanup control that had never worked on any supported Python, covered by a test that monkeypatched out the very function it was named for. Measuring the general form of that found 59 refusals no test executed at all.

    The safety net had holes I’d have sworn weren’t there.

    A skipped test reading as a passed test

    Exit-code satisfaction let a skipped or vanished test read as success. The measured failure destroyed 27 tests while the remainder stayed green, and the gate said fine.

    Absence had to become a first-class outcome.

    What you can take from this

    A check nobody has tried to break is a check nobody knows works. Every one of those defects survived a review that felt thorough at the time. What found them was mechanical re-measurement: mutation testing, re-running gates against disk, treating absence as failure.

    The correction that mattered wasn’t a better prompt or a smarter model. It was structural: read the artifact, not the report. That one costs you nothing to adopt: open the diff instead of the summary, re-run the check yourself instead of reading the log somebody handed you.

    Which, if I’m honest, is the same lesson from my Pantheon years. The incident is rarely what the dashboard says it is.

    Why I publish the failures

    Two reasons.

    First, a verification product that hides its own defects is arguing against itself. If I’ll paper over my failures, why would you believe my PASS?

    Second, this is the genuinely useful content. Nobody needs another post explaining what CI is. The specific way a green light lied to me. That’s worth your ten minutes.

    The kernel is MIT-licensed. Every slice above is in docs/slices/done/ with its own record, including the parts that make me look bad.

    Questions people actually ask

    What does ‘Ranex gates Ranex’ mean?

    The kernel evaluates this repository’s own test suite: 943 frozen test IDs, run provisioned, sealed and offline against the real current commit. Every catch is evidence the thing works; every miss is a defect in the product.

    Why publish your own failures?

    Because a verification product that hides its own defects argues against itself. If I will paper over my failures, there is no reason to believe my PASS.

    What was the worst thing you found?

    All 47 checks green while the gate had actually run none of them. A perfect score from a gate that executed nothing.

    So here’s your move: take the check you trust most and try to make it lie to you. Delete a test file and see whether the gate notices it’s gone. Point a required check at nothing and see whether it fails or shrugs. Whichever one refuses to go red is the one that was never protecting you.

    Try it. Break it. Tell me what broke. If you find a hole in the code that decides pass or fail, that’s a contribution, and I’d rather hear it from you than find it in production.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Why AI-Built Software Needs an Accountability Apparatus

    Why AI-Built Software Needs an Accountability Apparatus

    When you hire an architect, you don’t rely on them being good. You rely on licensure, liability insurance, building codes, and inspectors: an accountability apparatus that exists entirely outside the architect.

    AI labour has none of that. Ranex is my attempt to build the missing apparatus: the code, the inspector, and the record.

    This is the argument behind the whole project, and it’s deliberately not the argument you usually hear.

    The framing that gets weaker every year

    The more capable the agent, the larger its blast radius, so bounded authority matters more, not less.

    Most tooling in this space rests on a premise like:

    AI is unreliable, therefore you need guardrails.

    I don’t build on that, because it’s a bet against the models improving. And they will improve. They’ll need fewer retries, understand larger repositories, write better tests, and hallucinate dependencies less often. Any product whose value comes from compensating for temporary model deficiencies is on a clock.

    The framing that gets stronger

    The premise I actually build on:

    Trust is the wrong axis. What matters is whether claims can be independently established.

    Follow the chain:

    1. AI is nondeterministic delegated labour.
    2. Actions have consequences.
    3. Therefore authority must be bounded.
    4. Claims require independent evidence.
    5. Consequential transitions require deterministic gates.
    Diagram: a four-step chain from nondeterministic delegated labour to bounded authority and independent evidence.
    Each step follows from the one above it. Better models don’t break any of them.

    A ten-times-better model doesn’t break a single arrow in that chain.

    The thought experiment

    Suppose we get a coding model that’s correct 99.999% of the time. It understands the whole repository, writes excellent tests, understands security, rarely hallucinates, and can deliver six months of engineering overnight.

    Would you hand it production database admin, deploy keys, npm publish rights, payment credentials, and customer PII, with the instruction “do whatever you think is appropriate”?

    I wouldn’t. And notice why: not because I doubt its competence. Because its blast radius grew. The more capable the agent, the less comfortable unbounded authority becomes.

    This isn’t a hypothetical position. Sandboxing, constrained network access, approval requirements for higher-risk actions, and agent telemetry for auditing are how serious coding agents are already deployed. Capability is increasing at the same time execution controls are increasing. Those aren’t contradictory trends; they’re the same trend.

    Separation of duties survives intelligence

    Imagine an AI genuinely better at software engineering than any human alive. Why should the same actor be able to:

    define success
    implement solution
    modify verification
    run verification
    declare success
    publish result

    Intelligence doesn’t solve separation of duties. Banks don’t abandon accounting controls when they hire smarter accountants. This is why no self-approval is an invariant in Ranex rather than a configurable policy, though approver identity is unauthenticated today, so that check compares unverified strings.

    Provenance matters more as generation gets cheaper

    If agents produce far more code than humans inspect, the useful question changes shape. It stops being “did someone review line 912?” and becomes:

    What requirement caused this code to exist, what agent created it, what authority did it have, what tests established conformance, what exact artifact was evaluated, and what allowed it to ship?

    That’s a supply-chain and verification problem, not an LLM-quality problem. It doesn’t get solved by a better model. It gets solved by binding evidence to artifacts and keeping a record that resists ordinary rewrites, though it doesn’t yet catch a rollback or truncation of the journal itself.

    The green light that proves nothing

    The deepest version of this problem isn’t about AI at all.

    While building Ranex I hit two states worth sitting with. In one, 37 of 47 checks were green while secrets remained readable. In another, all 47 were green while the gate had actually run none of them.

    The issue in neither case was “was the AI smart enough.” It was:

    Does this green light actually prove the proposition we think it proves?

    That problem exists with AI, humans, CI systems, security scanners, tests, and auditors. It’s the reason I think the durable idea here is meaningful evidence: not guardrails, not linting, not hallucination detection. I wrote up the specific incidents in the slice log.

    What this means for the shape of the tool

    If the thesis is right, then Ranex shouldn’t care which coding agent you use. The worker port is replaceable by design; the kernel stays outside the loop regardless of what fills it. Better AI creates more autonomous labour, more autonomous labour creates more delegated authority, and more delegated authority increases the value of a trustworthy boundary around it.

    That’s the version I’m building. It’s pre-release and honest about it, but the thesis is the part I’m most confident in.

    Questions people actually ask

    Why do AI coding agents need governance if models keep improving?

    Because the problem is authority, not capability. The more capable an agent becomes, the larger its blast radius, so bounded authority matters more, not less. Better models do not remove the need for separation of duties.

    Isn’t ‘AI is unreliable’ the real argument for guardrails?

    That argument gets weaker every year. The durable one is that claims require independent evidence: AI is nondeterministic delegated labour, actions have consequences, so authority must be bounded and consequential transitions gated.

    What is the difference between trusting AI and verifying it?

    Trust is a prediction about future behaviour. Verification is a statement about a specific artifact. Ranex is built on the second because it holds regardless of how good the model gets.

    Try it. Break it. Tell me what broke. The kernel is MIT-licensed. If you find a hole in the code that decides pass or fail, that is a contribution.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • How Ranex Judges AI-Written Code: The Kernel, Explained

    How Ranex Judges AI-Written Code: The Kernel, Explained

    Your agent reports “done: all tests pass.” Do you believe it?

    Nothing in that sentence is evidence, and the cost of finding out lands on you, later. I’ve been building with AI coding assistants for years, and the failure that kept costing me time was never that the model wrote bad code. It was that the model told me it was done, and I believed it. This post is the mechanism I built so I don’t have to, written out in enough detail that you can judge whether it would hold up against your own agent.

    Ranex is a kernel (ordinary, inspectable code) that stays outside the AI’s loop and judges every step of its work. It never asks a model what to do next. Rules an agent can read are suggestions; rules compiled into code are constraints.

    The problem is not that AI writes bad code

    An AI writing software is a blindfolded dart thrower with a guide shouting coordinates. Two things go wrong, and they’re separate problems:

    1. The thrower is blind. It cannot perceive whether its own dart landed, so it reports success either way.
    2. The guide is bad. The coordinates were wrong or vague before the throw.

    There’s a third failure, and it’s the most common one:

    Most tools let the thrower paint the bullseye around the dart after it lands.

    One actor writes the code, writes the test, and declares success. That’s why “all tests pass” from an AI means so little: the target moved to wherever the dart went.

    Notice that none of this gets fixed by a better model. A more capable agent paints a more convincing bullseye, so upgrading the model you point at your repo does not touch this. That’s why I stopped trying to improve the throw and started working on the scoring.

    Three ports, and only one produces a verdict

    The architecture is deliberately boring:

    • Model port: one completion, forced structured output. Intake, review, translating machine state into plain language. Stateless.
    • Worker port: an agent with its own loop and tools, running in an isolated git worktree. Returns a diff. Replaceable by design.
    • Check port: the only thing whose output counts.
    Diagram: model, worker and check ports feeding a central Ranex kernel, which emits a verdict into an append-only journal.
    Three ports around one kernel. Only the check port produces a verdict.

    Models appear in exactly three roles (proposer, critic, translator), and none of them can pass a gate. A proposer produces a proposal. A critic produces a finding. A translator produces text. None of them decides.

    It is make for a nondeterministic compiler. make invokes gcc; nobody asks gcc what to build next.

    What a verdict actually is

    A verdict is a pure function of (gate, evidence, subject, approver). Same inputs, same verdict, always. That’s not a design goal, it’s the thing that makes the rest possible. I wrote more about why a verdict has to be a pure function separately.

    Four properties hold on every evaluation:

    • Absence blocks. A required claim with no satisfying evidence is FAIL, never a default, never a skip.
    • Evidence is bound to a subject digest. The same command run against a different commit proves nothing about this one.
    • No self-approval. Whoever produced the evidence cannot approve it, though approver identity is unauthenticated today, so that check compares unverified strings.
    • A gate that cannot block is refused at construction. A non-blocking gate is decoration, so the kernel won’t build one.

    And the invariant that keeps me honest about all of it: removing every model credential from the machine must not change a single verdict. If it would, something in the verdict path is asking a model for its opinion, and that’s a bug. That one is portable, by the way: pull the credentials out of whatever grades your agent today and see whether its answers move.

    The loop, end to end

    take the next ready task
       → create an isolated git worktree
       → spawn a worker with the task envelope
       → wait for it to exit
       → read the DIFF ON DISK  (the worker's own summary is discarded)
       → run the checks         (code, not a model)
       ├─ pass → THE KERNEL merges          (workers never merge)
       └─ fail → retry ×3 with the failure output
                   → still failing → escalate to a human in plain language

    Ranex never trusts the worker, including its own loop’s. It doesn’t need to control what happens inside the loop, only what’s allowed out of it. Containment is a smaller problem than control: the exits are enumerable, the interior is not.

    What stops an agent editing its own tests

    An agent that can edit its own tests will always pass. If the same actor writes the code and the test in your setup, you already know how that ends. Four rules prevent it:

    1. Tests are frozen before building starts. Generated, digested, read-only. Any diff touching a test file fails the gate instantly.
    2. Red-then-green, enforced. Every generated test must fail against the pre-implementation tree. A test that passes before the code exists is not a target. It’s a circle painted around a dart.
    3. Edge coverage as a gate, not a metric. Not “80% of lines” but “every edge in the approved graph has at least one passing test.”
    4. No self-approval. The task that implements a scenario never authors or judges its test, and approver identity is unauthenticated today, so that check compares unverified strings.

    This project applies those rules to itself. The SLICE-001 tests were committed red at b495e3635, before any implementation existed, red-then-green as a fact in the git history rather than a claim in a document.

    What a passing build actually proves

    Precisely this:

    Every behavior on the graph the owner approved has at least one executable test. Every test ran. Every test passed. Here is the evidence, pinned to this exact code digest.

    And these are not claims Ranex makes:

    • That the plan was right. Only the person who owns the target can judge that, and only by using the thing.
    • Anything off the plan. Unspecified behavior is unconstrained. Absence of a requirement is absence of a guarantee.
    • Non-functional properties (performance, accessibility, security) unless you add gates for them.

    I go into this boundary in more detail in what a gate verdict asserts, and what it refuses. “Conformant to an approved specification” is real, defensible, and deliverable. “Correct” is not a claim anybody can make.

    Ranex does not improve aim

    Not by one degree. It makes misses visible and cheap, and hits provable. I don’t claim more than that anywhere: not in the docs, not in the program output, not here.

    Where this actually stands

    Ranex is pre-release. It is not a usable product yet. It’s a kernel with a working verdict path and very little else, and the README says exactly that before it says anything else.

    What works today: evaluate(), subject-bound evidence, absence-blocks, no-self-approval, an append-only hash-chained journal, Ed25519-signed evidence, ranex run, and worker dispatch. The kernel already gates this repository’s own test suite: 943 frozen test IDs, run provisioned, sealed, and offline against the real current commit.

    What doesn’t: flow graphs, scenario compilation, budget, escalation. Designed, not built.

    Roughly speaking: the hardest part to get conceptually right exists, and almost none of the surface around it does. The kernel is MIT-licensed. Read the code that decides pass or fail, and try to break it.

    Questions people actually ask

    What does Ranex actually do?

    It judges AI-written work by evidence and executable checks instead of by the model’s own report. A kernel of ordinary code sits outside the AI’s loop, reads the diff on disk, runs the checks, and produces the verdict. The agent’s summary is discarded.

    Does Ranex make my AI write better code?

    No. Not by one degree. It makes misses visible and cheap, and hits provable. Ranex optimizes the scoring, not the throw.

    Can I use Ranex today?

    Not yet. Ranex is pre-release and the README says so before it says anything else. The verdict path works and already gates Ranex’s own 943-test suite, but flow graphs and scenario compilation are designed rather than built.

    Does it need a model or cloud service to run?

    No. A stated invariant is that removing every model credential from the machine must not change a single verdict. Checks run locally against your code.

    Three questions you can put to whatever grades your agent right now, whether or not it is this one. Pull the model credentials: does the verdict change? Remove a required piece of evidence: does it fail, or shrug? Ask who signed off: is it the same actor that wrote the code? The answers tell you what your green light is worth.

    Try it. Break it. Tell me what broke. The kernel is MIT-licensed. If you find a hole in the code that decides pass or fail, that is a contribution.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.