Author: Anthony Garces

  • Why a Verdict Has to Be a Pure Function

    Why a Verdict Has to Be a Pure Function

    TL;DR: If your green result depends on the clock, a network call, or a model’s opinion, you cannot replay it; Ranex makes the verdict a function of gate, evidence, subject, and approver. Read how the kernel works.

    You inherit a green check, then a release breaks. Now you have a hard question. Did the check prove anything, or did the conditions happen to be friendly when it ran?

    A verdict must not have a mood. The kernel’s working evaluator takes a gate, evidence, a subject, and an approver. Give it those same inputs and it returns the same verdict, always. That is a narrow promise. It is also the foundation beneath a result you can inspect later.

    In this note

    A verdict must not have a mood

    A pure verdict answers from its inputs alone. It does not ask what time it is, call a service, read a model response, or change its answer because a different operator ran it.

    That sounds strict because it is. A verdict that reaches outside its inputs is partly a report about the moment it ran. The code could be unchanged while the result moves because a service was down, a credential was present, or a model phrased an answer differently. You are left arguing about atmosphere instead of examining proof.

    Ranex states the rule in executable terms. evaluate() is a pure function of (gate, evidence, subject, approver). The repository also names a useful test for the boundary: removing every model credential from the machine must not change a verdict. If credentials can move the outcome, a model is inside the judging path. That is not an assistant helping explain a result. That is an unrecorded decider.

    The topology matters here. The model port can propose, criticise, or translate text. The worker port returns a diff from its isolated worktree. The check port produces the outputs that count. None of those model roles can pass a gate. The kernel reads the evidence and applies the rule. Your agent can be inventive where invention belongs; it cannot make the ruler bend.

    This does not make a verdict friendly. It makes it legible. When it says FAIL, you can ask which input failed to satisfy the gate. When it says PASS, you can preserve the inputs and ask the same question again later. There is no confidence score to interpret and no model memory to guess at.

    Purity draws a hard boundary

    Purity does not mean every part of building software is predictable. It means nondeterministic work has a boundary, and the judgment on the other side does not inherit its chaos.

    An agent can write different code from the same request. It can take different amounts of time. It can fail before it succeeds. Ranex does not claim otherwise. The determinism ledger separates those facts from the mechanical chain: graph to covering paths, paths to scenario text, test plus code to result, results plus rules to verdict, and journal to full replay. Each left-hand transformation is intended to be a pure function.

    Generated code is not the proof. Evidence tied to a subject and rules applied to that evidence are the proof. The approved flow graph is the root of trust in the designed larger loop; graph compilation and the surrounding product flow remain designed rather than built. The current kernel has a working verdict path, not the whole picture around it.

    That boundary stops a familiar trick. A worker can say that the diff is correct. A model can say the tests look convincing. Neither statement is a verdict. The kernel needs evidence that satisfies the gate for the precise subject being judged, and it applies the same evaluation rather than trusting a self-report.

    It also keeps a hard distinction between what a passing build proves and what it does not. A passing build supports the claim that approved graph behavior has executable tests, those tests ran and passed, and the evidence is pinned to an exact code digest. It does not prove the graph was right, behavior outside that graph, or properties such as performance, accessibility, or security unless separate gates check them. Saying less is not a weakness. It is how the verdict keeps its meaning.

    Check your own verdict path

    You can inspect whether your own pipeline has this boundary before replacing anything. Start with the green result you least trust. Ask what could change it without changing the commit.

    • Can the verdict read the clock, a network response, or a model opinion?
    • Can you name the gate, the evidence, the exact subject, and the approver it used?
    • Would the same recorded inputs return the same answer on another machine?
    • Does evidence from an older subject stop counting after the tree changes?
    • Can the person who produced evidence also approve it?
    • Can you remove model credentials and get the same verdict?

    That list is not a purity contest. It is an incident drill. If one answer is unclear, your next disagreement over a release will be harder because the system did not retain a stable question to ask.

    Notice the link between this rule and absence blocks. A pure evaluator is not useful if missing evidence silently becomes success. The evaluator needs stable inputs, and it needs to fail when a required input cannot satisfy a required claim.

    Proof is a record you can replay

    Purity buys replay, audit, and freedom from a judge’s mood. It does not bless bad evidence; it makes the question of evidence impossible to hide.

    Subject-bound evidence is part of that discipline. The same command against a different commit proves nothing about the current one, so stale evidence stops counting. No self-approval is another part: whoever produced the evidence cannot approve it, though approver identity is unauthenticated today, so that check compares unverified strings. These are admission and gate rules around the pure evaluator, not decorations attached after a green result.

    The repository’s Status section says the project is pre-release and describes the evaluator as working today. It also says the full governed loop is not all built. Read that as a boundary, not a sales pitch. A small, repeatable verdict path is useful precisely because it does not pretend to settle every question about a system.

    When someone challenges a PASS, the useful answer is not “trust the process.” It is “here are the gate, evidence, subject, and approver; run the evaluation again.” If the answer changes, the record or the evaluator changed. Either way, you have something concrete to investigate.

    Questions people actually ask

    What is a pure-function verdict?

    Ranex evaluates a pure-function verdict from gate, evidence, subject, and approver, so the same inputs produce the same verdict.

    Why must a software verdict be deterministic?

    A deterministic verdict lets an operator replay the same recorded inputs instead of trusting the conditions of one run.

    Is AI-generated code deterministic?

    Ranex says generated code is not deterministic; determinism describes the process and verdict, not the code an agent writes.

    Try it. Break it. Tell me what broke. Read the Ranex repository, then try the credential-removal test on a verdict you rely on.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Gitea 1.27.3, SiYuan 3.8.2 and Keycloak 26.7.3: upgrade checks

    Gitea 1.27.3, SiYuan 3.8.2 and Keycloak 26.7.3: upgrade checks

    Running Gitea, SiYuan, or Keycloak? These three releases contain security-related changes worth checking against your installed version and configuration. Start with the relevant row, then use the upgrade worksheet to record your rollout decision.

    This is a dated release brief originally published September 3, 2026, with release sources reopened and a Gitea startup test added September 8. These are the versions covered by the article, not a claim that they remain the newest available releases.

    Release covered What to investigate Next action
    Gitea 1.27.3 Package access, repository permissions, and Actions trust boundaries. Map the linked SECURITY entries to your features; test representative repository workflows in staging.
    SiYuan 3.8.2 The stored-XSS advisory described below. Compare your version and install method with the advisory’s affected range and fixed release.
    Keycloak 26.7.3 Authentication, token exchange, and delegated administration fixes. Read the advisories relevant to your configuration and exercise your login and authorization flows.

    Download the Gitea check and upgrade worksheet (ZIP). It includes the exact test script, image digest, recorded result, and an editable rollout checklist.

    What the upstream sources say

    Gitea v1.27.3

    Published 2026-08-29T17:42:17Z, not a prerelease. The release notes lead with a SECURITY section, not ENHANCEMENTS: package token-scope handling, attachment path enforcement, same-repo issue access in markup, Actions artifact signatures, hiding limited users, fork-PR trust boundaries, hook permissions, and repo-creation token authorization all get fixes. No CVE is assigned on the page, and there's no CVSS score attached.

    If you self-host Gitea, this one is yours to read. The release page doesn't enumerate which older versions each fix covers, so open the SECURITY list and check it against what your instance actually exposes. Gitea Cloud instances auto-upgrade in a maintenance window, so that's not your call to make. The license is unchanged: MIT at the v1.27.3 tag.

    Primary source: Gitea v1.27.3 release notes

    SiYuan v3.8.2

    Published 2026-08-30T04:36:15Z. The changelog lists one bugfix line ("Some security vulnerabilities") linked to issue #18837, which indexes several GHSAs closed into this release. The one worth naming is GHSA-7h8j-qw37-w46g: Critical, CVSS 9.0, an incomplete asset blocklist that let a previously-fixed stored-XSS issue come back. Affected versions are <= 3.8.1; 3.8.2 patches it. The advisory was published 2026-08-19 and carries no CVE ID.

    The advisory notes that desktop builds bundling pandoc carry higher impact than the official Docker container, which ships without pandoc at all. That distinction matters if you're choosing between install methods, and not only whether to upgrade. License is unchanged: AGPL v3.

    Primary source: GHSA-7h8j-qw37-w46g advisory

    Keycloak 26.7.3

    Published 2026-08-31T09:29:37Z, not a prerelease. This is a CVE batch, not a single fix. CVE-2026-35563 covers an LDAP client that doesn't verify the server certificate hostname. CVE-2026-16093 is a signed-JWT assertion policy bypass through unsigned headers. CVE-2026-16089 lets authorization codes get retargeted to a different client session. The same notes list token-exchange tenant and hosted-domain bypasses (CVE-2026-18215 and CVE-2026-18214), plus bugs in FGAP v2 authorization.

    The release page doesn't say which earlier patch levels each CVE reaches. What it does name is components: the security-fix rows carry labels for OIDC, LDAP, organizations, token-exchange, and FGAP v2. If your Keycloak leans on any of those, this is your release. Keycloak links a migration guide alongside the notes. I didn't re-fetch the individual GHSAs for this one; the release page was the primary source I checked.

    Primary source: Keycloak 26.7.3 release notes

    What we actually tested: Gitea fresh startup

    On September 8, 2026 at 12:31 UTC, an automated local check started a fresh Gitea 1.27.3 container on amd64 with Docker 29.7.2, SQLite, one CPU, and a 512 MB memory limit. The test container had no external network, published ports, or host-directory mounts. Its data was disposable.

    The image was pinned to:

    docker.gitea.com/gitea@sha256:87a67ee09d3ae0d1df5fda5dcda3e2a1f9236a45b0a59025d6e00e46adc43bef
    Check Observed result
    GET /api/v1/version {"version":"1.27.3"}
    GET /api/healthz Overall status pass; database:ping and cache:ping passed.
    Cleanup The test script removed its disposable container and anonymous volumes.

    To repeat this startup check, download and inspect the kit, then run the following from its extracted directory. It requires Python 3, a running Docker daemon, and network access to download the image.

    python3 gitea-smoke.py

    Test boundary: this demonstrates fresh startup and the reported endpoint responses. It does not verify an upgrade from an older database, backup/restore, repository workflows, or any security fix. SiYuan and Keycloak were not run locally for this article.

    Before upgrading an existing installation

    1. Record the starting point. Save your installed version, image digest, deployment method, enabled integrations, and the exact target release.
    2. Read the matching guidance. Follow the release notes to relevant advisories and migration documentation. An affected-version range from one advisory does not describe every fix in a release.
    3. Prove recovery. Follow the project’s backup procedure and restore into a separate instance. For Gitea, start with the official backup and restore guide.
    4. Exercise your workflows in staging. For a Git forge, include login, clone, push, permissions, and any Actions or package workflows you use. For identity software, include the login and authorization paths your applications depend on.
    5. Write the rollback and monitoring plan. Record the owner, maintenance window, recovery steps, and post-upgrade checks. Do not assume that switching back to an old container image reverses a database migration.

    A startup check is a useful first gate. Your staging results and demonstrated recovery path should determine whether the rollout is ready.

    Evidence and disclosures

    The release sections report upstream documentation and advisories; the separate Gitea section reports a local automated test. AI assistance was used for research, editing, and the test. Human review of the September 8 revision has not been recorded. Ranex publishes this site; no production-readiness or vulnerability-remediation certification is implied.

    For a project you have not adopted yet, start with six repository checks beyond stars. Browse more guides at Open Source.

    FAQ

    Does the Gitea startup test prove an upgrade is safe?

    No. It checks a fresh SQLite startup, the version endpoint, and the built-in health endpoint. Upgrade, restore, repository workflows, and security fixes were not tested.

    Are these the latest releases?

    This article covers the named releases from its September 3, 2026 brief. Check the upstream release pages and advisories for newer versions before planning an upgrade.

  • Before You Adopt a New Repository: Six Checks Beyond Stars

    Before You Adopt a New Repository: Six Checks Beyond Stars

    A repository link lands in your team chat the evening before sprint planning. A teammate asks whether the team should depend on it. The only signal anyone has mentioned is its star count.

    You have no inside knowledge and no time for a full audit. What can you bring to the meeting that is more useful than another opinion?

    GitHub's own records can give you a compact evidence sheet. The point is not to manufacture a score. It is to replace one attention number with six observations whose limits are visible.

    Attention is not an adoption decision

    A star count cannot tell you what depending on a repository will be like. Neither can any single replacement metric. Commit activity can be shallow or imported. Issue totals can hide very different conversations. A license label can describe one detected file without settling the status of every file in the tree.

    So keep the evidence separate from the decision. Reports come from the worker. Verdicts come from a judge.

    In this case, your team is the judge. Your constraints (where the dependency would sit, how costly replacement would be, what licensing review you require, and how much uncertainty you can absorb) determine the verdict.

    The six checks

    1. History depth

    When does the visible history begin, and what shape does it have? Start with the repository creation date. Then inspect the newest and oldest entries in the default branch's commit listing.

    If you request one commit per page, the pagination header can show how many pages GitHub enumerated at that moment. Treat that number as a snapshot, not a permanent total. History can change, and a reachable branch does not prove that no other roots or orphaned history exist.

    2. Issue and pull request evidence

    Issue and pull request totals are inventory. They do not explain what the discussions contain, how reviews work, or what a typical response looks like.

    Could you open at least one merged pull request and read it end to end? Note who submitted it, the number of commits, the files touched, the recorded review comments, and the open-to-merge timestamps. One example still cannot establish a pattern, but it gives the totals concrete context.

    3. Release record

    List the releases and record exactly what GitHub returns: versions, dates, and draft or prerelease status. Resist the urge to turn the list into a claim about stability or discipline. A release record proves that releases are recorded. Compatibility and support require different evidence.

    4. License classification and scope

    GitHub's license endpoint can return an SPDX classification, a path, and a blob SHA for the detected license file. That is useful: it identifies the file your team needs to inspect and gives you a specific object to recheck.

    It does not prove that every file in the repository is covered by the same terms. It is not legal advice. If licensing matters to your use, this check defines the next question instead of pretending to answer it.

    5. Contributor-attribution concentration

    The contributors endpoint attributes contributions to accounts. Add the returned contributions, then calculate the share held by the top account and top few accounts.

    Write the result without turning it into a resilience claim. Accounts are not necessarily unique people, anonymous work may be missing, and an attribution snapshot cannot tell you who can maintain the project tomorrow.

    6. Initial-commit provenance clues

    Open the oldest listed commit and read the object itself. How many parents does it have? How large are the recorded additions and deletions? Does GitHub mark its signature as verified? Does the commit message contain trailers?

    These are provenance clues. Record them precisely. On their own, they do not prove authorship, misconduct, or code quality.

    Animated diagram of a repository health check: a commit graph is scanned while six verification checks tick into place

    Worked example: THU-MAIC/OpenMAIC

    Checked at: September 2, 2026. The following is a point-in-time reading of public GitHub sources, not an adoption recommendation.

    The repository metadata reported that OpenMAIC was created on 2026-03-11, uses main as its default branch, was not archived, was not a fork, and had a pushed_at date of 2026-09-02. Its README describes an immersive multi-agent learning and course-creation platform and includes a v1.0.0 section. That is the project's documentation describing itself; it does not verify what the software can do.

    For history depth, the commit listing at one item per page linked to page 503 as its last page and returned f760f58 as the newest commit, dated 2026-09-02, with message feat(render-service): machine-readable admission state (#1351). The last linked page returned 0d20abf, dated 2026-03-12, as the oldest listed commit in that snapshot. This is evidence of the visible history GitHub enumerated then, not a forever-exact commit count.

    The issue search returned 537 issues, while the pull request search returned 763 pull requests. Those volatile totals say how much inventory GitHub search reported; they do not say what the inventory means.

    To ground that inventory, pull request 1273 was contributor-authored and recorded 13 commits, 21 changed files, eight review comments, and two comments. It was opened on August 28 and merged on August 31. That is one sampled merged pull request. It cannot establish the repository's typical response time or review process.

    The latest-release endpoint returned v1.0.0, published on 2026-08-27, with both draft and prerelease set to false. The returned releases list contained nine non-draft, non-prerelease releases, from v0.1.0 on 2026-03-26 through v1.0.0. That is the release record visible at the check time, nothing more.

    The license endpoint classified the detected LICENSE file as MIT and returned its blob SHA. The classification does not establish license coverage for every file, and this article does not provide legal advice.

    The contributors endpoint returned 62 accounts and 481 attributed contributions. The top account held 50.9% of that returned total; the top three held 72.3%. Those percentages describe concentration in that API snapshot. They do not count unique humans or establish future maintainability.

    Finally, the oldest listed commit object showed zero parents, 125,393 additions, zero deletions, a signature status of verified: false, and a message trailer naming Claude Opus 4.6 as a co-author. Those facts describe the object GitHub returned. They do not tell you who wrote which code, whether anything improper happened, or whether the code fits your needs.

    Ranex has no disclosed affiliation with OpenMAIC.

    What do you take into sprint planning?

    You now have six bounded observations instead of a single attention signal:

    1. the visible history and its dates;
    2. issue and pull request inventory plus one inspected example;
    3. the release record as listed;
    4. the detected license classification and its limits;
    5. contributor-attribution concentration at one point in time; and
    6. initial-commit provenance clues.

    Could you run those checks before the sprint commitment, then decide against your own constraints? If an AI assistant supplied any repository facts along the way, close that loop too: here is how to verify which source your AI cited.

    FAQ

    How do I assess a new GitHub repository before adopting it?

    Replace the star count with six recorded observations: history depth, issue and pull request evidence, the release record, license classification and scope, contributor-attribution concentration, and root-commit provenance clues.

    Why aren't issue and PR counts enough on their own?

    Counts are inventory, not evidence. Open at least one merged pull request and inspect its commits, changed files, review comments, and timeline. One sample cannot establish a pattern, but it gives the totals concrete context.

    What does the license check actually tell you?

    GitHub's license endpoint classifies a detected license file and returns its path and blob SHA. It does not confirm that every file is covered, and it is not legal advice.

    What are initial-commit provenance clues?

    They are facts read from the root commit, such as its parents, additions and deletions, signature status, and message trailers. Alone, they support no conclusion about authorship, misconduct, or code quality.

    Research and drafting were AI-assisted. Human review is required before publication.

  • Notion alternatives: compare AFFiNE, Docmost and SiYuan

    Notion alternatives: compare AFFiNE, Docmost and SiYuan

    Start with the workflow you need to move. A team wiki, a visual planning space, and a personal knowledge base need different things from a Notion alternative. Use the comparison below to shortlist AFFiNE, Docmost, or SiYuan, then test a small export before moving your workspace.

    Evidence scope: this is a documentation-based comparison, not a hands-on migration benchmark. The original feature table is a September 2, 2026 snapshot. The September 8 revision adds a migration acceptance kit and rechecks the linked Docmost import guide and AFFiNE project overview. No candidate has been run through the full acceptance kit for this article.

    Which candidate should you investigate first?

    • Team documentation: investigate Docmost if moving a page-based workspace is the priority. Its import guide documents a Notion ZIP import path. That makes it a useful pilot candidate; it does not establish that your databases or permissions will survive.
    • Documents and visual planning: investigate AFFiNE if the workflow combines writing with a canvas. Its project overview describes documents and whiteboards together. Include canvas content in your exit test.
    • Personal knowledge work: the original SiYuan assessment below focuses on a personal workspace. Verify its current collaboration and access model against your team requirements before extending that fit to shared work.

    Download the workspace migration acceptance kit (ZIP): an editable CSV worksheet and a step-by-step pilot plan. All result cells start as “not tested.” This is not a Notion export file.

    What "alternative" actually has to match

    An alternative does not have to match Notion's feature page. It has to match your workflows: the meeting notes, the wiki, the database your team actually opens every day. Before you compare candidates, it is worth knowing exactly how portable your current content is, from Notion's own documentation.

    Workspace export covers HTML, Markdown, and CSV for databases, plus your uploaded files. That is the good news. The fine print matters more. Notion states plainly that you can't instantly recreate your workspace by reuploading exported content. Download links expire after 7 days, and exports can take up to 30 hours to process. On Enterprise plans, workspace or teamspace owners can disable member exporting entirely. And workspace PDF export is being removed, with rollout through 31 August 2026.

    None of this is a complaint. It is documented behavior, and it sets the bar: any candidate you evaluate should export your data in a form you can re-import elsewhere, on your schedule.

    The original product observations below were checked on September 2, 2026; they are not a claim about the latest release or edition today.

    A checklist that works for any candidate

    Run these five checks before you fall in love with a demo. They apply to any self-hosted tool, the three below included.

    1. License. Read the LICENSE file, not the badge. Monorepos sometimes split licenses by directory, and "core" licenses can hide enterprise-gated features.
    2. Deployment model. Find the official install path. Docker Compose, Kubernetes, single binary. Note every dependency it drags in: databases, caches, proxies.
    3. Data export. What formats come out, and can a restore get your content back? A backup that only restores one component is not a full exit.
    4. Project activity. Check the releases page and issue tracker. Look at dates, not stars. Recent, steady releases beat a big number that stopped moving.
    5. Maintenance burden. Who upgrades it, backs it up, and watches it at 2 a.m.? If the answer is "nobody yet", that is your answer.

    Animated comparison of three open-source candidates: feature bars fill across three equal cards under a settling balance beam

    Documentation comparison: September 2, 2026 snapshot

    I applied the checklist to three candidates on September 2, 2026. Sources: the AFFiNE repository and self-host docs, the Docmost repository and installation docs, and the SiYuan repository with its releases.

    AFFiNE Docmost SiYuan
    License CE described as MIT, but split LICENSE: backend carries a separate license; EE needs a subscription for production AGPL-3.0 core; enterprise features in ee directories under an enterprise license AGPL v3; most features free for commercial use, some paid
    Deploy Docker Compose recommended; Postgres only, Redis required; 4 cores, 2 GB RAM minimum Docker Compose recommended; Postgres 18, Redis 8, APP_SECRET of 32+ characters; claims air-gapped operation Official Docker image b3log/siyuan on port 6806, workspace bind mount, accessAuthCode required
    Data export Official backup is pg_dump plus blobs and config; database-only restore loses uploaded files Markdown or HTML ZIP export; Notion ZIP import; Confluence, PDF, DOCX import are Enterprise Markdown with assets, PDF, Word, HTML; .sy JSON on disk. But Docker hosting cannot export PDF, HTML, or Word
    Latest release v0.27.4, 18 Aug 2026; 72.1k stars, 614 issues v0.95.0, 3 Jul 2026, security fixes; 21.6k stars, 223 issues v3.8.2, 30 Aug 2026; 46.1k stars, 79 open issues
    Self-host burden Postgres, Redis, volumes, config; 0.27.4 changed Compose paths and dropped .env, so guides rot fast Postgres, Redis, WebSockets at the proxy for the realtime editor Lightest of the three; but auth is a lock-screen access code and it is positioned as personal-first, not a multi-tenant workspace

    Notion stays as the context row in your head: whatever you pick, the migration path out of it is the export behavior described above.

    Why the usual suspects are missing

    You will notice some famous names absent. That is the checklist working, not an oversight.

    AppFlowy-Cloud was archived on 1 September 2026. The README calls it legacy and unmaintained, and says production self-hosting is a closed-source commercial fork with a free tier of one user seat. Outline's LICENSE is Business Source License 1.1, which states in its own text that it is not an Open Source license (its change date to Apache-2.0 is 2030-09-01). Wiki.js 3.0.0-beta.537 is explicitly marked not for production use; v2.5.314 is the production latest.

    Run a small migration acceptance test

    Use a disposable workspace containing no sensitive material. Record the candidate version and edition; repeat the same checks for every candidate. Decide which checks are mandatory before testing.

    1. Prepare a representative sample. Include a parent page, two child pages, internal links, an image, an attachment, and a small database using the views and relations you actually need. Record the original counts and permissions.
    2. Export and import. Use the products’ supported flows. A successful import message is only the beginning: open the imported pages and inspect their content.
    3. Check the workflow. Follow internal links, open attachments, search for a known phrase, and repeat a normal database task. Record missing features and repair time.
    4. Check access with separate users. Confirm that a viewer cannot edit and an uninvited user cannot open private material. Do not assume source permissions migrated.
    5. Test the exit and recovery. Export again and inspect the files independently. Separately, restore a backup into a fresh instance and confirm that text, attachments, and required settings return.
    Decision Evidence to record
    Does the content survive? Counts, broken links, missing attachments, and screenshots of differences.
    Can the team still work? Required workflow results and viewer/editor/private-page checks.
    Can you operate and leave it? Export inspection, restore result, repair minutes, and a named maintenance owner.

    Mark each worksheet row pass, fail, or not tested. If a mandatory workflow fails, keep that candidate out of a full migration until the gap is resolved. An untested requirement stays an open question.

    Our recommendation: pilot the candidate whose documented workflow fits your main use case, and let the mandatory checks decide whether to continue. A feature checklist cannot establish migration fidelity or operating cost.

    For the earlier adoption decision, use the repository evaluation guide. More comparisons and practical guides live at Open Source.

    Caveats

    • Feature tables rot fast. The checked-at date of September 2, 2026 bounds every claim here, including versions and star counts. Re-verify before you rely on any of them.
    • License nuance: Docmost and SiYuan are AGPL-3.0, which carries network-copyleft implications a company should run past counsel. This is not legal advice. AFFiNE CE is described as MIT, but the backend carries a separate license.
    • No candidate here is "the winner". Fit depends on your constraints: team size, compliance, who is on call.

    Disclosures

    • First-party interest: ranex.dev is my site, and this series covers the open-source ecosystem I work in.
    • AI assistance was used in the research and drafting of this post.
    • The September 8, 2026 revision was AI-assisted. Human review of this revision has not been recorded. The migration worksheet is an unexecuted test plan, not a report of a completed migration.

    FAQ

    Is AFFiNE really MIT licensed?

    Mostly, for the Community Edition. AFFiNE CE is described as MIT and free for self-hosting, but the repository LICENSE is split: packages/backend and packages/common/native carry a separate license, and the EE server code requires a valid Enterprise Edition subscription for production use. Read the LICENSE files yourself before committing.

    Which candidate is closest to Notion?

    None is a full replacement. Docmost is the most team-oriented of the three, but its Bases feature (Table and Kanban) is Enterprise-only as of v0.95.0. AFFiNE's official backup procedure is a Postgres pg_dump plus blobs rather than a document export. SiYuan describes itself as a privacy-first personal knowledge system, and its official Docker hosting cannot export PDF, HTML, or Word.

    What about AppFlowy?

    The AppFlowy-Cloud repository was archived on 1 September 2026. Its README calls it legacy and unmaintained, and says production self-hosting is now a closed-source commercial fork with a free tier of one user seat. That is why it is not in the table.

    Will my Notion content migrate cleanly?

    Partly. Non-database pages export as Markdown and full-page databases as CSV, with callouts becoming HTML. But you cannot export all database views at once, Form views cannot be exported, and Notion states you cannot instantly recreate a workspace by re-uploading exported content. Databases, views, and permissions will not round-trip.

  • A Signed Record Is Not a Fact

    A Signed Record Is Not a Fact

    TL;DR: Ed25519 proves who signed specific bytes, not whether the claim is true; verdict policy must come from the evaluated commit. This slice log started with the question that matters: what exactly did that green light prove?

    Your build says PASS, the record is signed, and everyone wants to move on. Before you do, what exactly did that green light prove?

    In this note

    A signed record can be real and still be used to tell you a false story. The failure is in the trust root, not the cryptography.

    I learned this by closing the same slice, reopening it when audits showed the tests were narrower than reality, closing it again, then reopening it a second time. That cost belongs in the record because it gives you something useful to check in your own pipeline.

    A signature proves origin, not truth

    An Ed25519 signature proves that the holder of a private key signed specific bytes. It does not prove that the claim inside those bytes is true.

    That distinction is the whole post. Your CI record can tell you who produced an observation. It cannot, by signature alone, tell you whether the rule that admitted it was trustworthy, whether the command meant what the claim says, or whether an approver was independent.

    In Ranex, evidence is signed and bound to a producer in a committed public keyring. The verifier has public keys, not the private keys used to sign. That is useful when evidence is produced on one machine and verified on another.

    But the first version still had a dangerous opening: the keyring and gate catalog could be read from the working tree. An uncommitted edit could decide a verdict. The project said review of the committed keyring was the control. The code had not made that true.

    Review cannot control bytes that were never committed for review.

    That is a shape worth noticing. Security language can make a weak boundary sound solid: signed, verified, trusted, approved. Ask one plain question instead: which exact bytes decided this result, and who was allowed to choose them?

    The second reopening was the real scar

    The second reopening showed that checking a committed file is not enough. You must also refuse a trust-root path the commit does not carry.

    SLICE-002 was reopened a second time on 2026-08-02 because the trust-root check skipped itself when the evaluated commit had no such path. The path came from a flag. The party being gated could name a catalog or keyring that the commit did not carry, and Ranex would read it unchecked.

    This is the attacker selecting the rulebook after the game, not an edge case.

    The slice record reproduced an attacker-named gate catalog that could rewrite the gate after the work. It also reproduced a keyring at a gitignored path, where a producer could register itself while git status stayed clean. Another route used a committed symlink at a reviewed name: resolution followed the link before Git was asked, so the reviewed name was never the thing being checked.

    All of those routes landed on the same mistake: absence got a pass. The code returned without comparing anything.

    ADR-002 closed that opening by refusing any trust-root path the evaluated ref does not carry. It compares the bytes Git records for the path as named with the bytes that would decide the verdict, then returns committed bytes for the loaders to parse. The loader does not reopen a mutable path afterward.

    That last part matters. A compare followed by a second read is two chances for different bytes to appear. The record measured the change with strace: one open per trust-root file after the fix, where there had been three and two.

    Read the policy from the committed state, not from wherever a flag points. The kernel model only helps when the inputs to its decision are pinned too.

    What to hunt for in your own pipeline

    Look for the files and paths that decide what a PASS means. If a worker, build, or command under test can choose them, your verifier has an opening.

    • Configuration that declares required checks, trusted identities, allowed commands, or approval rules.
    • CLI flags or environment variables that select policy files, keyrings, manifests, or catalogs.
    • Loaders that check one path and then reopen that path later.
    • Symlinks, normalized paths, and ignored files near a policy lookup.
    • Code that treats a missing policy file as an empty policy, a default, or an innocent absence.
    • Tests that edit a committed policy file but never name a policy file the commit does not carry.
    • Green controls that prove ordinary work can still proceed after a refusal is added.

    Do not stop at “the file is in the repository.” Ask whether the evaluated commit carries that exact path. A working tree is a place people work. It is not automatically a trustworthy source of policy.

    And do not mistake a clean status output for evidence that nothing changed. Gitignored inputs can be invisible to the habitual check while still deciding the result.

    Why four audits found what the tests missed

    Independent audits exist because your first test suite is narrower than reality. SLICE-002 closed after 17 defects across four independent audits, not because the first green run had been sufficient.

    The initial close was premature. Two independent auditors found defects and showed that two completed criteria were false. One test compared the same string to itself. Another used a single-claim gate even though the real repository gate required two claims.

    They are the lesson, embarrassing or not. A test can be green because it tested an easier world than the one your tool inhabits.

    The second reopen sharpened it further. The tests had caught an edit to a committed trust root. Nobody had asked what happened when the path was absent from the commit. The test proved the answer to one question. The code needed to survive another.

    There is a simple review prompt hiding here: what input shape did this test not ask about? Missing path. Ignored path. Redirected path. File swap after check. A policy that is syntactically valid but weak. These are not exotic when the input selects the control.

    Ranex is pre-release. The closed slice does not mean the whole system has become a source of truth. Its own record states limits plainly: a signature does not stop a same-user attacker who can read the signing key, and the approver identity is an unauthenticated string. I do not have evidence that either is solved by signing evidence.

    Stating limits that plainly is part of the control. A signed record that looks stronger than it is can cause more damage than an unsigned one, because people stop asking questions.

    Questions people actually ask

    These questions separate signature origin from committed trust-root policy.

    What does an Ed25519 signature prove in an evidence record?

    An Ed25519 signature proves the record was signed by the holder of a registered producer key over the signed fields. An Ed25519 signature does not prove the claim is true or that the approval is authentic.

    Why must a keyring and gate catalog come from the commit?

    The keyring and gate catalog files decide which evidence counts. If the party being judged can point the verifier at unchecked bytes, it can choose the policy that judges it.

    What should happen when a trust-root path is absent from the evaluated commit?

    The verifier should refuse. A file no commit carries was reviewed by nobody, whatever its contents say.

    Make one trust decision visible this week

    Start with one green check you already trust. Identify its policy file, identity list, manifest, or rule catalog. Then answer: does the verifier read committed bytes, or does it read whatever the current process can reach?

    If the answer is the second one, make absence block. Test an uncarried path. Test an ignored one. Test a symlink. Test a swap between validation and load. Keep the ordinary committed path as a green control so “refuse everything” cannot impersonate safety.

    The slice records are in the Ranex repository, under docs/slices/done/. Read the failure before you borrow the fix.

    Try it. Break it. Tell me what broke. If this helped, star the repository and send an honest critique. The useful reply is the one that finds the next unchecked input.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • Your AI Cited Source Code. Can You Prove Which Source?

    Your AI Cited Source Code. Can You Prove Which Source?

    Your coding agent just handed you a fix and backed it up with a quote: "based on the
    implementation of the retry logic in your HTTP client library." It even pasted the
    function. The fix looks right. Then a reviewer leaves one comment on your pull request:

    Which version of the library is that from?

    And there it is. The citation names a project, but not a version. Not a tag, not a
    commit, not an artifact. The function might be from the release you have pinned, from
    main at a different point in time, or from a fork with the same file name. The
    citation alone does not tell you which.

    This isn't a story about careless review. A citation and provenance are two different
    things, and the distinction matters whenever version-specific behavior affects a
    change. That is why this series starts here, alongside the site's other
    field notes.

    A citation is a claim. Provenance is evidence.

    When an agent says "this comes from library X," that's a report. Useful, often correct,
    and unverifiable on its own. Provenance is what turns the report into something a
    reviewer can check: an exact version, a specific artifact, a checksum that either
    matches or doesn't, a commit hash you can visit.

    Reports come from the worker. Verdicts come from a judge.

    The agent is the worker here. The judge is any verification step that someone else can
    rerun and get the same answer. If your review process accepts the report without a way
    to reach a verdict, the question isn't whether a mismatch will slip through. It's
    whether you'll notice when it does.

    Five questions that establish provenance

    You don't need any particular tool for this. For a dependency your AI cited, can you
    answer these five questions with links?

    1. Does the version resolve to exactly one release? "The requests library" is not
      an answer; a specific resolved version is. Range specifiers need to collapse to the
      one version your lockfile actually pins.
    2. Are you looking at the registry artifact or a repo snapshot? The package
      published to a registry and the repository at a matching tag are different evidence
      surfaces. Which one does your build consume, and which one did the citation use?
    3. Does a published checksum verify, and what happens when it doesn't? A check
      that warns and continues is a formality. A check that fails closed, refusing to
      proceed on mismatch, is a control.
    4. If the code comes from git, is the commit pinned and are the file contents
      verified?
      Can you point to the commit hash and check the individual blobs against
      the host's tree data?
    5. Is there a record you can attach to the review? Provenance that lives in your
      terminal history helps nobody next week. A manifest, even a pasted one, makes the
      verdict portable.

    If you can answer all five, the reviewer's "which version?" has an evidence-backed
    answer instead of another assertion.

    Worked example: what the primary records show

    Disclosure first: Leitir is maintained by Anthony Garces, who publishes ranex.dev.
    It appears here as a transparent first-party example of the checklist above: a
    methodology walkthrough, not a ranking win or an independent endorsement.

    Checked on 2026-08-31 (Asia/Manila) against the project's primary sources: the

    Leitir README,
    latest-release endpoint,
    tags,
    project metadata,
    and LICENSE:

    • Leitir describes itself as a deterministic, provenance-bound dependency-source corpus
      plus a deterministic code-search kernel for AI coding agents.
    • Its README documents exact-version resolution, bounded registry source-artifact
      fetching, fail-closed published-checksum verification, pinned git commit trees, blob
      checks, and provenance manifests, the same five questions above, mechanized. I did
      not run those paths for this article.
    • Its README lists support for GitHub, GitLab, Bitbucket, Codeberg, Sourcehut, npm,
      PyPI, crates.io, and Go sources.
    • The README documents leitir info <spec> as one JSON response containing provenance,
      an API summary, examples, trust, and parity.
    • GitHub's latest-release endpoint returned v0.1.6, published 2026-08-25. The tags
      listing maps v0.1.6 to commit
      91e93e16466405bb95fcfb762e9319a5f716de09; that observation does not guarantee the
      tag can never move.
    • The README says distribution is through GitHub releases and tags rather than PyPI.
      Current-main pyproject.toml sets version 0.1.6 and dependencies = []; that does
      not prove anything about optional extras or the contents of the release artifacts.
    • The checked LICENSE file contains MIT License text. GitHub's
      repository metadata reports
      NOASSERTION, so this article does not claim GitHub classifies the repository as
      MIT or that the checked license text covers every file.

    Now the scope of all that, stated exactly: the verification methods described above
    are designed to establish origin binding only. When you run them successfully,
    they can show that specific bytes match a specific published artifact or commit. They
    do not prove software quality, security, correctness, suitability for your project,
    or the identity of the publisher. A perfectly provenanced dependency can still be a
    bad choice. Verification narrows what you have to trust; it doesn't remove trust from
    the picture.

    That's also why the series will use careful wording from here on: a citation either
    has "A scoped Leitir record is linked" or "No Leitir record is linked." Absence
    never means failed verification. It means unverified: a neutral state, and an honest
    one.

    What could you verify this week?

    Could you pick one dependency your AI assistant cited recently and run the five
    questions on it by hand? Can you resolve the exact version, find the registry artifact,
    and verify one checksum? Where does the trail go cold? That cold spot is your
    workflow's actual provenance gap, and it is worth knowing before a reviewer finds it.

    Caveats worth carrying with you: facts above were verified on 2026-08-31. The README
    behavior was not runtime-tested, and release assets were not independently hashed.
    If publication happens later, human review must recheck the version, release, activity,
    and license facts that morning. Follow the primary links rather than trusting this
    snapshot. Treat the method as the takeaway, and evaluate any tool against your own
    constraints.

    Research and drafting were AI-assisted. Human review is required before publication.
    Primary sources are linked so you can check the claims yourself.

    FAQ

    Does a checksum match prove the code is safe?

    No. A verified checksum proves the bytes you examined match match a published artifact:
    origin binding, nothing more. It does not prove quality, security, correctness,
    suitability for your use, or who the publisher really is. Those need separate review.

    What does "No Leitir record is linked" mean?

    Only that no scoped provenance record accompanies the citation, so origin has to be
    established another way. Absence never means failed verification. It means unverified,
    which is a different and honest state.

    Can I check AI source code provenance without any tool?

    Yes. Resolve the exact version, download the registry artifact, verify its published
    checksum, or pin the git commit and compare. It is manual and slow, but every step uses
    primary sources you can link in a review.

    Why disclose that Leitir is a first-party project?

    Because the example only has value if you can weigh the interest behind it. Leitir is
    maintained by the author of this site, so this article is a transparent methodology
    walkthrough, not an independent ranking or endorsement.

  • Cause Is Structure, Not Prose

    Cause Is Structure, Not Prose

    TL;DR: A failure message should explain a cause that already exists as data, not force the next system or person to guess the cause from a sentence. The Ranex slice log keeps the evidence behind those distinctions.

    You open a failed check and get one sentence: “this claim is not satisfied.” You still have to ask the question that matters. Was the work never done, was the evidence for another revision, was it forged, or did the checker refuse to admit it?

    A sentence can be useful to read. It is a poor place to hide the only category that tells you what to do next.

    In this note

    A sentence cannot be the interface

    A cause must survive as structured data because different failures demand different responses. “Absent” is not another spelling of “forged,” and neither is another spelling of “stale.”

    ADR-020 records the specific Ranex failure. The kernel’s _diagnosis() already partitioned unsatisfied claims into five kinds: contradicted, failed, mismatched, stale, and absent. The admission layer added refused and unattributable. Seven events could reach a caller as the same broad outcome: a claim was not satisfied.

    Then the old path discarded the distinction. _diagnosis() joined the kernel buckets into English, Evaluation exposed claim IDs without their cause, and the CLI recomputed then discarded the admission results. A later renderer had two bad choices: parse prose or invent a category it had not received.

    This is not a copywriting complaint. ADR-020 ties prose parsing to a defect that reopened SLICE-002: a forgery could be reported using wording reserved for honest absence. If the words are the interface, an attacker who changes a field can influence the story the operator sees.

    The cause must survive the boundary

    Ranex’s decision is to compute one partition, return it as data, and render the human sentence from that same partition. The structure and the prose then have one source.

    The five kernel claim causes remain at the evaluation boundary. The two admission causes do not enter Evaluation, because pure evaluate() cannot see admission rejections; the projection composes them once. Self-approval is also separate: it is an evaluation-level refusal marker, not a claim cause, and it must render even when missing_claims is empty.

    That placement protects the kernel’s purity. evaluate() remains a pure function of gate, evidence, subject, and approver. It does not learn about UI wording or admission state merely because a screen needs to explain a result.

    The design also preserves compatibility where it earns it. reason must remain byte-identical for existing inputs, because people read it and the journal records it. The structured field is additive, but adding it changes the evaluation record digest for new evaluations. Old journal rows keep their own digests; comparing digests across that declared boundary is not valid.

    Unknown is not the nearest known cause

    An unknown cause must block and render as unclassified. It must not be rounded into the closest familiar category.

    ADR-020 treats the causes as unordered. There is no severity ranking that lets a renderer keep one cause and erase the others. A claim that is contradicted and missing is named once under contradiction; suite detail belongs on a failed cause, not in a new category; and a nullable admission claim_id stays null rather than being coerced into honest absence.

    That discipline also gives the reader an honest response to future change. The wire accepts an unknown tag, the reader shows unclassified, and the result still blocks. The renderer does not guess. It shows the operator that the system has encountered something it cannot yet name safely.

    One exit code cannot offer that honesty. A successful process can conceal skipped, missing, or unexamined work; a failed process can represent many distinct causes. A skip is not a pass follows the related move from exit codes to structured test outcomes.

    What to keep structured

    You do not need a governance kernel to apply this rule. Any boundary that turns a machine result into a person-facing explanation needs a stable category before it needs polished wording.

    • Define the closed causes. Name the states your consumer has to handle rather than relying on message fragments.
    • Attach the cause where it is discovered. Do not calculate it again in an API, CLI, dashboard, or report.
    • Render prose from the data. Keep human language useful without making it the protocol.
    • Make the mapping total. Test every known cause and reject a default arm that silently swallows a new one.
    • Preserve unknowns. Block, disclose, and investigate them instead of assigning the nearest familiar label.
    • Keep distinct layers distinct. A claim failure, an admission rejection, and self-approval can be related without becoming one bucket.
    • Test old wording deliberately. If compatibility matters, assert the existing sentence rather than trusting an incidental refactor.

    Why ranking does not fix it

    Ranking causes does not preserve them. It chooses a winner and throws information away.

    ADR-020 rejects a severity rank over the seven causes. It notes Knative’s approach of ranking many causes into one and discarding what falls below the selected severity. That can be a useful presentation choice in another system, but it is the wrong data model when each cause tells an operator a different next action.

    A renderer also needs exhaustive handling. ADR-020 cites typed string states that downstream code handled inconsistently because a switch’s default arm did not report an unfamiliar value. The corrective move is not to add more prose. It is to validate the closed set at deserialization so no renderer receives a known value it cannot handle.

    Where Ranex stands

    Ranex is pre-release, and the status must be read carefully. ADR-020 is accepted and describes the kernel decision. The README says SLICE-020 closed structured five-kind evaluation causes and self-approval, and its projection composes refused and unattributable rejections.

    That does not make every interface or future board feature complete. The ADR leaves presentation, colour, glyphs, layout, and new causes out of scope. What it establishes is narrower and more useful: the cause should reach the consumer as structure, while the sentence remains a readable rendering of that structure.

    Questions people actually ask

    What does cause is structure, not prose mean?

    ADR-020 requires Ranex to compute a per-claim cause as structured data and render reason from that same partition, rather than asking a later consumer to recover the cause by parsing English.

    What failure causes does Ranex distinguish?

    ADR-020 identifies five kernel claim causes (contradicted, failed, mismatched, stale, and absent) and two admission-layer causes, refused and unattributable; self-approval remains a separate evaluation-level marker.

    Why is parsing a failure message unsafe?

    ADR-020 forbids parsing reason because changed wording can mislabel a forgery as honest absence; the machine-readable category must survive to the renderer.

    Does Ranex ship structured causes today?

    The Ranex README says SLICE-020 closed structured five-kind evaluation causes and self-approval, while its projection composes refused and unattributable rejections; Ranex remains pre-release.

    The next time a failure arrives as one tidy sentence, ask what structured state produced it and whether every consumer can handle that state. Try it. Break it. Tell me what broke.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • A Skipped Test Is Not a Passed Test

    A Skipped Test Is Not a Passed Test

    TL;DR: Exit code zero does not prove required tests ran; compare structured outcomes against a frozen manifest and fail on missing or undeclared results. The parent slice log documents why.

    Your test command exited zero. Which required tests actually ran? If your answer is “the job was green,” you have a hole in the gate. Exit-code satisfaction allowed a skipped test, or a vanished one, to read as success.

    In this note

    The measured failure was not theoretical. Twenty-seven tests were destroyed while the remainder stayed green. The gate had no way to distinguish that partial absence from a healthy run.

    A passing exit code is a promise about process completion. It is not a complete account of the test outcomes you needed.

    Zero exited cleanly while the suite got smaller

    Before SLICE-009, Ranex accepted exit_code == 0 and nothing else for its tests-executed claim. A test that asserted False, then received a skip marker, reported one skipped test, exited zero, and satisfied the claim.

    Pytest refuses a suite with total absence through exit code 5. Partial absence is different. A suite can skip itself into silence or shrink without changing its exit code.

    That exact shape had already happened: two agents in one worktree destroyed 27 tests, and the remainder stayed green. The problem was not that pytest lied. It returned the process status it was designed to return. The gate asked that status to prove more than it could prove.

    An exit code is a promise about nothing your gate has defined. Until you define the required outcomes, zero only tells you that the command chose zero.

    This is broader than agents. A changed test selection, a conditional skip, an environment difference, or a deleted file can all turn your familiar green badge into a smaller test run. A test that runs and sometimes fails is a different disease; that one is worth a read too. If the only evidence you retain is a status code, you cannot compare what ran with what should have run.

    Judge outcomes against a frozen manifest

    The fix is to bind structured test outcomes into the evidence and compare them against a manifest frozen from the suite. SLICE-009 binds junitxml into the digest-bound command argv, then signs structured outcomes as evidence v3.

    The signed summary includes outcome counts, sorted non-passed IDs with their kinds, and a full-outcome digest. The artifact is read from the hermetic sample before teardown. An older evidence v2 row is refused loudly rather than treated as a skip.

    The manifest is generated by a freeze ceremony from a hermetic run’s junitxml and committed beside the gate catalog. It is outcome-blind: it records the expected test IDs, not a convenient claim that every test passed. Expected skips are declared by ID with a reason.

    Then the manifest diff becomes the rule.

    • Every manifest ID must appear as passed.
    • A declared expected-skip may skip or pass.
    • An undeclared skip blocks.
    • An xfail, xpass, error, or missing ID blocks.
    • An absent, unparseable, oversized, or duplicate-ID artifact blocks.
    • Extra IDs neither satisfy nor block the claim.

    This is not a parser trick. The slice requires a real pytest run to produce the artifact, then freezes and judges that same artifact through the same path. A handwritten XML string can exercise a parser while proving nothing about the pipeline that produced it.

    The manifest also comes from the dispatch-time base tree during delegated judging, never the candidate. A candidate that can edit the measuring stick cannot be expected to judge itself strictly.

    Go hunt for disappearance in your pipeline

    You do not need to build a new kernel to test this idea. Start by asking whether your gate can name every required test and say what happened to each one.

    • Run a required test with a skip marker and see whether your release gate still approves the result.
    • Delete a required test file in a disposable branch. Confirm the gate fails for the missing test, not only when the suite crashes.
    • Add a selection rule that deselects a required test. Check whether the gate notices the missing ID.
    • Collect a structured result artifact such as junitxml from the exact command your gate judges.
    • Freeze the expected test IDs before the change being judged. Keep declared skips explicit and reasoned.
    • Compare outcomes to that frozen list, including skips, xfails, xpasses, errors, and duplicate IDs.
    • Keep the manifest outside the candidate’s authority when a delegated worker is being judged.

    A test that cannot be found is not a test that passed. That sounds obvious until a zero exit code waves it through.

    What the proof covers and where it stops

    SLICE-009 proves that an undeclared skip blocks, declared expected skips are explicit permission rather than an obligation, non-passing outcome kinds block, and a missing ID blocks. It also proves this against a deleted test file, the same class of disappearance as the 27-test incident.

    Ranex now gates its own repository using this rule. The README says the current gate judges signed structured outcomes against a manifest diff rather than exit code alone. It PASSes only against that manifest and flips to FAIL when a frozen test file is deleted.

    The close-out record gives the implementation’s then-current measurement: a 736-ID manifest, 67 declared expected skips, and 669 passed in the sealed sample — the manifest has grown since; governance/suite_manifest.json currently carries 943 IDs and 113 expected skips. The source also states why the skipped tests exist: harness-fork, cold-start by design, dependency and provisioning conditions, one OpenRouter credential, and one mount namespace. Those skips were declared at freeze time with reasons. They were not silently accepted.

    There is a boundary here. A hostile tree can fabricate an all-pass artifact. The slice has a passing test that states this forgery boundary rather than hiding it. Structured outcomes improve what the exit code could establish; they do not make an untrusted producer truthful.

    That is why this belongs with how the kernel works: a verdict is only as strong as the evidence and authority around it. Ranex is pre-release, and its README lists a working verdict path alongside gaps that remain designed rather than built. The source records are available in the Ranex repository under docs/slices/done/.

    Questions people actually ask

    These questions help you inspect skipped and missing tests before a green status approves them.

    Why is exit code zero not enough for a test gate?

    A skipped or vanished test can leave pytest exiting zero, so exit-code satisfaction alone cannot show that required tests ran.

    How can a test gate detect skipped or missing tests?

    A gate can compare signed structured junitxml outcomes with an outcome-blind manifest frozen from the suite and block missing IDs.

    What happens when a frozen test file is deleted?

    Ranex’s own gate flips from PASS to FAIL and names the missing test ID.

    Make your next green run account for every test

    Take one required test file and delete it in a disposable branch. If your gate stays green, stop treating that gate as proof that the suite ran. Bind a structured outcome artifact to the command, freeze the required IDs, and compare the result against that list.

    Try it. Break it. Tell me what broke. Star Ranex on GitHub if this gives you a sharper test gate, and send the honest critique. The missing test your pipeline catches next is the point.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • An Environment Variable Can Choose the Repo You Judge

    An Environment Variable Can Choose the Repo You Judge

    TL;DR: A git safety flag cannot protect you if inherited environment state chooses a different repository first; build the child environment deliberately. The Ranex slice log shows why the observed artifact must be the artifact you intended.

    You can write “refuse” in a threat table and still leave the door open. That happens when a control answers one question while an unexamined input answers the more important one.

    For a tool that judges a repository, the first question is painfully basic: which repository did it actually read?

    In this note

    The flag did not choose the repository

    --no-replace-objects stops git object substitution inside the repository it is querying. It does not choose that repository.

    ADR-004 records how Ranex learned that distinction. An earlier ADR said every git query passed --no-replace-objects, so a git replace substitution would refuse. The statement was only half a boundary.

    At nine of ten call sites, Ranex’s shared git() helper passed env=None. That inherited the caller’s GIT_* variables. A relative GIT_DIR could therefore direct subject-digest reads, trust-root reads, and scratch-index status at a shadow repository the observed party controlled.

    The attack was reproduced end to end without forging anything. Genuine signed evidence produced in the shadow tree passed the governed repository’s gate even though the governed repository’s own committed check exited 1. The evidence was signed correctly. The measuring instrument was aimed at the wrong thing.

    Ambient state is an input

    Inherited environment variables are inputs, even when no function parameter mentions them. If the observed party can set them, they can influence what your tool observes.

    This was not a git-only lesson. PATH can choose the binary behind a command name. HOME can choose global configuration. A process that silently inherits either has accepted a decision from outside its explicit interface.

    That does not mean one fix closes every environment risk. ADR-004 scopes its change to the environment of git queries Ranex itself makes. It says plainly that HOME still selects ~/.gitconfig, PATH can still select a git shim, and a repository’s own .git/config remains in scope for git. Naming the limit is part of the control.

    This same shape sits behind several false-pass routes: a toolchain or import path selected by the party being measured can make a check run somewhere other than the code you thought you checked. Six roads to a false pass covers the related PYTHONPATH and sitecustomize bypasses.

    The fix builds the child environment

    Ranex now removes every GIT_* key from os.environ before launching git, then applies only explicit call-site overrides. Ambient GIT_* values no longer choose the repository.

    The helper’s parameter changed from env to overrides, because the old name suggested a complete environment when the intended meaning was a small deliberate addition. One call site needs that addition: uncommitted_paths passes the scratch index it computed as GIT_INDEX_FILE.

    The choice is broader than a denylist. A future variable beginning GIT_ is excluded as soon as it exists. Git’s own environment helper was not copied: it intentionally lets configuration variables through so -c works across submodules. ADR-004 records that a copied allowlist would have preserved that gap.

    The backing test does not settle for “the poisoned run did not pass.” It asserts that the poisoned evaluation returns the same verdict as the clean evaluation. A crash could satisfy the weaker assertion and still hide a broken boundary. Restoring inherited environment makes the reproduction red again while the clean controls stay green.

    What the boundary still misses

    The GIT_* boundary closes the reproduced redirect, not every input git honors. Repository-local config can still inject a filter; ADR-004 leaves that as a strict expected failure. HOME, PATH, and future non-GIT_* variables are also outside this specific control.

    That disclosure is not an apology for the fix. It is what keeps a narrow control from becoming a broad claim. The earlier defect was not merely an unhandled edge case; it was a sad-path row that implied it covered more than it did.

    If your own tool shells out, inspect the process boundary as closely as the command-line flags. Ask which inputs select the target, which inputs select the executable, which inputs select configuration, and which of those inputs came from the caller by accident.

    An audit you can run

    You can find this bug shape before an incident. Start with the boundary that decides what your tool reads, not only the boundary that decides what it may do.

    • Locate every subprocess wrapper. A single shared helper is useful only if every relevant query uses it.
    • List ambient selectors. Check environment, current directory, config discovery, executable lookup, and inherited file descriptors.
    • Separate target from action. A flag can constrain an action while saying nothing about the target.
    • Construct the child environment. Start from an intentional set and admit only values the program calculated or explicitly owns.
    • Poison the environment in a test. Assert the clean and poisoned runs produce the same result.
    • Write down exclusions. A boundary that cannot see an input must not claim to control it.

    Where Ranex stands

    Ranex is pre-release. The README describes a kernel with a working verdict path and a larger designed surface that is not yet built. ADR-004 records a closed control in that kernel boundary; it does not turn the project into a finished product or erase the separate risks listed above.

    Your next review should begin one layer earlier than the flag. Confirm the process is asking the right repository before you celebrate the answer it received. Try it. Break it. Tell me what broke.

    Questions people actually ask

    Can GIT_DIR make a tool judge the wrong repository?

    ADR-004 reproduced a relative GIT_DIR redirect that made Ranex read the subject digest, trust roots, and scratch-index status from a shadow repository controlled by the observed party.

    Does --no-replace-objects choose the repository git reads?

    No: ADR-004 says --no-replace-objects constrains object substitution inside a repository, while inherited environment variables had separately chosen which repository Ranex queried.

    How did Ranex stop ambient GIT_* variables?

    Ranex changed git() to build a child environment from os.environ minus every GIT_* key and then apply only explicit call-site overrides, including the computed GIT_INDEX_FILE override.

    What does the GIT_* fix still not cover?

    ADR-004 leaves repository-local configuration, HOME-selected global gitconfig, PATH-selected git binaries, and non-GIT_* variables outside this boundary.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.

  • 59 Refusals, Zero Tests: What Mutation Testing Found

    59 Refusals, Zero Tests: What Mutation Testing Found

    TL;DR: A test can pass without executing the control it names; mutation and path coverage found 59 refusal paths no test reached. The parent slice log carries the failure.

    Your safety check has a test with a reassuring name. Does that test execute the safety check? You cannot answer by reading the name, the coverage total, or a green suite. One cleanup control had never worked on any Python the project supports, while its test replaced the very function it claimed to cover.

    In this note

    That is how a green suite becomes theater. The test is present. The control is present. The binding between them is absent.

    When the general form was measured, 59 refusal paths had no test execution at all. None of those 59 are failures or weak assertions. 59 refusals that no test reached.

    The test that proved less than its name promised

    The named test never executed the cleanup function it claimed to cover. It replaced _remove_materialisation with a stub, so it proved only the surrounding precedence logic.

    SLICE-004 isolated the runner, its toolchain, and its environment so Ranex could observe a materialisation of the subject commit rather than a tree chosen by the party being measured. Its first close claimed that its controls had been mutation-checked.

    Then a cleanup path broke the claim open.

    The materialisation must be removed after a run, including after a refusal. The intended cleanup error handling failed on Python 3.11, 3.12, 3.13, and 3.14. On Linux, the removal path could call an error handler with os.open, which needs a second argument. The handler called it with one. That raised TypeError, escaped from a finally block, and replaced the original refusal.

    The control had never worked on any supported Python. This was not a new regression. The first closure had declared a fixed behavior that did not exist.

    The named test was the worse part. It monkeypatched _remove_materialisation out and replaced it with a stub that raised SubjectError. So it exercised precedence logic in materialise_subject. It never ran the cleanup function it was named for.

    That is a shape you should learn to fear: a test whose setup removes the thing its title says it proves. The test can be impeccably formatted. It can pass for years. It can still certify a path no real run takes.

    The original mutation check was also run by hand by the actor who wrote the code. The slice says exactly why that matters: it was a self-report, and it missed the defect. This is not an accusation of bad intent. It is a statement about the limits of a workflow that asks one actor to build the control, choose how to break it, and summarize the result.

    Measure the general form, not your favorite example

    One broken cleanup handler was a symptom. Measuring refusal paths across src/ranex/ found 59 raise statements and except bodies no test executed.

    That is the move. Do not stop at “did this one example fail?” Ask the larger question that could embarrass the entire pattern.

    The unreached paths included kernel input validation, five pinned-toolchain refusals, unsafe-path and duplicate-entry guards in the materialiser, and one branch of the journal chain check. Any one of them could be wrong in the same way the cleanup handler was wrong.

    After reopening, SLICE-004 added mutmut and diff-cover. diff-cover prevents a future change from adding a line no test executes. mutmut replaced the hand-run claim with recorded output.

    The whole-package mutation run produced 2,596 mutants: 1,636 killed, 73 with no tests, 7 timeouts, and 880 survivors. The larger figure is a map of debt, not a victory lap.

    Fifteen of the 59 unreached refusals were closed. Forty-four remain. Some of the 880 surviving mutants are in the kernel. They are recorded instead of being rounded into a tidy conclusion.

    A check nobody has tried to break is a check nobody knows works.

    Mutation testing does not make a test suite omniscient. The slice documents that mutmut excludes several subprocess-driven tests, which makes its signal for parts of the CLI noise rather than evidence. That limitation is a reason to say where the measurement applies, not to abandon it.

    Go hunt these shapes in your own safety net

    You can start without adopting Ranex or redesigning your test stack. Look for the distance between a control, the test named after it, and the behavior your system actually takes.

    • Choose a refusal, validation, cleanup, or recovery path that matters. Confirm a test reaches that exact function rather than a stub standing in for it.
    • Delete or invert the control in a disposable branch. Watch the covering test fail. If it stays green, the test is not bound to the control.
    • Search for tests that monkeypatch the function named in the test title. Read what remains after the patch replaces it.
    • Measure all error and refusal paths in the relevant package, not only the one you planned to fix.
    • Add a changed-lines coverage ratchet so a newly added path cannot enter unreached.
    • Record mutation survivors, timeouts, and exclusions. Do not call a tool’s blind area coverage.

    Yes, this is slower than saying “the suite is green.” A safety net is supposed to slow a fall, not decorate the floor.

    What the recorded proof says

    SLICE-004 closed with the cleanup behavior made version-independent and proven against a real mode-0 directory. It replaced a hand-run mutation claim with tool output, and diff-cover reported the newly added cleanup at 100%.

    It also says what remains unclosed. The kernel’s verdict.py had 47 surviving mutants with zero unreached mutants. One inspected survivor inverted the success comparison inside the contradiction check, and no repository test detected it. That is named as a test gap, not treated as proof the behavior is wrong.

    This distinction matters. A surviving mutant is evidence that a test did not distinguish a code change. It is not automatically evidence that the system fails in production. A green test suite is not automatically evidence that the relevant branch ran either. Keep the claims separate.

    Ranex is pre-release. Its README says it has a working verdict path and very little else, while significant production hardening remains unstarted. The accountability apparatus is not a slogan if it hides its own blind spots. The slice records are in the repository under docs/slices/done/ precisely so the unresolved counts stay visible.

    Questions people actually ask

    These questions help you check whether a named test reaches the safety control it claims to cover.

    What can mutation testing reveal that line coverage misses?

    Mutation testing found 59 raise statements and except bodies in src/ranex that no test executed at all.

    Why was SLICE-004 reopened?

    A cleanup control had never worked on supported Python, and its named test monkeypatched out the function it claimed to cover.

    What changed after reopening SLICE-004?

    The slice added mutmut and diff-cover, closed 15 of 59 unreached refusals, and recorded 44 remaining refusals plus 880 surviving mutants.

    Break the check before it blesses the code

    Pick one safety check this week. Make the guarded condition happen. Remove the control in a disposable branch. Verify its named test turns red. Then measure the rest of that control family, because one passing example cannot tell you the family is covered.

    Try it. Break it. Tell me what broke. Star Ranex on GitHub if the record helped, and send the honest critique. Especially if you find a measurement that is only pretending to measure.

    Disclosure: this post was drafted with AI assistance. Every factual claim traces to the repository’s README or slice records, the same fact gate the product enforces on code. It ships only after Anthony’s own review.