Release notes

Startup Studio release notes, newest first

Every release, newest first, written as what you get rather than what was touched. The most recent one is open below. Any term here that is not ordinary English is defined on the reference page.

Latest release

All seven pages now publish machine-readable data. Two did before. Six say what each page holds; the home page describes the product, as it always did. The how-to page publishes the five steps of its recommended flow as steps. The reference page publishes its nineteen glossary terms as nineteen definitions, each with its own address, so an assistant can answer a question about a stack card by quoting the one definition that covers it, not the whole page.

The copy matches the page today. All nineteen definitions and all five steps were checked word for word against what the page shows a reader, before release.

A check went in first, and was watched failing. Machine-readable data that stops parsing fails in silence: no error, nothing visibly broken, and the search engine ignores it until somebody notices. Our checks now refuse a page that carries no data, and refuse data that does not parse.

What they do not check is that the page and its copy stay in step: edit a definition and its copy will not follow.

Six of the seven pages had no main heading at all, so a search engine had almost nothing to go on about what each page was for. All seven have one now, written to say what the page actually covers rather than to chase a phrase. Nothing else about the pages changed: the type is the same size, the home page opens the way it always did.

Nothing here was watching for this. Our own checks now fail when a page goes out with no main heading, or with one that is only the product name. They were watched failing on all seven pages before anything was fixed.

The robots file told a story that was not true. It opened by describing a block that does not exist and then reasoned at length about working around it. That is gone. What is there now is what was actually measured, with the date it was measured.

There is also a plain-text map of the site for assistants, and the GitHub page finally links to the site.

The tool that tidies your project history can no longer lose the trail back to your decisions, and what you were shown when you decided is now checked against what was written down.

The tidy-up could lose the trail back to your decisions. It did not recognise every signpost it writes, so a document collected duplicates and older decisions could end up with nothing pointing at them. Clearing those duplicates by hand was the move that lost them. All are recognised now, and tidying by hand is safe.

Nothing checked that the options stored against your answer were the ones you saw. Now something does. If the list you clicked was ordered differently from the list on the record, you are told while it can still be fixed. If the two cannot be compared, it says so, because silence from a check reads like approval.

A tidy-up that made a document longer called it a saving. It now says the document grew.

Cheaper sessions wherever a project has history to archive, a tidy-up that cannot lose your text, and claims about this method you can rely on.

A project pays for its own past every time you ask it anything. What a session loads is sent again on every request, so old history is a bill that keeps arriving. One project now loads 196,000 characters less, roughly 50,000 tokens returned on every request there. Nothing was deleted: it moved to a linked archive.

The tidy-up could damage what it was tidying. It could lose a sentence you added, or cut the end off a document while reporting a saving. Both are fixed.

It could not read some documents at all, walking past two hundred thousand characters of one project's history. It reads them now, and stops rather than guess.

A new project loaded the wrong shared document on every request, four times the size of the right one.

The health check now reports one total across every project.

Twenty-five sentences claimed more than they delivered. Every one now says what is true.

Fewer abandoned releases, release notes written at the point they cost least, and a health check that will not quietly report on a different project.

Releasing kept getting stuck. Editing the notes that describe a release threw away the review of code those notes never touched, so the whole change had to be read again from the beginning. Two releases in a row were abandoned inside that loop. The two are now kept apart: changing what you say about a release no longer discards what was already checked about the code.

The notes are written last now, on their own, once everything else is finished, rather than first. That is the order that wastes the least, and it applies to every project rather than being advice that one of them follows.

The check that reports on a project task board could pick up a different project's board when run from somewhere unexpected, and report on it without ever saying so. It now declines to answer rather than answer about somebody else.

A project record you can put in order, a test run that leaves your machine as it found it, and release notes that stay short and say what you get.

Your project record can now be put in order. Different parts of the studio wrote times into the same record that disagreed by ten hours, so a later reader could not tell what happened first. They agree now, and every new entry says which clock it came from. Older entries are untouched.

Test runs no longer leave anything on your machine. One pass used to create 341 folders in your temporary directory and remove none of them.

A reason for waiving a rule stays with the file it was about.

A question that events overtook can be closed. Your record stops asking about decisions that no longer matter, and never shows a ruling you did not give.

Release notes are capped at 200 words. Anything longer needs the owner's approval first, and the approval is recorded next to your own changelog rather than somewhere central.

A studio session opens already knowing which faults have returned and carry no ticket. Before this, the record was written and the rule refused, but nothing put either in front of the session that plans the work, so the loop was enforced and not closed.

It says nothing when there is nothing to act on. Not a summary of what passed, not a count: nothing. This block is injected once and then re-sent with every request for the rest of the session, so a reassuring line costs exactly what a useful one costs and carries no information.

It is capped twice, on lines and on characters, and it says when it cuts. A block that quietly stops at its limit tells the reader the list ended.

It can never take a session down. The reader is wrapped on its own and logs its own outcome, so a fault here loses the summary and not the session start.

Honest about what is proved. The reading tool is covered by 51 assertions and a mutation sweep. The wiring into session start was proved by running the real hook end to end and watching the block appear and then not appear, not by an automated test. That is written on the ticket rather than counted as coverage it does not have.

The studio reads every project's doctor rows in one pass and refuses when the same class of fault has returned in a later session carrying no ticket. That is the exact failure this whole line of work was raised about: a fault found, written down, found again, written down again, and never turned into anything.

Two sittings, not two rows, and the difference matters. The same fault written twice by one session is one bad afternoon. The same fault returning in a later session is a standing defect. Counting rows would have refused the first and missed nothing useful.

It finds a board the way the board defines one. A directory holding project.json beside a tickets directory, discovered rather than assumed. The previous generation of this kind of tool looked only for a directory called .board, found nothing in a project that runs 154 live tickets, and wrote that nothing down as a fact about somebody else's project.

It is read-only and it needs no network. Each project writes its own rows at its own wind-down. This opens those files and writes nothing anywhere.

The correction it made to itself on its first real run. It reported that fourteen projects "run no board at all". One of them runs a database-backed board, which a tool that reads files cannot see. The line now says what it actually measured, that no file board was found, and says plainly that this is a limit of the instrument rather than a finding about the project. A negative result carries the search as a premise.

What a session sees at the start is capped, says how many lines it left out, and prints the command that shows the rest. The full record is read on demand and never loaded automatically, because anything loaded automatically is paid for on every single request.

A record, one row per finding, written at the end of every session and read at the start of the next. Each row carries a class, and a class is what lets anybody count how many times the same thing has gone wrong. Two of those rows exist already and both are about the session that built the record.

The number that justifies it. No tool here had ever opened a review transcript. So every finding any reviewer ever made existed only in the conversation that produced it and vanished with it. That is why the same faults keep returning: one ran four sittings in a row, another three, each time found fresh by somebody with no way of knowing it was not the first time.

What changed. tools/doctor-record.js writes, reads, gates and archives the record. It refuses a row with no class, no evidence command, or a finding too short to say what went wrong, and it refuses before it writes anything, so a rejected row leaves the file untouched. A wind-down that wrote no row now blocks the commit, with a written-reason escape that is printed on every run rather than hidden behind a flag.

One row per line. Sessions only ever append; archiving is the one routine that rewrites. Two sessions writing into this repository at the same time is ordinary here and happened twice in the last two sittings. A single record rewritten by both merges silently and wrongly. Whole lines appended by both produce an ordinary conflict a person is shown. Losing a row loudly beats keeping the wrong one quietly. The rewrite is called out rather than glossed because an earlier draft of this entry said the file was only ever appended to, which was a stronger promise than the code kept: archiving rewrites it, and a row another session appended while that ran used to be lost. It is now removed by key from a re-read of the file rather than by position from a stale snapshot, and a row that arrived in between is kept and reported.

The defect found by running it, which reading the code would not have caught. Archiving rebuilt the file from the rows it could parse, so any line it could not parse was deleted with nothing said. The line most likely to be unparseable is the half-written last row of a session that crashed, and archiving is meant to run at every wind-down. The routine that runs most often was the one destroying the evidence the record exists to keep. Reading a damaged record is now tolerant, because a reader that refuses locks out every project at once; rewriting one refuses, because a writer that refuses costs one person one repair.

One helper, tools/patch.js, that every patch script calls. A shell collapses two backslashes to one even when quoted, so a search string containing an escape arrives mangled: it either matches nothing, or writes a corrupted string. That has happened seven times across five sessions. Twice it produced a file that would not parse, and once the studio's own board was down until it was repaired. Two of those sessions responded by writing a note telling the next session to watch for it, and the note is nought for seven.

What changed. The helper refuses an empty anchor, refuses a replacement carrying a stray control character, requires the anchor to match EXACTLY the expected number of times, parses the result in memory before anything reaches disk, and writes atomically then reads back. Nothing is written unless every one of them passes, so a refused patch leaves the file byte-identical.

The part that is new rather than collected. When an anchor matches nothing, it prints the nearest text in the file and the first character that differs, by code point, and says ESCAPE COLLAPSE when that difference is a real control character sitting where the file holds a backslash and a letter. Nobody made that diagnosis in seven instances, because nobody was shown the two strings side by side.

How it was proved. Sixty-four assertions, and fourteen mutations run against them: thirteen killed, one survivor recorded in the file as unprovable rather than left as a silent gap. One mutation exposed a real hang in the occurrence counter that no end-to-end case could reach.

Two rules were reported to the founder as delivered and were not. The em-dash ban was in no shared fragment at all. The reply cap was in seventeen of seventeen role files in each of two projects and in none of the documents either project's session loads, because roles govern dispatched subagents and the session that writes to the founder reads something else. Four sittings of delivering that rule changed nothing.

What changed. tools/check-rule-delivery.js reads a manifest of every rule this studio has claimed to ship, with the destination the claim rests on, and refuses when the text is absent. It understands two kinds of destination: a named file, and the LOADED SET of a CLAUDE.md, which is that file plus its imports resolved all the way down. The second is the one nothing checked. It also refuses when wording a rule REPLACED is still being read, because while both are present the roles and the governance say opposite things. Every refusal prints the full list of files it searched, since an absence reported by an instrument that looked in one place is a claim about the instrument. The em-dash ban is now written into the shared brevity fragment for the first time.

The founder asked for it in their own words: a session should say its goal and the business value at the start, wind-down should compare against that, and the goal should be finished rather than carried. Twenty-seven sittings had happened before one of them stated its goal out loud.

What changed. tools/check-session-goal.js reads a ## Session goal section carrying the goal, the value, when it was stated and a verdict of MET, PARTLY MET or NOT MET, with a reason required for anything carried. It then reads the BOARD, where the same goal is written as a note before work starts, and refuses when the goal first appears after the first ticket of the day. A goal recorded at the end is a description of what happened, and the two are indistinguishable once written down unless something reads the timestamps.

And the warm-start skill no longer prints the whole resume prompt. Its stated reason was that the founder should see what the record claims against what is true. The corrections table printed two paragraphs above it already IS that gap, computed and labelled. So roughly two thousand words per session start, in every project, served neither reader. That is why a three hundred word cap never worked: a cap cannot beat a direct instruction to print the manual.

tools/check-print-anchors.js finds every string a tool prints from more than one place, then every assertion matching it with nothing pinning it to a site. An assertion over a whole output is an assertion about the union of everything that produced it, and a union hides exactly the single-component failures a suite exists to find. This had happened three times, in two files and two languages, and each fix was one more hand-written clause on the one instance somebody noticed.

How it reads a suite. The first version read only quoted strings, found nine results and missed every instance it was written for, because the assertions that matter are written as regular expressions with no quotation mark on the line. It now recovers the literal runs from a regular expression as well. It is a ratchet: forty-one existing findings are recorded and printed in full on every run, and a RISE refuses.

On 2026-09-18 the releases page check went into the record as status ok, exit 0, in the release set, with its own text reading that one heading matched no release and was dropped. Five entries were one command from publishing under a five-day-old headline, because the release note is taken from a DATED heading and an undated one silently selects the previous section. A person caught it. No check did, and one was watching.

What changed. A check that exits 0 while printing a warning is recorded as warned, and the gate refuses on it and quotes the warning. Separately, an undated heading CARRYING release content now refuses outright, the way a near-miss date already did; an undated heading with nothing under it still only warns, so a reader with prose dividers is not locked out.

The check had two modes and neither could answer whether a project's conduct changed. Pooling every session ever written means months of history swamp the sessions since a rule arrived, so the number never moves. The other mode asks the host which transcript is the current session, and against another project the honest answer is that it cannot tell.

What changed. --per-session scores each transcript on its own and prints a row per session, newest first, with --since to bound the window. It never claims to be the current session and says so in its own heading.

The session budget guard stops a session when it has spent more than a segment of work is worth. That segment was a single number shared by every project, measured from three real sessions, and it was right as a default. It was also the only place an answer could go, so a founder who wanted a different ceiling for one project could not have one without silently raising it everywhere.

One project asked for that ceiling in twenty-five sittings out of twenty-six and never got an answer, which is what a question with nowhere legitimate to land looks like.

A project can now name its own in .claude/session-budget.json:

``json { "firstWeighted": 8000000, "setBy": "who and when", "why": "what it is based on" } ``

Read from the working directory upwards, so it works from a subdirectory. The second stop stays half a segment further on, derived rather than configured. A value BELOW the default is honoured too, because a setting that only ever permits more is not a control.

A missing, unreadable or malformed file falls back to the shared default and is not an error. A guard that stops firing because somebody mistyped a config is exactly the failure the guard exists to prevent. Proven rather than asserted: measured at the project root and from a deep subdirectory (both the override), at another project (still the default), and with the file holding broken JSON and then a negative number (both the default). Existing tests 60 passed, 0 failed.

What you need to do. Nothing, unless you want a different ceiling. Every project without that file behaves exactly as before.

Four roles in the base roster carry rules they did not have before. A project had been holding them in its own rulebook while that rulebook said, in its own text, that they belonged here. Some had said so for eight sessions. They are stack-neutral, so by the routing test they are studio lessons, and every project picks them up on its next session.

  • code-reviewer. A test mutation that changes a VALUE cannot detect a change to whether something is emitted at all. Cover each thing a change adds on both axes. Came from five assertions and three mutants that all altered one element's value while a reviewer's untried mutant, making a second element conditional, survived every one of them.
  • security-reviewer. A probe that asks for the record back after writing it cannot tell a refused write from a refused read. The write had succeeded and the read of it was refused, under an error naming the write, and the probe reported a wide open hole as closed. Ask for exactly the thing under test, and after a fix check the refusal CHANGED SHAPE rather than merely staying red.
  • security-reviewer. A test for a control whose job is to REFUSE must assert the refusal, never the success of the happy path, or it is measuring the fixture. The original could never have seen a revoke that silently disabled the guard, because a disabled guard makes the write succeed.
  • backend-engineer. A name is not a body. An object existing in a live system does not mean the file defining it ever ran there; two routines present by name were 2 of 20 and 0 of 18 lines when compared by contents. Keep a positive control in the same run, and never report PRESENT as "applied".
  • frontend-engineer. An error is not a verdict. A path that turns "we could not ask" into an answer about the thing being asked about has deleted a state. This one told buyers a live event was closed whenever the page could not reach the database.
  • frontend-engineer. A new bound needs its own check, because the rule that looks like it covers the case is often the one that cannot. A price of 0 passed a "cheaper than standard" rule and the buyer was told to bring Free with them.

What you need to do. Nothing. Projects pick these up at their next session start.

The release command now finds out whether it is allowed to publish before it changes anything, instead of halfway through. Nothing about what it publishes has changed.

  • The order was wrong, and only a refusal could reveal it. Releasing does two things: it commits and pushes the private repository, then it builds the public export and pushes that. The scan that blocks a publish for naming something private ran inside the second half. So a release that tripped the scan had already pushed the first half. You would be left with a private repository one release ahead of the public one, no publish, and a status command that has never been able to report that gap. The command exists to keep those two in step, and this was the one path where it did the opposite.
  • What changed. The export is now staged and scanned first, before a single commit. If the scan finds something, the release stops with both repositories still in step and tells you what to fix. It costs one extra staging pass and no network.
  • The preview still cannot see this, and that is said here rather than left to be discovered. Asking for a preview prints that it would publish and returns before it reaches the scan, so a preview can still tell you a release is fine when the real run will refuse. Making the preview run the scan was the first attempt and it was reverted: the preview would then have to stage the whole export to disk, and a check exists to guarantee a preview stages nothing. A preview that writes in order to find out whether writing is safe has stopped being a preview. Closing it properly means scanning somewhere other than the publish staging folder, which is a change of its own. Until then, the preview covers what would be committed and not what would be published.
  • Found by trying to ship. Four entries had been waiting since 13 September. The release was attempted, the check that reads the changelog refused, and pulling that thread found eleven places across five published files carrying a client project's name in passing, in comments and test fixtures where it served no purpose. All eleven are reworded. Two of those projects are named and linked deliberately on the showcase page, so this is about not scattering their names through unrelated internals rather than about hiding them.
  • The eleventh was found by a reviewer and not by the scanner, and that is the useful part. Ten matched the blocked pattern. The eleventh was a directory name with two letters transposed, which no pattern anchored on the correct spelling can ever catch. The scanner reported clean and was right about the question it was asked. This is the same shape as an older defect in this file where a name with a suffix slipped past a word boundary: a scanner proves the strings you thought of are absent, and never that a name is.

The document this project reads at the start of every session is a third smaller, and it will stay that way without anybody remembering to make it so. Nothing was deleted.

  • The old rule made the limit into a target. History was only ever moved out of the loaded document when a budget check refused, so the file was cut back to just under the ceiling, grew to it again by the next sitting, and never once went below. Six sittings in a row describe that move as the fallback rather than the plan. A saving you take only when you are forced to is a saving you never compound.
  • The two sections that dominate the file had no tool. A record of each sitting is written in two places every time work stops, and neither is read again after the sitting that follows it. The decisions table already had an archiver that runs unattended. These two were moved by hand, at the end of a session, by whoever had least budget left to do it carefully.
  • What changed. A new tool moves all but the most recent sitting out of both sections into the archive files that already sit beside the document, and the session start now runs it every time rather than only when something refuses. Measured on this project: the loaded set went from 135,142 characters to 96,791, which is about 9,600 tokens off every single request for the life of every session. Every one of the 1,178 distinct lines it started with was proved present at its destination and read back from disk before a byte was removed.
  • It refuses rather than guessing, and says what to do about it. The dated history in that document is followed by live state that looks exactly like more history. Moving that by mistake would delete facts a session needs today, so the tool works out where history stops from the pointer an earlier archive left behind, and if there is no pointer it writes nothing and tells you how to name the boundary yourself.
  • A no-op is reported as success. Its older sibling reports "nothing to archive" with the same failure code it uses for a real fault, which teaches every reader that the code means nothing. This one exits clean, because a healthy document is not a failure.

The rule that a decision reaches you as a prompt you click is now enforced in the one other project on this machine running a board, as well as here. It was meant to reach every project already. It could not, and the reason had nothing to do with where the guard was installed.

  • The guard was already running everywhere. It is registered once for this machine, so it fires in every project on it, and no project overrides it. There was nothing to install.
  • It was looking for the board in one place. A project keeps its tickets in a folder, and the board itself has four rules for finding that folder. The guard had one. That project keeps 154 live tickets where the guard's single rule could not reach, so the guard found nothing, decided it could not tell, and allowed every answer through. The one other project running a board was the one project the guard was blind to. It now finds the folder the same way the board does, working from the command it is holding, so the two cannot disagree about where a board is.
  • A quotation mark switched it off completely. Every folder path on this machine contains a space, so the command that records an answer is quoted. The guard recognised unquoted paths only, so in those cases it never fired at all and nothing appeared anywhere to say so. A guard that a quotation mark turns off looks exactly like a guard with nothing to report.
  • What it still cannot do. Four projects keep no tickets at all, so there is no record of a decision for any guard to read. Installing something there changes nothing until those projects track their work somewhere it can be checked.

When a session is stopped for going over budget, it is now told how many characters it has sent and how much of that was writing files. Until now the guard could say what a session had cost and never what the cost was made of, although every tool call hands it the information.

  • The obvious fix was measured first and is not worth building. Refusing calls above some size looks like the answer. Across 62 sessions of this project, the twenty largest calls out of eleven thousand carry four per cent of the total. The cost is spread thin rather than piled up, so a size limit would refuse correct work to save almost nothing. That finding is written into the stop message itself, so the next session does not spend a day building it.
  • What the spend is actually made of. Half of everything sent is shell commands, and nearly half of those are long inline scripts. Writing files is another quarter. Classified by what the characters are for rather than by which tool carried them, the largest single thing is this project writing its own records: state documents, changelogs and ticket notes together are 42 per cent. That is the method describing itself, not work on anything you use.
  • And the biggest calls pay for the same thing twice. Seventeen of the twenty largest are a script being written whose only job is to then write a document, so both the script and the document are paid for. Editing the document directly is the cheaper path and is now named in the stop message.

When an agent is about to record your answer to a decision it never showed you as a clickable prompt, the tool call is now refused before it writes anything. Until this, the rule was measured only at the end of a session, which is after every decision had already been put. The measure could report a breach and could never prevent one.

  • Where it sits is the whole design. The refusal lands between the moment a question is written to the board and the moment an answer is recorded against it. That is the only point where the remedy can actually be carried out: raise the prompt, run the command again, and it passes. A refusal at the release gate would have been unrecoverable, because a question already written down with no prompt behind it cannot be un-written. That is the same defect that had the older brevity check blocking four releases in five, and it was removed for that reason.
  • It is not in the board program, and that is deliberate. The board is published for anyone to install. Putting a refusal there that has to read a session transcript would either break for readers who run a different agent, or quietly do nothing for them, and the published board would have grown a dependency on one specific coding agent. The refusal lives in this project's own hook instead.
  • It stays out of the way, and that was tested as hard as the refusal. A hook that blocks the wrong call is worse than one that never fires. Four separate checks prove it leaves alone: a board command that is not an answer, the question-writing command itself, an unrelated shell command that merely contains the word, and a ticket with no open question. It also allows the call on any error at all, including a transcript it cannot read.
  • It caught its own author first. The first version silently allowed the very case it was written to block, because it read the session identity from the environment instead of from the payload that names it. A hook does not reliably inherit that variable. Without a test that watched the refusal actually fire, it would have shipped as a gate that never fires.
  • Two comments in the gate definition were wrong and are corrected. Both said the brevity check is one a release runs. It has not been for several releases. Neither of the two measures you named as your top priorities gates a release today. That is a decision taken on evidence rather than an oversight, and it is now written where a reader will see it instead of the old claim.

Three things, in the order they will matter to you. When an agent needs a decision from you, it has to raise your tool's interactive multiple-choice prompt so you answer by clicking, rather than typing a list into a reply and hoping you scroll. There is now an instrument that measures whether that actually happened, and it is the first one this project has ever had for it. And the release gate no longer names a reviewer role that nothing anywhere proved was real.

  • The rule you were given told every agent to do the wrong thing. The shared rule about asking for a decision said "numbered options, so the reply can be a single character". That describes a list typed into a reply. It is now the interactive prompt, with the prose list demoted to a fallback for the one case that genuinely has no prompt available, which is a dispatched subagent handing its options back to the session that dispatched it. The rule is one file composed into seventeen roles and into the shared governance, so it changed in one place and arrives everywhere.
  • Nothing had ever measured it, and that is the part worth knowing. A search for the prompt by name across every document under the venture root found it in no governance file and no instrument. The companion rule about brevity has had a measuring tool for weeks. So one of the two rules about how you are spoken to was enforced and the other was decoration, and no report would have told you which.
  • tools/check-decision-shape.js is the new instrument. It reads the session's own record and the project's board and reports how many decisions were put to you against how many prompts were actually raised, marking each decision as PROMPT or PROSE. It refuses on the board record, which is unambiguous, and only reports on whether a reply merely looks like a list of options, because numbered steps followed by a question is how anybody writes ordinary instructions and refusing on that shape would fail correct work. It runs in the wind-down set. Twenty assertions, and a project with no board reads as cannot-tell rather than as a project in breach.
  • The role that reports on your sessions now leads with those two. Brevity and the clickable decision are sections one and two of the doctor's review, ahead of the instruments and the process, because those are the only two surfaces you actually see.

The release gate demanded a role name and nothing proved the name existed. The gate required a reviewer called doctor and matched it by exact name against a hand written constant. Deleting that role from the repository moved no number anywhere: the session start set, the gate's own suite and the roster count were all identical either way, because the roster check counts FILES and the gate matches NAMES. Two surfaces were being compared to each other and neither to the roster, so they could agree perfectly about a role that did not exist.

  • It was live rather than theoretical. The roster installed on the machine that publishes this project held the role under its old name and no doctor.md at all, so the gate was printing "start one of: doctor" to an install that could not start one.
  • And a sync could not have fixed it, which is why the rename had quietly been blocked. The global install wrote every file the base defines and removed nothing the base had dropped, so a sync would have installed the new name and left the old one in place for ever, still dispatchable and still carrying claims that had just been corrected. The install now prunes what it placed and no longer recognises, and it keeps anything hand edited unless you pass -Force.
  • The split between refusing and reporting is deliberate. The test suite REFUSES when the gate's constants disagree with the roles in this repository, because that is ours and a disagreement is a defect. The tool only REPORTS on the roster installed in your own home directory, because you may be running the method with no roster or with one you wrote yourself, and there is no way to tell a broken install from a different setup by looking at it.

A sentence in the shared governance had been false for months and was about to be copied into every project. AGENTS.md said that each reviewer returns PASS or FAIL and that only then is a deploy permitted. No reviewer's verdict has ever blocked anything here, and that includes the work reviewers: nothing in this repository opens a review transcript, so no PASS and no FAIL has ever been read by the gate. What the gate does refuse on is whether a review HAPPENED, which is a different question and the only one it can answer from the record it keeps. A forced governance sync was about to push the old sentence into five projects, so it was corrected first. It now says plainly that no verdict stops a deploy, and that the teeth are the ticket a finding gets written to rather than the gate.

The method reviewer is called the doctor everywhere. The role that checked the installation across projects and the role that checked whether a session followed the method were doing one job from opposite ends, and neither was read by anything. They are one role now. The rename touched the roster, the shared governance, the gate tool and the published site, and every edit was applied by a script that refused unless its pattern matched exactly once, because a plain search for the old name returned more than fourteen hundred hits across a hundred and seventy five files and almost all of them were the word directory.

A review of this release found three things in it that should not have gone out, and the worst one would have deleted most of your roster. None of them reached you, because the review ran before the publish rather than after it, which is the whole reason that step exists.

  • -Sync -Only <role> deleted every role it was told to skip. The prune added earlier in this same release read the filtered list of roles rather than the full one, so every role you asked it to leave alone looked exactly like a role that no longer exists. Running it for one role reported sixteen others as dropped from the source and would have removed all sixteen from your machine, giving a reason that was untrue of every one of them. -Update went through the same code. It now refuses to prune at all on a filtered run and says why, and it refuses again if the source directory is empty, which would otherwise have turned a sync into an uninstall.
  • The note describing that fix repeated the false claim it was describing. The entry above about reviewers and deploys said the blocking rule was true of the work reviewers. It is true of neither. Nothing in this tool has ever read a reviewer's verdict, so no PASS and no FAIL has ever stopped anything; what the gate checks is whether a review happened at all. The sentence has been rewritten from the code rather than from memory of it, which is the third time in two releases that correcting a false claim produced a new one.
  • The new check for clickable decisions accused a session of something it had not done. When a decision had no prompt inside its own window it said the decision had reached you without one, printed directly under a line counting the prompts that were raised. The two causes are different faults: no prompt at all, or a prompt raised before the question was written down. It now names which, and for the second it says what to do next time rather than printing an instruction that cannot be carried out now.

The set that runs when you open a session dropped from eleven checks to five. The set that runs before a release dropped from fourteen to nine and from about thirty six minutes to about four. Nothing that protects you was removed. What was removed had, between all of it, refused once in seventy nine recorded runs, and that one refusal was this project tripping over its own paperwork.

  • The measurement that prompted this. Every check writes a row to a ledger, and that ledger is committed, so there are seventy nine recorded versions of it to read. Counting refusals across all of them: health-report, hook-wiring, decision-keys, board-doctor, resume-pointer and session-brief have never refused once between them. comment-shape refused a single time, and that time it was this repository's own new comments raising its own baseline. Each of those was standing in the path between writing a change and being able to test it.
  • They are not deleted, they are in a set called deep that nothing gates on. Run them with node tools/run-checks.js --set deep whenever you want the full picture. A check that has earned its keep as a diagnostic has not automatically earned a place in front of your work.
  • mutation-coverage left the release path for the same reason, and it is the big one. It takes about thirty minutes. Its findings are real and they are never urgent: the two open when this change was made were fragments of a diagnostic message that no user of this tool will ever see. A slow check belongs on a schedule you choose, not between you and a release.
  • health-report alone was eighty two per cent of the time your session start took, 3,351ms of 4,065ms, for a row that always exits zero and that no gate can read.
  • The method review now reports and never blocks. When a release has been reviewed for its work but not for its method, that is recorded as advisory rather than refused. It used to exit 1 and stop the release. It now exits 4, which the runner treats as a notice. A release with no code review of any kind still stops at exit 1, and so does one where a commit landed after the review read the tree.
  • The first version of that change quietly removed a second guard, and a reviewer caught it. The two places that report a missing method review sit above the check asking whether the tree moved after the review started. While both returned a refusal the order could not matter, because either way the release stopped. Making one of them advisory made the order load bearing: a session with a code review and a commit landing after it reached a green gate with nothing having looked at the newer commit. The missing method review is now carried to the end and answered last, so the tree question is always asked. Four assertions cover it and go red when the early return is put back.
  • Why that split, rather than trusting the reviewer less. Nothing in this tool has ever read a reviewer's verdict. No instrument opens a review transcript, so a review's conclusions have never gated anything mechanically. What gated a release was whether a review had been STARTED. Making the method half advisory removes a refusal that was never measuring what it appeared to measure, and the findings still reach a ticket, which is where they were always meant to live.
  • reply-shape-recent left the release set. It refused a release when a single banned character appeared in the session's own replies, with no way to clear it, and by this project's own record it blocked four of the last five releases. Nothing a reader of this tool sees was ever at stake. The absolute count is still recorded at every wind-down, so no slip is hidden.
  • board-doctor left the session-start set for a different reason. It exits 1 in a tree that has never run init, so a fresh clone of the published export reported it red before the reader had touched anything. It still runs before a release, where a board is expected to exist.
  • What is kept, and why. changelog-leak, because it is the one check here that has ever stood between a private name and a public repository. suite, because it is the tool you install. board-audit, roster-count, published-counts, governance-core and releases-page, because each polices a number or a page that a reader actually sees.
  • Every assertion that encoded the old behaviour was restated rather than deleted. Eleven assertions across two test files named the old exit code or the old set membership. Each one still makes its claim, against what the tool does now. Both suites are green: 78 and 137.

You can keep your changelog in your own house style. The releases page check now tells you it has nothing to compare rather than refusing, unless your configuration declares that this repository is the one publishing the page.

  • The rule and the check disagreed, and the check was the one you met first. This project's own non-negotiable rule is that nothing ships without a changelog entry written first. Write one without the **What this gives you.** marker this repository happens to use, and the generator could build no page at all, which read as a failure at exit 1. That check runs in the set every install runs at session start, so the refusal arrived before any work did, every time, and the remedy printed alongside it told you to rewrite your own file.
  • Whether the page is yours is now declared, not guessed. Set $PublishesReleasesPage = $true in your configuration and the check holds your page to your changelog as before. Leave it unset, which is every install that has not deliberately opted in, and a changelog the generator cannot build is reported as nothing to compare. An earlier version of this change inferred the answer from whether the tool was configured at all, which is true of every install, so the refusal came straight back for anyone who set the tool up. A reviewer measured that before it shipped.
  • The guard for this already existed and could never run. A test asking whose tree this is sat at the bottom of the page comparison. Getting there means the changelog parsed into a page first, and the ordinary case does not parse, so the guard only ever protected readers who did not need it. It now also covers the failure to build, and it is one named function rather than two copies of the same condition, so both refusals move together when the evidence changes.
  • Only the comparison is affected. Run the generator yourself and a changelog that cannot build still fails at exit 1 with the real reason, because then you asked for a page and the error is the answer. And in the repository that does publish the page, a page that disagrees with its changelog is still a refusal, which is the entire reason the check exists.
  • Measured in a reader's layout rather than in ours. In a tree holding the tools, a changelog and no publisher configuration, the ordinary entry moved from exit 1 to exit 3. Built out to the full export from the publish manifest, the session start row moved from FAILED to ADVISORY. Three assertions cover it and three separate mutations each break at least one of them.

The two checks that read session transcripts now ask the host which session is running, instead of taking whichever transcript was written most recently.

  • Two sessions in one repository is normal, and it broke the measurement. Both checks ranked every transcript for the project by modification time and took the newest. Where a second session is open on the same repository, the newest file is whichever session typed last, so a check could report another session's conduct as yours. On the machine this was found on, the project directory held 54 transcripts and a second session had written into the repository in four consecutive sittings.
  • There is deliberately no fallback. Where the session cannot be identified, both checks now say so and exit advisory rather than guessing. Falling back to the most recent file is the fault itself, and a guess that reads like an answer is worse than no answer: it is the number somebody then quotes.
  • Nothing here gets stricter. Both checks already treated this exit as advisory, so a tree that cannot identify its session gets a row saying why and is not blocked.
  • Two of the first four mutations proved nothing and were redone. One left the refusal in place so the fallback was never reached; the other made the match looser in a way no test fixture could see, because no session in them was named after another. A fixture was added for the second and the first was restated as the edit that actually performs the fallback. All four now break assertions.

The role that actually picks work now carries the queue rule the governance document publishes, rather than the one it retired.

  • The contradiction shipped inside a single release. base/governance/GOVERNANCE_CORE.md says a CEO instruction jumps the queue only if it belongs to the initiative in flight or is a defect we introduced, and says in so many words that this replaces the older rule that a direct CEO instruction is always top of the queue. base/agents/tech-lead.md still carried the older rule, word for word, in the same range.
  • Why the role file is the one that matters. A role reads its own definition as its instructions and does not go looking through governance for contradictions. That file composes into five projects and publishes in the export, so the studio was distributing a rule it had just published a decision to retire, to the one agent whose job is deciding what to work on next.
  • Measured rather than asserted. Searching base/ for the retired wording returned two files before the change and one after, and the survivor is the sentence that does the retiring, which has to keep the words or the retirement stops being findable.

Two roster changes, both found by a real defect on a live project the same night, and both of a kind that will be silently wrong on any project rather than only the one that found them.

  • The mobile gate now measures element rectangles instead of a document width. An app shell that hides its own horizontal overflow at a mobile breakpoint also hides it from the instrument, so the gate reads exactly the viewport width no matter what the content does. The roster already told the gate to lift both overflow axes before trusting a width, which works but only if you correctly guessed every element that needed lifting; on the project that found this there were 144 of them. Reading each element's bounding rectangle against the window width needs no lifting, no restoring and no guessing, and cannot be defeated by a stylesheet nobody has written yet. Lifting stays, as the way to show how much is hidden and to prove the instrument can see, rather than as the verdict.
  • The frontend role now knows which of the two wrapping rules actually works. A row containing a long unbroken string, such as an email address used where a name was expected, runs off the side of a phone even when the container carries the flex settings everybody reaches for. overflow-wrap: break-word looks like the fix and is not: it wraps the text visually and still lets the unbroken string set the minimum width, so the row overflows by exactly as much as it did with no rule at all, verified to the pixel. word-break: break-word and overflow-wrap: anywhere are the two spellings that work. The trap is somebody later tidying the one that works into the one that does not, so the rule now asks for a comment saying which is load-bearing.

The tool that builds the releases page reads your changelog by date heading. A heading that is nearly a date, but not exactly one, matched nothing, and everything under it was quietly discarded while the build carried on and the staleness check reported the page current.

  • This was not hypothetical. Someone wrote a second entry for the same day and spelled the heading in a way that got past the duplicate-date rule. Their whole entry was in the file, was going to ship, and would have been announced on no page and in no release note. Nothing anywhere could see it, because a parser that stops looking reports exactly what a parser with nothing to report does.
  • A heading that is nearly a date now refuses, naming the heading and the date to write instead. That case is never deliberate.
  • Any other unrecognised heading is counted and warned about, not refused, because your changelog may carry headings this tool knows nothing about and refusing on those would lock you out of your own file.

The front door measure reports work that reached a commit without its ticket going through the board, and it refused. There was no exemption, no waiver, no recorded escape of any kind, and a release is gated on it, so one breach anywhere in the recent commits blocked every release until those commits scrolled out of range, including for a session that had done nothing wrong.

  • Every other check here already had a way back, and this one did not. The comment checker takes a recorded rise with a reason. The published numbers checker takes a dated exemption. The coverage checker takes a baseline entry with a reason a stranger can read, and refuses when that reason no longer applies. This measure said no and stopped, which turns a rule into a wall, and a wall does not produce compliance, it produces a queue of work that cannot ship.
  • A waiver names a commit, so nothing can be excused in advance. A commit that does not exist yet cannot be waived. The act has to have happened and somebody has to have written down why it stands, with their name on it.
  • A waived breach is still counted and still printed. The line stating the answer counts it as a breach and says how many were waived, so the summary cannot contradict the detail below it, and each one carries its full reason and the name of whoever granted it. Worth being exact about the limit, because the first version of this note was not: what a reader sees is honest, and the exit code on its own does not distinguish a waived run from a clean one, since it was already advisory for a different reason. Making that machine readable is written down as its own work rather than claimed here.
  • Every waiver must resolve to a commit that already exists. This is the part that stops it becoming a way to ignore the rule, and the first version did not have it. A reviewer wrote one waiver that turned the whole measure off, another that excused nine unrelated commits at once, and a third that sat dormant and then silently excused a commit made 79 commits later. An entry is now checked as a real hexadecimal hash and resolved against the repository, so it cannot be written ahead of the thing it excuses.
  • A waiver that no longer matches a breach refuses. Point one at a commit that is not a breach and the check stops and says so, rather than leaving a standing exemption nobody can account for.

The budget guard runs before every tool call in every project, and it kept a small file of its own in your system temp directory to remember what it had already told you. Two faults in that arrangement, both fixed, and the first one is the serious one.

  • Losing that file turned the guard into a wall. The call count, its place in the transcript, and the record of what it had already announced all lived in that one file. Lose it, to a temp cleaner, a lock, or a directory it cannot write to, and all three reset: the record of what it had announced went back to nothing while it re-read your whole session and got the full total. On any session that had already passed the first threshold, which is most long ones, that meant a stop on every call for the rest of the session. A stop on every call also stops you winding down, so the very thing the guard exists to protect, the state written at the end of a session, was what it would have cost you. Measured against the version before it: 60 of 60 calls blocked, where the older one blocked none. It now stays quiet on any call where it could not read that file, and also on any call where it could not write it, because a stop it cannot record is a stop it would make again on every call after that one. The honest bound: where the file is lost once, this costs one delayed warning. Where the directory can never be read or written, a full disk being the realistic way in, it costs every warning, which is silence. That is the trade, taken deliberately, because a guard that says nothing is recoverable and a guard that blocks you winding down is not.
  • Being stopped for a long session hid the moment you went over budget. The guard has two thresholds, one on what the session has cost and one on how many calls it has made, and they shared a single mark for what had already been announced. So a stop on the call count also marked the spend threshold as dealt with, and the session then crossed it in silence. The two now keep their own marks, and the message names whichever one actually moved rather than whichever number happens to be larger, which is what it was doing when it told this session it had stopped because spend could not be read while spend was readable and over.

Every refusal below was correct about there being a problem and wrong about what to do next, which is worse than staying quiet, because a remedy that cannot be performed sends you looking for a fault in your own repository.

  • The board could refuse every command and print three ways out, none of which worked. If an eviction is interrupted it leaves a journal file behind, and every command that writes refuses while that file exists, which is right: the board is half moved and you need to see it before you build on it. If the file itself is damaged, the escape was a loop. Commands that write sent you to the rollback command. Rollback refused, because it cannot know what to put back, and sent you to a command that writes. And the health command, which is where people actually look when a board is behaving oddly, told you to restore the file from version control. It is excluded from version control on purpose, so that was impossible by construction. All three now name the same real escape, deleting the file, and they name it FIRST, because every other step they suggest is refused until it is gone. There is now a test that performs that escape and checks the board is free afterwards, rather than checking the sentence was printed.
  • The leak check could report your changelog clean against half of its own rule list. It reads the list of private names from your configuration, and it found the end of that list by looking for the first line starting with a closing bracket. A rule carrying its own list of exemptions ends on a line of exactly that shape, so every rule below it was never read. The guard meant to catch this counted the rules inside the text it had already cut short, so it agreed with itself. Watched failing on a three rule config: the old version says clean against two rules with the third name sitting in the file. It now reads the list by matching brackets, ignoring anything inside a quoted value of either kind and anything in a comment, because the rules are patterns, patterns are full of brackets, and a stray bracket in a comment ended the list just as effectively. The first attempt at this fix skipped only single quotes, and a reviewer broke it again in two lines: a smiley in a comment, and a bracket inside a double-quoted value. Worth knowing which way it fails now, because the two are not the same: a bracket that never closes makes it say it cannot read your configuration, which is safe, and that is the only direction left.
  • Asking to check a release without naming which one reported success and checked nothing. Typing the gate option with nothing after it silently fell back to running a different set of checks, overwriting the very record it had been asked to read back, and exiting zero. Naming an option and getting the default is now a usage error, for that option and for its partner.

The check that reads your session and refuses a release over an em-dash now asks a narrower question at release time, and the same absolute question as before at the wind-down.

  • The problem was that a slip could not be undone. The check reads the session transcript, and a reply that has been sent cannot be unsent. So one slip in the first minute condemned every reply after it, however clean, for the rest of that session. Four of the last five releases here died that way, including one session that went 45 consecutive clean replies after its slip and still could not publish.
  • The obvious fix was measured and thrown away. Moving the release to the front of a session, where the transcript is short, sounds right. Across 50 stored sessions, 43 of 50 would still have breached inside their first five replies. That is a change worth two sessions in fifty, so it was not built. The number that mattered pointed the other way: sessions recover once they notice.
  • So a breach ages out at release time and never in the record. The release asks whether the last twenty replies are clean. The wind-down still counts every slip in the session, because that number is read into the compliance table and a slip you recovered from still happened.
  • They are two separate rows rather than one row with two settings, because the gate keys its record on the row name. One row would have let a windowed pass overwrite the absolute count, which is the record disappearing through the change meant to protect it.
  • Twenty is a judgement and is written down as one. Long enough that you cannot slip and publish a minute later, short enough to reach inside one sitting.
  • A pass after a slip says so. It reports what the session did earlier rather than printing the same line a spotless session gets, because those two are not the same and the one line most people read should not pretend otherwise.

The check that reports whether work reached a commit without its ticket being started had been refusing since a week earlier, and both reasons were faults in the check.

  • Raising a ticket counted as work. The check already sets aside a commit that touches only the board, because raising, parking and annotating tickets is administration. But raising a ticket also writes the handover document the next session reads, and that one extra file flipped the same act back into work. The set-aside now covers the board and the project's own record, and the record is derived from what your entry document actually imports rather than from a list of names.
  • It is not a hiding place. A commit is only set aside when its entire footprint is the board and the record, so bundling a source change into a handover edit still counts.
  • A commit was judged once per ticket it named. A commit that closed one ticket and parked another was refused for the parked one, so naming both tickets failed where naming one would have passed, for identical work. The check was rewarding a thinner commit message. It now asks whether any ticket the commit names went through the door, and names all of them when none did.
  • Found by bisecting rather than guessing, and the ticket's own first hypothesis about the cause was wrong. Worth saying because the session that broke it published its tree as green with this check inside the suite it was quoting.

Everything in the three entries below was reviewed before it shipped, and the review failed it. This entry is what changed as a result, because a changelog that only records what worked is a sales document.

  • The check that stops a review certifying the wrong code could not be cleared. It refuses when a commit lands after your reviewers started, and it prints one remedy: start them again. The window opened at the first reviewer of the session, and a session log only ever grows, so starting them again added a later entry and could not move it. Anyone who reviewed, fixed something and reviewed again was refused for the rest of that session with nothing they could do. The window now opens at the earlier of your two most recent reviewers, one of each kind, so the remedy works. Earlier rather than later is the whole point: taking the later of the two would let re-running one reviewer clear a window the other has not looked at, so both have to be started again before the window moves. There is now a test that performs the remedy and checks it clears, and another checking that re-running only one kind does NOT clear it, which is what was missing: the old test checked that the sentence was printed.
  • The same check passed when it had not looked at all. A timestamp it could not read went straight to the tool that lists commits, which accepts anything it cannot parse and quietly reads it as the current moment, producing an empty result and a clean pass. It now says it cannot tell.
  • The leak check reported clean on text the publish refuses, two different ways. It dropped rules written in grammars it did not expect, without a word, so the scan reported clean because it had stopped looking; it now counts the rules it found against the rules it read and refuses to report on a partial scan. And it was matching case exactly where the publish ignores case, which was wrong for nine of twenty-one rules including keys, tokens and file paths, always in the direction that lets something through.
  • Twenty-five new checks were being run by nothing. The list of test files is kept by hand and the new one was not added to it. That is this project's own oldest finding, in the sitting that cites it.

One report was itself wrong and is recorded as wrong: one of the three grammars said to be dropped is read correctly, and the test now asserts what the tool does rather than what the report said.

  • The front door now names its own model. An assessment is six disciplines taking positions and writing a verdict. Building is code, tools and tests. There was never a reason those two wanted the same model, and until now there was no obvious way to say so.
  • The setting sits on the dispatch, not on the roles. Those six leads get sent for plenty of work that is not an assessment, in every project that runs the roster. Naming the model in the role file would have changed all of it. Naming it where the assessment starts scopes the choice to the room it was decided for, and it is one place to undo.
  • The thing that had to be established first was whether this could be scoped at all, and it turned out to be already wired. Every role carries a model field, set to inherit, and it survives composition into the installed roster. So this was never going to require anybody remembering to flip a setting at each boundary, which is the version of the idea that would have been skipped within a fortnight.
  • A renamed model does not take the front door with it. If the name is refused, the leads are dispatched without it and the verdict says so. A gate that dies because a vendor renamed something is worse than a gate running on the wrong model.

Not claimed: nobody has yet assessed the same idea on both models and compared them. This puts the lever where it belongs and sets it. Whether it is the better choice is unmeasured.

  • The check now runs when the entry is written, not when you publish. The leak scan that keeps private project names out of a public repository lives in the full test suite and in the publish step. A changelog entry is written at the wind-down, by the session most likely to be quoting another project by name, because that is the session summarising what it just learnt elsewhere. So the gap between writing a private name and hearing about it was a whole sitting, and it has now closed on a real name three times.
  • It reads one file, and says so out loud. The authoritative scan reads the entire publish manifest and has not moved. Widening a slow, broadly scoped check into a set that strangers run is how a lockout ships, which this project did last week and spent a sitting undoing. This is the cheap early warning for the one file whose writing schedule guarantees the problem.
  • It does not print what it finds. The matched text is the private name, so printing it to prove the tool worked would copy that name into a terminal, a transcript and a log. You get the line number and the kind of rule that matched, which is enough to open the file.
  • On a copy of this repository it says it cannot tell, and never blocks. The rules are a list of exactly the names to look for, so the list is itself the leak: it lives in the private half of the publisher, which is not published. A reader therefore has no rules, and a scan with nothing to look for passes everything. It reports that it cannot tell instead, and the same answer covers a missing changelog and a rule that will not compile.

Measured: 25 assertions, 0 failed. Two defects in it were found by running it against the real rule list rather than by reading the code: one PowerShell spelling of case sensitivity would not compile at all, and the string parser stopped at the first half of a doubled quote, silently reading 20 rules where the file holds 21. Both now have their own fixture. A third was in a test rather than the tool: the usage fixture supplied a valid argument in front of the empty one it meant to test.

  • The banned character had three definitions and one of them was wrong. The check that reads every reply counts two code points, U+2014 and U+2015 horizontal bar, because the two are indistinguishable at every size a reader sees. The check that reads the founder brief borrowed that tool's prose limit and its preamble test, then tested the character by hand against one of the two. So the rule the changelog published was enforced on replies and half enforced on the one piece of text nobody can edit afterwards.
  • What that cost, concretely. A brief carrying a horizontal bar passed the wind-down, was printed verbatim by the session-start hook as the first thing the session said, and then refused the release from inside the first reply of the session, where there is no override and no way to take it back. That is the exact failure the brief checker was built to prevent, arriving one character over.
  • It is now the same definition, not a matching one. The hand-written test is gone and the borrowed predicate answers all three questions it can refuse on. The refusal also says why it cannot be fixed later, because a reader who does not know that will try.
  • The test that certified the rule could not fail on half of it. The fixture used one code point while the assertion was named for the rule, so no change to the other branch could ever redden it. Both characters now have their own fixture.

Measured: 107 assertions, 0 failed, up from 104. Three mutations against the same tree: putting back the code that shipped 2, deleting the refusal 5, dropping the sentence that explains the timing 1. One of the two new assertions was itself falsified on the first run, returning 1 where it should have returned 2, because it matched wording the passing branch also prints. It was restated rather than removed.

  • A review is now evidence about a particular tree, and it was not before. The gate that asks whether anybody reviewed the change proved only that review agents were STARTED. It said nothing about when, so a reviewer could read one version of the work and the release could publish another, with every check green and no surface anywhere reporting it.
  • This is not hypothetical and it was not caught by an instrument. Two review agents were dispatched at 07:52 against one commit. A SECOND coding session, working on a different project on the same machine, committed into this repository at 07:59:23 while both were still reading. Four files moved. Two were the source and the tests of an instrument the release itself depends on, and one was the changelog, which the publish reads to build the releases page. The release was one command away. It was found by a person reading the git log by hand.
  • The rule now is that the reviewers must have been dispatched after the last commit. If any commit lands after the first reviewer starts, the gate refuses and prints the commit: its hash, its time and its subject. The remedy is one line and the refusal states it, which is dispatch them again.
  • It does not ask whose commit it was, and that is deliberate. Committing your own work after dispatching your reviewers leaves the release exactly as unproved as somebody else doing it, so the refusal says in as many words that it is not an accusation. A reader who takes it for one goes looking for a person who does not exist.
  • Two designs were tried and abandoned, and the reasons are worth more than the fix. Forgiving the session's own commits by reading the session identifier off each commit cannot be built: the identifier written into a commit and the one a session can read about itself are different kinds of identifier, with no mapping between them on the machine. Counting how often a foreign identifier appears cannot be built either. One identifier was found spanning 80 commits across four days, so it does not mark a sitting, and 76 of 377 commits carry none at all.
  • Not being able to tell is a third answer, not a pass. Outside a git repository, or where a dispatch carries no timestamp, it says it cannot tell and stands down as advisory. A reader running the method without git is never locked out, and is never told the tree held still when nobody looked.

Measured: 116 assertions, 0 failed. Six mutations, each the exact edit named, against the same copy: removing the refusal 5, making the window always empty 5, treating an absent repository as a pass 2, treating an absent timestamp as a pass 2, dropping the line that states the tree did not move 1, and no longer naming the offending commit 1.

  • The em-dash ban now has an instrument. The ban is permanent, retroactive, and named in the governance core as applying to every string a reader sees. The check that reads every reply a session wrote to the founder measured prose blocks and preamble, and did not count the banned character at all. The rule was being enforced by whoever happened to be looking.
  • The measurement that forced it. A method reviewer caught one em-dash by hand and reported one. The instrument, the moment it could count, found FOUR across two replies in the same session, three of them inside the very reply that was reporting a content gate's em-dash findings. A rule enforced by hand is enforced at whatever rate the hand is having a good day, and here the hand belonged to the reviewer whose whole job was catching it.
  • Quoting a tool still works. Fenced blocks are excluded, because pasting output that happens to contain an em-dash is showing your working, which these rules ask for everywhere else. Refusing a session for that teaches people to stop pasting the numbers, which costs more than the dashes it saves.
  • No documented way around it. U+2015 horizontal bar is counted alongside U+2014, because the two are indistinguishable at every size a reader sees. The en-dash is not counted, so page ranges and score lines are safe.
  • Eight new controls, and the count is asserted. Each edge is pinned by its own fixture, including both sides of the fence rule, and the suite refuses if the total moves. 66 assertions to 74, 0 failed.
  • The fence exclusion is now reported rather than silently dropped. The same method gate that approved excluding fenced blocks also named it as a real hole: quoted tool output is evidence, but a fenced paragraph still reaches the founder carrying the banned character. In-fence em-dashes are counted and printed as a second number marked REPORTED, NOT REFUSED ON, which is the pattern this tool already uses for prose share. It becomes a refusal when the person watching the number decides it should. 74 assertions to 76.

Found by another project's round-four method gate.

  • The width check was structurally blind to text running out of its own box. Overflowing text does not enlarge the box, so the document scroll width and every bounding rectangle read exactly the same whether a title fits or overflows by nine thousand pixels. Measured on the tree that found it: document scrollWidth 375 in both cases, element bounding right 353 in both, and only the element's own scrollWidth against its clientWidth moved, 331 against 331 staying silent versus 9534 against 331 firing. A page could be unreadable past thirty characters and pass.
  • The gate now asserts scrollWidth <= clientWidth + 1 on every visible text-bearing element, and reports the node, the overflow in pixels and the offending string.
  • It also feeds absurd fixtures on purpose. A 1000-character title, a long URL, a string with no spaces at all. Realistic test data is precisely what hides this class of defect.
  • A second blindness, in the shell rather than in the text. An app shell setting overflow-x: hidden at a mobile breakpoint clamps the reading to the viewport width whatever the content does. Lifting that axis alone changes nothing, because overflow-x: visible beside overflow-y: auto computes straight back to auto. Both axes must be lifted together, and every reading now states whether it was taken lifted or unlifted, because the two are not comparable.
  • And the sticky bar that pins on desktop but never on mobile. The obvious explanation, that the element has no slack to travel through, is usually wrong and is worth disproving before it is written down: if the same rule pins at desktop width it is not about slack. The common real cause is html, body taking overflow-y: auto at the breakpoint, which makes body the scroll container and therefore makes the sticky view rectangle the whole document. Where it was found, lifting overflow and nothing else moved the control 866 pixels with the slack untouched.

Found by another project's round-four mobile gate, which proposed the fix by running it rather than by suggesting it. All three are the same root cause wearing different symptoms, so a project that has one very likely has the others.

  • You can keep your own changelog without our tooling turning red. The releases-page check holds a generated page to the changelog it came from. It was moved into the set that runs at every session start, and the export ships our changelog, our page and the tool together, so the moment you followed our own rule and wrote your first changelog entry the check called it drift and refused. Every session, with one documented remedy: regenerate the page, which writes our marketing domain into your repository.
  • The check now asks whose page it is holding. If the tree does not publish that page it says so and stands down as advisory. Your changelog is yours and the page arrived with the export.
  • It has not been blunted where it does apply. In a tree that does publish the page, a page that disagrees with its changelog is still a refusal. That direction is asserted on its own, so a scope that quietly stopped the check working would fail the suite rather than pass quietly.
  • The check that finds unproved code had itself not run for six commits. It lives in the release set only, and no release had happened, so a guard added last week sat unproved and nothing said so. That guard is now proved, along with the line that tells you which session file was read when a review is reported missing. Lines proved by mutation in that tool went from 95 to 97, and the four remaining exempt lines are explanatory prose in one message, each carrying a written reason rather than an exemption nobody has to justify.
  • Preamble is now caught behind the bullet character we actually write. The rule that refuses an opening line of throat-clearing stripped the asterisk, underscore and hash markers before reading the line, and not the hyphen. So the same sentence was refused as a bullet written one way and waved through written the other, and the immune form was the common one. Both the reply rule and the founder brief rule read that one predicate, so both were holed and both are fixed.

Why. Two faults with one shape: a rule written where its author lives, then shipped to people whose folders look different. The page check had never once been wrong in the repository that wrote it, because that repository always has a page it generated itself. The preamble check had a fixture for every marker except the one this project's own writing rule produces, so the suite certified a branch no real document could take. Both were found by review rather than by any instrument, and neither would have been visible from inside this tree.

What was deliberately not done. No second copy of the shape rule was made. The founder brief borrows the reply checker's predicate, so widening the strip set fixed both at once, and the measurement was taken before widening rather than after: across 49 transcripts and 3,822 replies, one opened with a hyphen bullet and none would be newly refused. The scope on the page check asks one question about the tree rather than trying to tell two identical files apart, because a reader's drift and ours are the same bytes in the same relation and nothing inside either file can separate them.

  • The brief you read at the start of every session is now written in point form. Same eleven lines, 171 words of paragraph down to 151 words of bullets, and none of them in a prose block. Two clauses were dropped in the rewrite rather than reshaped, which is worth saying plainly in an entry whose whole argument is that the rule is met by writing better and not by cutting: the rule does not require it, and the next brief should carry its full content.
  • A brief may not open with throat-clearing either. The same instrument refuses a reply whose first line is preamble, and the first version of this only read the paragraph rule, which left the identical fault reachable through the other field.
  • Publishing is unblocked. The check that reads what a session actually sent you sits in the release set and has no override, so while the two disagreed nothing could go out at all.
  • The rule is enforced where it can still be fixed. The brief is now held to its shape at the wind-down, against the document, before it is ever handed over. The other check reads a transcript, so it can only tell you after the message has been sent, and a sent message cannot be unsent.

Why. One instrument capped the brief in LINES. The other capped an unbroken run of prose in WORDS. Both were right inside their own unit and neither could see the other, so an eleven-line brief was simultaneously within its cap and a hundred and seventy-one words past a limit of eighty. The session-start hook orders that brief printed word for word, so the breach arrived in the first reply of every session, was nobody's writing, and could not be edited away afterwards.

What was deliberately not done. A second number was not chosen here. Picking one would have reproduced the fault one level down: a third unit, agreeing with the other two only for as long as somebody remembered to keep it agreeing. The predicate and the limit are imported from the instrument that does the refusing, so there is one definition and the two cannot drift.

The rule is a shape and not a length. The same content that fails as a paragraph passes as bullets, and there is a test that asserts exactly that, because otherwise this is a word cap wearing a different name and it would be met by hiding detail rather than by writing better.

The check comparing the published releases page against the changelog it is generated from now runs at session start as well as at release. If the page is behind, you are told in the first three seconds of the next session rather than the next time somebody publishes. A tree with no page, or a changelog with nothing dated in it yet, now reports advisory instead of red: that is somebody running the method without publishing a website, which is an absence rather than a disagreement, and every neighbouring check already draws the same distinction.

Why. It ran in the release set only. Every wind-down writes a release note into the changelog and does not rebuild the page, so the page broke at the exact commit that broke it and the only thing watching did not run again until a publish. That is not theoretical: the page was found drifted at HEAD, confirmed by running the suite against a clean archive of HEAD rather than assumed, and the same failure took an entire test run down with it. The check compares two files already on disk, wants no network and no git, and costs about fifty milliseconds.

  • One large item in progress instead of two, and the small work in flight has to belong to it. Start a small that belongs to nothing while an initiative is running and the board refuses, names the initiative, and prints the command that attaches it. The old limit was two large and three small, which was the right shape for work that has nothing to do with itself.
  • A ticket can be put UNDER another. add --under <ref> when you raise it, or under <ref> --under <ref> for anything already on the board. Nothing needs migrating: a ticket with no initiative is simply a ticket whose field is absent, and every existing ticket stays as it is.
  • An initiative cannot be marked done while its work is still open. It names the tickets that are, rather than counting them. Parking or killing it is still allowed, because deciding the rest is not being done is a real ending and refusing every ending leaves work nobody can abandon.
  • Dropping an initiative is one command. evict moves it and everything under it out together, with one reason recorded on every ticket. It either happens or it does not: a journal is written before the first ticket moves, and if a write fails part way the tickets that moved are put back. If the board is ever left part way through, every command that WRITES refuses and says how to undo it, while every command that READS still works, because a half-finished board is exactly the thing you need to be able to look at.
  • wip and the rendered board show the relation, so you can see which initiative each piece of work serves without opening it, and an initiative whose work is all closed says so.

Why. The limit on work in progress was a COUNT, and a count cannot tell a small that FINISHES an initiative from one that STARTS a fourth. Measured before any of this was built: 71 of 74 starts finished without being sent back, the median time in progress was under two hours, and the limit had been reached once in the board's life. So starting was never the problem. The board still reached ninety waiting items, and 73 of those named the work they belonged to in prose that nothing could read. Tightening the number would have changed none of that; making the relationship real is what lets the board refuse the thing that actually goes wrong.

A rule about how work is picked has changed with it. An instruction that arrives mid-session now jumps the queue only if it belongs to the initiative already running, or is a defect we introduced. Anything else becomes a ticket and you are told that is where it went, with a count of what is already running. It is not refused and it is not silently absorbed, because both of those end with work nobody chose. Reprioritising is still yours to do at any moment, and it now means evicting the initiative rather than starting a second one beside it. The older rule put every mid-session idea straight to the top, which is how the waiting list grew from arrivals rather than from starts: 2.45 things arriving for every one closed, over fourteen active days.

Two things worth knowing about the checks. Moving that limit turned nineteen published claims red at once across nine files, seven of which are handed to every project rather than being pages on the site, and all nineteen were rewritten in the same change. And the release gate that asks whether anyone reviewed the work had to be split first: the same role now both reviews the method and reports breaches, so being started is no longer enough on its own, and the dispatch has to say which of the two it is. A forgotten marker refuses a release rather than clearing one, which is the direction that fails safely.

  • A number published on a page is now held to the thing it counts. The reference page said it carried eighteen defined terms and carried nineteen. It also published the limit on how much work may be in progress as a sentence, with nothing anywhere comparing that sentence to the constant the board actually refuses on. Both are checked on every session start and before every release, in both directions: a page stating the wrong number fails, and so does a page that has quietly lost the claim altogether.
  • The eighteen is corrected to nineteen. That was live on the site.
  • The limit turned out to be published in nine places, not two. Two are on the website. The other seven are files handed to every project: two role definitions, the board specification, the shared session rules, two skills and the testbed notes. Changing that limit changes the instruction given to agents in five projects at once, and until now nothing would have said so.
  • The session guard stops a session on what it has SPENT rather than on how many times it has done something. It used to interrupt at the fortieth tool call regardless of cost. Measured across three sessions, the cost per call ranged from 38,200 to 69,600 tokens, so a session costing little was stopped exactly as hard as one costing nearly twice as much. The threshold is now the spend itself, derived from the same measurement the guard was originally built on, with a backstop on call count for the case where the spend cannot be read at all. The interruption now says which of the two fired, so a long cheap session is not mistaken for an expensive one.

Why it was wrong before. A number written once in prose, on a page nobody edits again, goes false on the day somebody changes the thing it counts, and the first person to notice is a reader. That had already happened twice on this site. The guard had a subtler version of the same fault: its own explanation argued entirely about cost, it computed the cost on every call, and then it decided when to stop by counting calls and printed the cost as decoration.

One thing found while doing it. The releases page had drifted from the changelog it is generated from: the previous release note was written and the page never rebuilt, so the newest entry was missing from the published page. It is rebuilt here.

  • The brief that greets you when a session opens now arrives as ordinary text. No warning colour, no prefix stamped on every line, and no disclaimer at the top explaining that it is not an error. On a session where nothing is wrong, the start-of-session step now prints nothing red whatsoever, where it used to print thirteen warning-coloured lines.
  • The previous attempt at this fixed the wrong thing, and the founder said so. They had complained about the colour. The change made in response altered the words instead, adding two lines to the top of the brief saying it was not an error. That spends two lines of a twelve-line brief apologising for a colour, at every session start, forever. Their words: "i was only complaning why it appears in red almost as an error".
  • It was built on a claim that turned out to be false, and the counterexample was in the same file. The reasoning recorded at the time said the warning-coloured channel is the only one a start-of-session step has, so the text was the only thing that could be changed. Half of that is true and does not matter: the colour genuinely is not configurable, checked rather than assumed. The other half was wrong. There is a second channel, it is neutral, and another part of the same file had been using it since it was written. The reasoning walked past our own code.
  • Findings stay red, because red is right for them. A finding is a warning about the project you are opening and belongs in warning colour. Only the standing brief moved. The two are now guarded by a test that fails if either ends up on the other's channel, which is a different property from each simply being present.
  • The cost of the trade, stated because it is real. The warning channel always prints. The neutral one hands the brief to the session, which then prints it, so delivery now depends on the session doing its job. That was put to the founder with the trade named and they took it, on the argument that a brief being read as a failure is a brief that is already not landing. A test checks that the instruction to print it verbatim is still attached, because shipping the brief without it would stop the founder seeing it at all, silently, which is worse than red.
  • One defect was created by this change and caught by running it rather than by reading it. The rebuild step contributes an entry even when it has nothing to say, so the warning channel was emitting an empty string rather than being absent. An empty warning-coloured line is still a warning-coloured line, which is the entire complaint. It could not have shown up before, because the brief used to live on that channel and was never empty.
  • The reading budget did not quietly get bigger. The cap on how much arrives at a session start was ruled on how much a founder reads, and they read both channels, so the cap still spans both rather than being reset to the smaller half. Dropping the brief out of that sum would have turned a real limit into a check that cannot fail.
  • The two tests that guarded the old behaviour were restated, not deleted. Both required the brief to lead with the words that are now gone. What has to be true changed from what the brief says to which channel it arrives on, so that is what they check now. Proved by putting the brief back on the warning channel in a full copy of the repository: seven tests go red, and the same run with the change in place has all seven green.

  • The studio hands you a short brief when a session opens, and it was being read as a stack of errors. It is not one, and it never was: the thing that produces it finishes cleanly, writes nothing to the error channel, and hands over a valid message every time. That was measured by running it, twice, on two different days. The problem is that the only channel a start-of-session message has is painted in warning colour by the terminal, with a prefix stamped on every line, and none of that is ours to change. So thirteen lines of ordinary text arrive looking like thirteen lines of failure.
  • The founder reported it as an error two sittings running, and the second report is the reason this is a change rather than an answer. The first time it was explained. Explaining it again would have meant explaining it every sitting until the brief stopped being read, which is exactly the fate this studio has already recorded for another line it prints: nine projects learned to ignore that one. A question asked twice is a defect in the thing being asked about.
  • The fix is one line and it leads rather than merely appears. The brief now opens by saying what it is not, before anything else. A correction printed underneath the thing it corrects is not a correction, so two separate tests guard it: one that the label begins with those words, and one that no line of the brief is read before the label. The second exists because the first was published as if it checked both, and it does not: move the brief above the label and the first test stays green, which is exactly the arrangement the fix exists to prevent.
  • It cost nothing you were already paying for. No line was spent to buy the label: one line of the message was replaced by one line, and the caps that keep this brief short and narrow are untouched.
  • The budget check was blind to anything a nested document pulled in. Every project loads a set of documents before a session starts, and every character in them is re-sent on every single request for the life of that session. There is a check that measures this and refuses when it is too much. It read the top document's own imports and stopped there. So a project whose top document loads another one that loads five more was charged for all of it and measured for none of it, and the check reported a clean bill on a project that was already eleven per cent past the limit it exists to enforce. It now follows imports all the way down.
  • It was found by another project's session reading our tool, not by us. That is worth saying plainly, because it is the second time this week that the useful finding came from somebody else's numbers rather than from our own review.
  • The same blindness was in the health report, and that is the command you are told to run. The number -Doctor prints is worked out by a completely separate piece of code that had the identical defect. Fixing only the check would have left the two disagreeing, with nothing anywhere that would notice, so both were fixed together and there is now a test that runs both and refuses if they give different answers. That test earned itself immediately: the repair to the health report went in with a mangled pattern that silently dropped every document whose path contained the letter s, and comparing the two answers side by side is the only thing that caught it.
  • Two rules decide the number, and both were measured rather than assumed. An import is resolved next to the file that declares it, not at the top of the project, which was confirmed against a real project where every second-level document is absent under the other reading. And a document reached by two different paths is charged once, not twice, which on that same project is the difference between a true number and one inflated by a hundred and twenty thousand characters. A loop of documents importing each other now reaches a verdict instead of running until it runs out of memory.
  • The health report had never been tested at all. The suite drives the tool as a program and had never reached the function that works this out, which is why a defect in it survived in a published script. It has two tests now.
  • The archiver refused to work for anyone who numbers their decisions with a dash. It reads the decision numbers to work out which end of the table is newest, and the pattern it used allowed letters straight against digits and nothing in between. So a table numbered S147 was read and a table numbered D-001 was not, which meant the tool worked perfectly in the project it was written in and refused in the one that needed it, whose decisions table had grown past the size where archiving is supposed to happen. Two separate sessions there recorded the refusal and moved on.
  • Worse, it told them the wrong reason. The single message it printed said it could not tell which end of the table was newest, and that condition had never been tested. Their table was unbroken from the first decision to the hundred and twenty-first, in order, with no gaps. Two sessions read that sentence and believed their own records were ambiguous. There are now three separate refusals, each naming what actually happened, and when it succeeds it says which signal told it the order. It also reads every row rather than only the first and last, so one unreadable line no longer refuses a table the rest of the rows settle beyond doubt.
  • A new check finds documents that look organised to a person and are unreadable to every tool. Everything here locates things by shape: a heading to find a section, a table header to find the decisions. A document can be perfectly clear to a reader and have none of that, and when that happens the tools do not fail loudly, they find nothing and say nothing, and the project carries on believing it is covered. One document was fifty-seven thousand characters with not a single heading. Three separate instruments had been silently doing nothing there for months. The check refuses four things, each one something a tool actually needs rather than a matter of taste, and it looks across every project rather than only the one it lives in.
  • That new check was scoped before it shipped, because the first version would have refused on your files. It reads the folder holding the studio, which here is a folder of studio projects and on your machine is wherever you happened to clone it. Measured before release: put any unrelated project beside it and the check reported that project's notes as broken, in the set that runs at the start of every session, on a document we did not write and you have no reason to change. It now only looks at projects that actually load the studio's own governance documents, and a fresh install with none beside it says there is nothing here to check and passes, rather than showing you a red line you cannot clear.
  • A rule about getting text safely through the shell now reaches every project, instead of one. Long commands typed straight into a shell break on an unclosed quote, and one project here lost sixty two tool calls to it over four weeks: fifty four unterminated single quotes, six double, two backticks. Backticks are the worst of the three because they do not announce themselves, the shell runs whatever sits between them and drops the output into your text, which silently deleted three references from a paragraph here and reported success. The fix was already written down and it was written down in one project's private notes, where no other project could ever read it. It is now a shared rule carried by all seventeen roles and by the governance document every project loads on every request, from a single file, so the two cannot drift apart.
  • It immediately found an invisible character in the file every session reads first. Three projects were loading a main instruction file that began with a mark you cannot see in any diff or review tool. That exact character stopped thirteen of sixteen team-role files loading a month ago; it was fixed for those and nobody thought to re-check the instruction documents. All of them are clean now.

  • A check that is supposed to notice deleted files stopped noticing whole deleted folders. The tool that keeps the written descriptions of the team honest refuses when a document it has a record for is no longer there, because a record with nothing behind it is a claim nobody can check. It had to be taught that a published copy legitimately does not carry every folder this source tree does. The version released this morning learned that lesson too well: it forgave any missing folder anywhere, so deleting an entire directory of shared rules made it report success and, in the mode both automatic runs use, print nothing at all. Deleting a single file out of that same directory still failed correctly, which is what made it hard to see. It now decides by which layout it is looking at rather than by what happens to be missing: in the source tree nothing is forgiven and an absent record is a deletion, and only an installed copy can have had a folder taken away from it.
  • The comment ratchet passed here and failed for everybody who installed it, and now it does not. The instrument that stops comments quietly bloating held its record against paths from this source tree, and publishing flattens those paths. So on every installed copy eleven published files looked like files the record had never seen, were held to the strict limit that exists for genuinely new files, and refused, in both the automatic run at session start and the one that gates a release. There was nothing the reader could do about it, because the obvious remedy rewrites a file they received rather than wrote. One record now serves both layouts, and the proof is that the two now report the same numbers: forty-eight files, four hundred and ninety-two control lines, the same ratio, in the source tree and in a rebuilt copy of the published export.
  • Both were found by pointing the reviewer at the repair rather than at the thing repaired. The first is a fault introduced by this morning's own fix, an hour old, that a suite of five hundred and eighty-three assertions could not see. Both now have assertions, each proved by putting the fault back and watching exactly the named assertions fail while the controls beside them stay green.
  • The page that sells the method said five checks and the tool that detects them counted a different five. Both listed five, and they were not the same five: the page drew the studio director as a gate and the tool did not count it, and the tool counted the QA tester which the page did not draw. So a session that started exactly the reviewers the published page advertises could be told that nobody had reviewed anything, and the release would refuse. Both surfaces now name the same six, and a check compares them in both directions on every test run, along with the number written in prose beside them, so neither can drift from the other again and the count cannot go stale a third time.
  • A reviewer is now two kinds, and a release needs one of each. Five of them read the WORK: the QA tester, the code, security and content reviewers, and the mobile check. The studio director reads the METHOD instead, meaning whether the process was actually followed and whether the replies you were sent are the shape this studio publishes. It never reads the change, so it can never stand in for the five, and holding them in one list would have meant a session that started only the director cleared a gate whose entire question is whether anybody read the work.
  • Being long-winded with you can now stop a release. The instrument that measures reply shape existed and only ran at the end of a session, where it could report and never refuse. It runs in the release set now, so a session that buried the answer in a wall of prose cannot ship until it is fixed. It measures shape rather than length on purpose: a long list of bullets passes and one dense paragraph does not, because a length cap becomes a target met by hiding detail.
  • A check that every installed copy failed and nobody could clear now passes. The roster check compared its record of exceptions against paths from the private source tree, and the published copy lays those files out differently and leaves some of them out entirely. So every reader who installed it got a permanent red mark with nothing they could do about it, while it passed here on every run. One record now serves both layouts, and an exception for a file the published copy does not carry is named rather than treated as a broken record.
  • An unfinished code block no longer hides the rest of a reply. The reply measurement treated a code fence as opening a block and never reconciled it at the end of the text, so a reply whose last fence was never closed scored nothing from that point on, and the densest paragraph in it was invisible.
  • The release command can now tell a branch with no commits from a detached one. It asked git a question that answers the same way for both, so on a branch with nothing committed it reported the wrong reason. Neither state is reachable on the ordinary path, and it now says which.
  • The coverage tool can be pointed at any tool, and at itself. It compared whatever you named against one particular tool's tests, which were hardcoded. So asking it about anything else gave a confident answer to a question you had not asked. It now works out which tests belong to the tool you named.
  • Smaller repairs. Loading the hook check from another program used to run it and then kill that program. The session guard left one small file behind in the temporary directory for every session ever run, and wrote Windows line endings on any machine. The coverage tool silently dropped any line of code that shared a line with the end of a comment, so its count of code lines was quietly short. And the published note about where the reply measurement gets its evidence claimed the whole derivation was shared with another tool when only part of it is.

  • Releasing can no longer leave your private copy behind the public one. The release command publishes to the public repository and, before that, saves and uploads your own private copy. The upload was written inside the branch that only runs when there is something new to save, so a session that had already saved its work as it went was told there was nothing to do, and the upload was skipped. The public repository then received work that the private one had never sent anywhere, and the command reported success. Measured on a real release: four saved changes sat on this machine only, at the moment strangers could read the same work publicly. The upload now runs whichever way the work got there, and it is checked by asking the server what it holds rather than by trusting that the upload said it worked. If the private copy cannot be uploaded, nothing is published at all, because the public copy is the one a stranger clones and it must never be the only copy that exists.
  • The shared rules now arrive as a short document instead of a long one, and every project actually loads them. The one file every project carried was about 12,200 tokens re-sent on every single request for the life of a session. Two projects held that file and imported nothing at all, so no rule written in it ever reached them. The rules are now in GOVERNANCE_CORE.md, roughly a quarter of the size, and the long document stays beside it as the reasoning behind each rule.
  • The reasons did not go missing, they moved. Each section of the short document names which part of the long one explains it. That matters because moving a rule out of what a session loads also stops anyone finding out why it exists, and a rule nobody can explain is the first one somebody deletes.
  • Nothing can be quietly dropped in the move, and that is checked rather than promised. A new check reads both documents and refuses if any section of the long one is neither pointed at by a rule nor declared as background only. It refuses in the other direction too, when a rule points at a section that has been renamed or removed, which is the half a hand-written list never has.
  • It also refuses when a project has the file and does not read it. Being delivered a document and loading it are different things, and until now nothing compared the two. That is the defect this whole change came from, and it had been sitting there for weeks.
  • A duplicate was found inside the shared rules themselves. One section appeared twice, word for word, in the file every project loads on every request. Editing one copy would have left the other saying something else. It has been removed and the check now refuses on repeats.
  • Three sub-projects are reported rather than fixed, on purpose. They inherit the file from their parent folder and would have to reach upward to import it, and nothing here has ever confirmed that works. The obvious fix is not applied while it is still a guess; it is written down as work with the verification attached.
  • A release can be made at all again, and this is why nothing shipped for five sittings. The record of which checks passed carries a fingerprint of the code they measured, so a green result from an hour ago cannot stand in for the code in front of you. That fingerprint included the commit. Releasing commits and pushes the work before it publishes, so the act of committing invalidated the record the publish then read, and every release refused with every check reported as never proved. The fingerprint is now taken from the content of the tree rather than the commit the content happens to be sitting on. Committing changes which commit the bytes sit on. It does not change the bytes.
  • A check that no reader could ever have cleared is fixed before anyone received it. One of the checks reads the shared governance documents, and those documents deliberately do not publish. On any copy installed from the public repository the check was therefore red for good, in the set that runs at session start and in the set that guards a release. It now reports that there is nothing here to look at, while a governance folder that exists with its documents missing is still an error, because that is a defect rather than an absence.
  • Four checks could return a verdict about a file they never opened. A flag typed with nothing after it was swallowing the next flag as its value. In one check that produced a clean bill of health, silently, about settings it had not read. In another it produced the exit code that means a duplicate decision key was found, with a stack trace where the finding should be. Each now refuses and names the flag it wanted a value for.
  • A release is refused when nobody reviewed the work. Every review here is carried out by a separate agent that a session has to choose to start, and nothing recorded whether one ever did. The honest baseline, measured before the check was written: of the 21 releases this project has made, 4 shipped with no review agent started at any point in the session that shipped them. Two of those four predate the vendor setting that was blamed for it, so this was an old hole rather than a new one. From this release forward, a release is refused when the session shipping it started no reviewer.
  • That check says plainly what it does not prove. It proves a reviewer was started. It does not prove the reviewer read the change, and it does not prove it came back clean. A session can start a reviewer, be told the work is broken, and ship anyway with this green. Nothing on this machine can see that, and a check claiming otherwise would be worse than none.
  • A test suite's own claim about what it proves is now worked out by the machine instead of written by hand. A paragraph at the top of one suite listed which parts of the tool were genuinely covered and which were not. Three separate reviews falsified that paragraph, and each correction was wrong again. It is now derived: every line of the tool is removed in turn, the suite is run, and any line nothing depends on is reported. An exemption has to be written down with a reason, and it is refused once the line becomes covered, which is the half a hand-written list never has.
  • A session can no longer end or run out of room with its own state written down and uncommitted. That is how a session's record of what it measured used to disappear.
  • The standing rules are re-delivered after a session runs out of room and reloads. For two weeks that repair reported success nine times and delivered nothing.
  • Every hook is held to the event that can actually deliver it. A hook registered against an event that never fires looks identical to one that works, in both of the two forms this tool writes.
  • Two sessions working on one project can no longer allocate the same decision number in silence. It happened: two sessions read the same table, each took the next free number, and one project ended up with two decisions numbered 91 and two numbered 92. Neither writer could see the other. Decisions are referenced by number for the life of a project and are never renumbered, so a duplicate is permanent.
  • A check that ignored dated archives by name now ignores them by shape. The name list could not name a file that did not exist yet, so the act of archiving a document turned a check red at the next session start on work nobody had changed.
  • The rule about writing short now has an instrument, and it fails on our own history. It measures the shape of a reply rather than its length, deliberately: a limit becomes a target, and a target gets met by hiding detail rather than by writing better. Two hundred lines of bullet points pass; one dense paragraph does not. Run over 3,020 past replies in this project, 120 were past the limit and 48 opened with throat-clearing.
  • The picture of the method was missing one of its own roles, and a count beside it was wrong. The page that explains how this works lists the team, then draws the flow they work through. One role was in the list and not in the drawing, so a single page disagreed with itself about who is involved. Reading the rest of it turned up worse: a sentence four lines under that list said four checks have to pass, where five roles are listed. Both are corrected.
  • The release preview now says what it will do with your own copy, and warns you when the release will stop. The command has a preview that shows what a release would do before it does it. It described the save and the publish and said nothing about uploading your private copy, which is the step that can now refuse the whole release. Measured on a real release: the preview said there was nothing to save and that it would publish, while the private copy sat seven changes behind. The preview now names the upload, or says it is not needed, and it says plainly when the release would refuse: no server configured, a server that does not answer, a detached checkout, or a server holding work your copy does not have. In those cases it no longer goes on to say it would publish. It reads the server to find out and never writes to it.

  • A release now refuses while anything says the checks were skipped, failed, or were run against different code. The checks themselves are not new. What was missing was any record that they had run at all, so a release could be made on code nobody had measured and nothing anywhere would say so.
  • The record is written by the tools, not by anyone reporting on them. Each check writes down the result it actually returned. Nothing in the chain summarises a run, because a summary is a description of evidence and the difference only shows up on the day the summary is wrong.
  • A pass expires when the code changes. Every result is stamped with a fingerprint of the code it measured. A green result from an hour ago is not evidence about what is in front of you now, and the release says so and stops.
  • A check your copy does not have is reported as missing, not as passing. The public copy of this project does not carry the full test suite, so on your machine that line reads absent and the summary says so out loud rather than reporting a clean bill of health over a partial install.
  • One check is reported as unproved, honestly. The health report always exits successfully and prints prose, so its result carries no information about what it found. It is named as unproved every time rather than counted as a pass it never earned.
  • A seventeenth role, the studio director. It runs the checks and reads the record back. It checks and records; it does not fix, build, or direct the work, and it is given no editing tools.
  • Every published claim about how many roles there are is now held to the roles on disk. That number was written by hand in page titles, share cards, headings and the tool's own output, and nothing compared any of them to the directory the roles live in. Adding this role broke twenty-eight of those claims at once and the new check named every one.
  • The note about running this on a different coding agent now says what is true. It said the roles are just markdown files and the only thing that needs changing is where they land. The roles are markdown. The tooling is not, and one role does its job by running it.
  • The comments in the code you install explain the code, and no longer name our internal work. Forty-five comment lines across the published files named a ticket, a role or a review round. That is provenance you cannot look up and do not need. There are now none, across thirty files, and a release refuses to go out if one comes back.
  • Closing a session now archives the decision table instead of reminding somebody to. The rule to do it was already written and already specific. The step that had never once happened was a person choosing to run it, so the tool runs it.
  • Every role now carries the instruction to read that archive, so a decision moved out of the document loaded on every request is still a decision anybody can find.
  • The limit on how much work can be in flight can now be overridden, and it cannot be overridden quietly. It refuses once, then asks for a reason, refuses a blank reason as hard as no reason, and writes what it was told into a file that is committed alongside the tickets. Three overrides inside fourteen days and it stops accepting them.
  • The health check reads the session hook log, which nothing had ever read, and reports when each hook last fired, where, and how it ended.
  • Every test suite states how many assertions it expects. Before this, a suite that lost half its checks reported a smaller number and stayed green.

  • The board's required contract now describes the board you actually get. It specified seven columns; the program has always drawn eight, one for each status. It required a field the board you install has never had. It listed ten commands the program does not have and left out all nineteen it does. Every one of those is corrected, and the reference page and the how-to page were redrawn to match.
  • The page that explains the board no longer says two different things about the one decision that is yours. Accepting work is the single move on the board that belongs to you and to nobody else. The board drawing labelled that column as the team's while the table two inches below it labelled the same status as yours.
  • A part of the toolkit that ships to you had never been checked for the fault that once broke thirteen of sixteen roles. The list of directories the health check reads was missing one, and that directory publishes. Nothing was wrong inside it, and nothing had been looking.
  • Three rules the contract states are now checked against the program rather than against another document. The board is rendered and compared. The commands it documents are compared to the commands it has. Two directories that are supposed to hold the same list are compared by reading both.
  • Every project on this machine was carrying a rule that told it to invent its own board shape, twenty-five lines above a rule saying the shape is fixed. That file loads at the start of every session in every project.

  • The message you see when a session opens is now written for you rather than for the assistant. It used to read out the whole operating manual: ninety six lines and 977 words, every time, whether or not any of it concerned you. It is now fourteen lines and 197 words that say what is being worked on, why it matters, and what is needed from you. Both numbers are counted from what the tool actually emits.
  • It cannot quietly grow back, because two caps are now enforced when you save your work. Twenty five lines for the whole message, twelve for the part addressed to you. This matters more than it sounds: the message grew twenty two lines in a single day before this, because every session had a reason to add to it and none had a reason to cut.
  • A project that has never had one of these is told about it, never locked out. The check reports a missing brief and lets the save through. An earlier version of it refused, and it turned out that almost every project on the machine that built it would have hit a wall it could not pass on its next save. You adopt the rule by writing the section once, and until then you are only reminded.
  • The check tells the difference between broken and absent, and only one of those stops you. A brief that is too long, too wide, or identical to the session text is a genuine problem and blocks the save. A project that simply has not written one yet is not.
  • What is written for you and what is written for the assistant are now separate text on separate routes, so neither can swallow the other. The founder-facing block is read to you at session start. The assistant's instructions reach it another way and are never read out.
  • Asking for a summary of where things stand now leads with the answer. The corrections come second and the full record third, instead of the answer arriving after eighty lines of process notes.

  • Large work cannot start until somebody has written down what it is for and how you will know it worked. The board now refuses to move a large item into progress without a recorded verdict and one named measure. It does not perform the assessment and does not pretend to. It refuses to let the work begin without one, which is the only part a program can honestly enforce.
  • A measure is required, not optional, and that is the point. An assessment with no measure is an opinion with a verdict attached, and the measure is the part that gets skipped.
  • Small work is deliberately exempt, and the exemption is tested rather than trusted. Making the gate fire on small items turns the test suite red on purpose. A gate that fires on everything gets routed around, and then it protects nothing.
  • A guard that was quietly reporting that nothing was in flight now reads the right place. The session cost guard counted open work from a folder that had been moved, so for a whole session it reported an empty board while three items were open. It was confident and wrong, which is worse than silent.
  • Two rules that were written down and guarded by nothing are now genuinely guarded. One is a decision about how the board may be written to; the other is a sentence published on this site. Either could have been deleted in full with every test still passing.
  • New projects are no longer sent to the board that was replaced. Two roles in the shared team still pointed at the older tool.

  • Your project now tells you when its own record has grown too big to load, at the moment you are writing that record. The health check has reported this correctly for weeks. It only helps if somebody chooses to run it, and the person who makes a document longer is winding down, not running a health check, usually on a different day.
  • It separates two problems that look identical and have opposite remedies. A record that has honestly grown needs archiving. A section that is supposed to be rewritten every session, and is being added to instead, needs rewriting. Archiving would return nothing there, and until now nothing could tell you which one you had.
  • The cost is reported as a running charge rather than a file size. A document of 165,000 characters is roughly 42,000 tokens re-sent on every single request for the whole session, before any work happens at all. Those are the same number and only one of them reads as a problem.
  • An archive stays findable. If a project moves older decisions into an archive file, the check refuses unless a document the session actually loads points at it. Decisions nobody can find get argued again from the start.
  • There is now a command that moves older decisions out of the file a session loads, and loses none of them. It reports what it would do and changes nothing at all until you ask it to write.
  • It refuses to guess. A decisions table can run newest first or oldest first, and both are in use here. If it cannot establish which end is which from the row numbers, it stops rather than archiving the wrong twenty.
  • It writes the copy first and reads it back before touching the original. If anything fails in between, you are left with a duplicate rather than a truncated record.
  • You can now run the same board this studio runs, in a repository you already have. One file per ticket in a folder, committed beside your code. No database, no account, no service to sign up for, and nothing to keep running.
  • Its rules refuse instead of reminding. Only QA moves work into acceptance testing, and only once test notes a person can follow have been written. Nothing closes over a question nobody answered. Work in progress has a ceiling. Each of those is a refusal in the program rather than a paragraph somebody is trusted to remember.
  • You can read it without opening a terminal. A rendered board file is rewritten on every change and committed with the tickets, so it opens on a phone in any code host. Acceptance testing is the one column that waits for you, and it would be a poor joke if that were the one thing you could not look at.
  • The documentation now matches the board. Two states the board has always had, parked and killed, were described nowhere. The vocabulary of record now defines Board, which was the most used word on this site and the only one never explained.

  • The command reference tells you when to reach for each command, not only what it does. Every row now opens with the situation you are in: after you change a shared role, the first time a project needs to differ, when a file was edited where it was installed.
  • What each command writes is spelled out in plain sentences. It used to be a list of paths, which only helps a reader who already knows what those paths are.
  • The two commands nobody should ever run by hand say so. They are hooks. Your tooling runs them for you, and the table used to describe them as though you would type them.
  • The three worked examples are numbered steps. What you are trying to do, the commands in order, then one sentence saying what changed as a result.
  • Releasing says what it ships and what it does not. It ships the studio itself: the shared roles, the skills and the website. It does not deploy the product a project builds, and a reader could previously have assumed it did.
  • One limitation is now stated on the page rather than left to be discovered. The roster installs into the directories one specific coding agent reads. The method does not depend on any particular agent, but this tool does, and other agents are not supported yet.

  • The page that explains how to run this now tells you what to type. It covered installing the team and tuning it, and never mentioned the four commands you actually use in a session. They were in the repository, working, and findable only if you already knew to look.
  • The command list no longer calls itself complete while leaving out the one that ships anything. It was headed as the whole workflow and did not include releasing.
  • Every command the tool accepts is now written down, with what each one changes on disk. Seventeen switches. The site had shown five of them in a code block with a one-line comment each, so anyone evaluating this had to clone the repository and read a file to find out what it does.
  • The two commands that point in opposite directions now sit next to each other. One sends your work out. The other pulls somebody else's in and rebuilds every project on your machine from it, publishing nothing. They are one letter apart in a terminal history, and only the first of them was named on the site.
  • Three worked examples, for the three things people actually do. Improve one role for every project at once. Make a single project behave differently without forking anything. Publish your own version.
  • The two pages you read before you type are much shorter. Same ground, written the way developer documentation is written: a table you can scan and code you can copy, instead of paragraphs you have to read in order.
  • Two tables on the site were wider than the screen and are now not. The stack table and the skills table were each forced onto a single line per row, so they scrolled sideways on a desktop monitor, not only on a phone.

  • Winding down now saves the record it just wrote. It used to write the state of a session and leave it sitting on the disk unless somebody thought to ask for it, and the session that would have noticed had already ended.
  • The check that should have caught that is now a measurement. It asks for the reference of the save or the count of lines still unsaved, rather than a judgement from the session being judged.
  • The prompt that gets you back into a project is handed to you when you open it. No finding the file and copying it out. There is a command, /warm-start, for asking on demand, and it tells you which parts have gone out of date rather than handing them to you as though they were checked.
  • The two commands that point in opposite directions are written down. One sends your work out; the other pulls somebody else's in and rewrites what your projects are built from. The published guide named only the first.
  • A long session now stops itself. Cost grows with the square of how long a session runs, so the expensive sessions are the ones that feel productive. It counts, and it interrupts, and it tells you what is still open before you go.
  • You can see what a project costs to open, and what that adds up to. The health check already reported how much a project loads. It now converts that to what you pay on every single request, and multiplies it out across a session, which is the number that actually decides anything.
  • Two test suites that nothing was running now run. They were sitting in the repository being nobody's job.
  • A resume prompt can no longer send you to state it has already replaced. Closing a session adds a new block of current state and marks the previous one superseded. The line telling the next session which block to start from is written by hand, and it could be left pointing at the old one. Nothing compared the two, so the document could contradict itself and still look finished.

  • You can see what each project loads before it starts working. Two projects had quietly grown past the point where a session can open without a warning, and the only thing that reported it was the session that hit it. The health check now shows the number for every project, so a project approaching the ceiling is named weeks before it gets there.
  • What a session has to read before it starts stops growing without limit. The record itself keeps everything, as it always has. What changes is that older entries move somewhere they are kept and can be looked up, rather than being loaded every single time.

  • A project tells you what is wrong with it, in the project. The health check has always run here and reported on everything else, so the session that could fix a problem was the one session that never heard about it. It now speaks up where the work happens, names the rule and the fix, and says nothing at all when nothing is wrong.
  • The rules are put back at the moment they are most likely to be lost. When a long session drops context, the standing rules are restated and the team is asked to prove its roster is actually loaded rather than assert it.
  • The board now refuses what it used to merely describe. Seven rules every project was trusted to remember are enforced by the tool, and the proofs ship with it.
  • Both new hooks record that they ran, which closes a gap that had been open for weeks: nothing could tell a hook that did nothing from a hook that was never running.
  • A web address that does not exist now says so. Every unknown address on the site used to return the home page and report success, so a mistyped or out-of-date link never told anyone it was wrong, and search engines could file the same page under any number of junk addresses.
  • The reference page says what it is for, and what each board status means. It opens with its purpose instead of listing what you are assumed to know, and a table now gives every status, what it means, and who is allowed to move it.
  • The infrastructure standard names its source control. It was relied on throughout the document and missing from the list of defaults.
  • A deploy that silently stops happening is now written down as a known failure. A git connection can stop triggering builds while every dashboard still reports it healthy, so the release reports success over a site serving the previous version.

  • Release notes generated from a single source of truth, so the page and the record cannot drift, and a stale page is detected rather than shipped.
  • Every question an agent asks you now arrives as numbered options with a recommendation and an explicit way out, so you answer with one character instead of doing the analysis yourself.
  • A check that every project's imports resolve. One project had been loading a pointer to a document deleted seventeen days earlier, silently, while reporting healthy.
  • A clean-checkout test run, so what another person receives is what we actually tested.
  • Thirteen dangling references fixed in the published docs, including the starter template we ask people to copy.
  • A reference page on the site. The board and its columns, the infrastructure the method runs on and why each part was chosen, and a glossary of every term that is not ordinary English. All of it was already written down, in a repository, where you could only read it by cloning.
  • Six silent failures in the tool, fixed. In each one the tool reported success while the disk held something else: a role composed to an agent with no instructions, a half-applied install left behind by a command that said it had failed, a session-start rebuild that failed without a word.
  • A copy of the studio can no longer reach your real projects. Copying the tree to work somewhere safe now does what it looks like it does, and says so when it cannot.
  • Forty-eight new checks, and twenty-nine deliberate breakages run against them to confirm they go red rather than assuming it. Three of the breakages found the check itself was faulty.

  • A defined intake path for anything you say. Captured before work begins and before the reply, inherited by all sixteen roles instead of the nine that happened to carry it.
  • A required acknowledgement back to you, carrying the reference, its queue, what it is blocked behind and when it starts. A receipt, not a promise.
  • Coverage on the dry run of the one irreversible operation, so a preview cannot execute for real or under-report what it will commit.

  • Rule inheritance across the roster. A shared rule is defined once and inherited at compose time. One edit instead of eleven, and a missing rule fails the build rather than composing an agent with a hole in it.
  • Your ideas get pressure-tested before anyone builds them. Six leads challenge a new idea from their own disciplines and can come back with a no. You find out an idea is weak at the cheapest possible moment, rather than after you have paid for it.
  • A safe mode for automated runs, refusing writes to live projects by default. Documented intent is not an access control; this is.

  • A concurrency limit, two large items and three small. Ideas arrive faster than anything finishes; without a ceiling the squad context-switches across five threads and completes none. At the limit you get the count and a question, not silent queueing.
  • Guaranteed capture of anything you say. Everything gets written down. Not everything gets started.
  • A self-assessment each session, read by the next before it plans, so standards slipping is visible from the centre.
  • The health check now covers the studio itself, which had been the one thing it never looked at, so a problem in the place that owns the rules is now as visible as a problem anywhere else.
  • Each project can now reach only its own data, so one project's credentials are worth nothing anywhere else.

  • A step that forces a project to look at what it earned, cost and attracted. One project here took card payments for six weeks with nothing in its record about whether a single order landed.
  • A monthly check on every project's numbers, enforced centrally. An owner can skip it, but the skip is recorded as their decision. Skipping is allowed. Saying nothing is not.

  • A load check on every agent in your roster, the set of sixteen each project runs. Thirteen had been failing to parse for weeks, in every project, while reporting as installed and current.
  • Presence is never accepted as proof of working, so you cannot be shown a healthy roster that no session can load.

  • Some data issues in the published repository fixed.
  • Verified removal. A correction is confirmed against the live source rather than reported complete and assumed.

  • Self-testing credential scans. Every pattern is asserted against a known-bad sample, and the publish refuses if any fails its own test.
  • The defects a first clean install finds, fixed at the source, found by provisioning our own board from scratch, so the next project does not rediscover them. They are named below.
  • A quota check before provisioning. Free-tier limits are account-wide, so one project cannot consume what the next one needed.
  • Deleting a ticket now hides it and keeps it, and that protection holds however the data is reached, not only through the board's own screens.

  • Five acceptance criteria before a review may pass an animated page, so nothing is signed off because it happened to render on a fast machine.
  • Skipped checks reported as having proved nothing, so you are never shown a green result that never ran.

  • A shared multi-tenant backend. Every project's board on one free instance, all visible in one place.
  • Row-level isolation at the database, applied to the page, the command line and any direct API call alike.
  • Per-tenant ticket numbering, so one project's sequence reveals nothing about another's volume.
  • Runnable proofs of that isolation, including the negative case, so you can verify it rather than trust it.

  • Direction-aware drift detection. The tool now tells stale apart from locally-modified, so a project cannot overwrite an improvement that exists in only one place.
  • A default infrastructure standard, so projects stop picking stacks independently. Every service in it is free until a project earns revenue.
  • A check that a project's state is actually loaded, so history is not written to a file nothing imports.

  • Four checks before anything is committed, so a project cannot put a secret into its history. Removing one afterwards is a rewrite, not a delete.
  • An atomic release. One command, both repositories, one changelog entry, so they cannot disagree about what shipped.
  • Build and verification separated. Nobody marks their own work ready, and test notes say what to expect.
  • Persisted session state, so a project's history survives the session and you can resume it without interrogating it.

  • Startup Studio, first release. Sixteen specialist agents, one shared roster, composed per project. Engineering, design, content, marketing, operations, QA and review, all working from one board and one set of rules, so a solo founder runs a product team instead of a prompt.
  • A fixed board schema, 8 statuses and 7 columns. Every agent is written against it, so a renamed column breaks the contract instead of quietly diverging.
  • Base-plus-overlay composition, so a project overrides a role without forking the roster, and any improvement propagates to everyone.
  • A single intake path. Work comes from a ticket, never from conversation.
  • A human approval gate. No agent marks its own work ready, and nothing reaches production without your instruction.

The full technical detail behind every release is in the changelog.