Fan swarm's external voices out per lens cluster (0.7.0) - #50
Open
gering wants to merge 5 commits into
Open
Conversation
codex and grok no longer run one broad multi-lens review each — they fan out over the SAME gated lens clusters as the Claude finders (one call per cluster; per lens under --max). The gate now prunes calls for everyone: a fully-gated-out cluster spawns nothing for any voice. - New adapter flag `agents.sh run --lens-instr <s>` prepends the workflow-supplied cluster briefs to the fenced-diff prompt in deterministic shell — the same "assembly is never an LLM step" contract the diff fencing follows, and backend-agnostic, so a future voice inherits per-cluster prompts for free. - LENS_BRIEF becomes single-source in swarm-review.js: the SKILL.md external-prompt HDR drops its hand-mirrored lens list. test_lens_sync.py flips that mirror check to a negative one and pins the --lens-instr wiring on both sides. - Lens tags become authoritative (the voice IS its cluster); an untagged finding from a single-lens external unit now resolves to that lens instead of falling back to 'unspecified'. - --max lifts every voice to per-lens granularity (previously Claude only). A claude:false control run keeps full-width external coverage. - Cost is logged at fan-out (live-backends × units, <=2x4 default, <=2x11 under --max) and never silently capped; an all-pruned gate logs why no external call ran. - Skill oversize threshold drops to 118784 B (4 KiB under the adapter cap) since exec now sees instruction + diff; briefs are asserted apostrophe-free and shell-quoted for the transport command. Grok needs no diff-only brief variant: since 0.6.0 both externals read project files, so both get the full cluster briefs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YWqcx14tMsuT7ptmywsMEQ
The 0.7.0 ensemble reviewed its own diff (12 voices, 17 findings). Applying the agreed ones: Gate blast radius (the critical finding). With externals gated too, a lens the low-effort haiku gate drops is now reviewed by NOBODY — the full-width external calls used to absorb a mis-gate, and the gate's only other protection is a sentence in its own prompt, i.e. injection-reachable via the diff it classifies. Add MANDATORY_LENSES (security, adversarial) as a code-level floor the gate cannot prune. Also materialize lenses the gate lists in neither `run` nor `skip` into `skip`, so the report can't under-report what was pruned — the self-review run silently dropped `adversarial` exactly this way. - agents.sh: reject a present-but-EMPTY --lens-instr instead of silently running a lens-free review whose findings the workflow still labels with the cluster's lenses (mislabeled coverage is worse than a hard error). An omitted flag stays legal for ad-hoc calls. Self-contained usage text; size messages render the limit from max_bytes instead of a hardcoded "120 KiB". - One unitBrief() now feeds both the Claude finder prompt and lensInstr(); the externals reuse finderUnits whenever a gate ran, so the two sides cannot drift. - backendErrors carry the unit + its lenses: every backend is multi-voice now, so "codex errored" alone hid WHICH cluster lost coverage. - LENS_BRIEF guard also rejects newlines/tabs/control chars — the transport command's one-line property depended on it. - SKILL.md: --max documented as splitting EVERY voice per lens (<=22 external calls), not just Claude; the external HDR is now lens- AND kind-agnostic, so a defect-only cluster is no longer invited to report design improvements. - test_lens_sync.py pins the skill's oversize threshold against the adapter's max_bytes and the largest instruction the briefs can produce. - Stale docs corrected: the workflow's drift warning and gate comment, the plugin.json description (now equal to marketplace.json), the README adapter synopsis, and the knowledge entry (per-cluster externals shipped, with the assembly/floor rationale recorded). Not applied: passing the instruction as a file payload (nobody can write that file — the workflow sandbox has no FS and the skill's prep runs before the gate exists; handing the write to the transport agent is the rejected LLM-assembly option), and keeping one external voice ungated (that would undo the cost control the per-cluster split exists for). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YWqcx14tMsuT7ptmywsMEQ
A codex+grok-only run (claude:false, 8 voices) reviewed the previous commit and
caught a real bug in its own mandatory-lens floor, plus test weaknesses:
- gate.run was still the gate's RAW pick after the floor re-added a lens, while
gate.skip had the floored entries removed — so a lens that actually ran
appeared in NEITHER column of the report's `Lenses: … — gated-out: …` line.
Write the floored set back to gate.run, and assert run+skip partition
CANDIDATE_LENSES (a mismatch logs a warning rather than misreporting silently).
- test_lens_sync.py derived its worst-case lens instruction from a Python COPY of
unitBrief()/lensInstr(). A sentence added to the JS template would have left the
copy stale and the headroom check green while real calls overflowed the adapter
cap. It now derives the fixed prose from the source template literals (${...}
stripped); verified byte-exact, and a 5 KiB block added to unitBrief() now fails
the check instead of passing it.
- The --lens-instr wiring assertion matched the bare string, which the new
comments satisfy on their own; it now matches the actual argument construction.
- Stale comment on the claude:false branch ("external voices review in full
regardless" — they inherit the gated set since 0.7.0) and on the LENS_BRIEF
loop (the SKILL mirror check is negative now). Restored the finder prompt's
colon→bullets order that the unitBrief refactor split.
Not applied: widening MANDATORY_LENSES to correctness (would pin 2 of 4 clusters
always-on and eat the cost control the per-cluster split exists for — the
coverage line now reports every pruned lens, so the tradeoff is visible), and a
content policy for --lens-instr in the adapter (its value is trusted caller
input; anyone who can set it already controls the whole argv).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YWqcx14tMsuT7ptmywsMEQ
- MANDATORY_LENSES gains `correctness` (user call): the gate prunes for every voice now, so the lens class most costly to miss must not depend on a haiku/effort-low call that reads the untrusted diff it classifies. Accepted cost, stated in code and CHANGELOG: `breakage` and `threat` are pinned always-on, leaving the gate only `design`/`consistency` to prune — a doc-only diff still pays 2 clusters × live voices (9 voices worst case, 6 if the gate returns nothing at all). - /swarm:review description 33 → 29 words, back inside the activation-surface budget check-structure.py enforces. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YWqcx14tMsuT7ptmywsMEQ
Third ensemble round (codex+grok only, 8 voices) over the full branch diff —
first external look at the mandatory floor itself. Two cross-family findings hit
the previous commit's own hardening:
- MANDATORY_LENSES is an explicit list, i.e. a MIRROR: renaming a lens in
LENS_CLUSTERS would leave a stale entry that can never match, silently voiding
the floor while every check stayed green. Assert the subset at startup and pin
it in test_lens_sync.py (proven: a rename now fails CI). Kept explicit rather
than derived from LENS_CLUSTERS.threat — which lenses are non-negotiable is a
judgement call, not a consequence of cluster membership.
- The gate coverage rewrite copied the gate's `skip` entries verbatim, so a
hallucinated lens name ("typo") would render as a gated-out lens that does not
exist. Filter skip to known lenses and assert run/skip stay disjoint as well as
complete.
- New `--lens-instr-bytes`: the empty-value check only caught a TOTAL loss, so a
transport that shortened or paraphrased the instruction still ran and had its
findings attributed to lenses the backend was never told to review — quietly
hollowing out "the voice IS its cluster". The caller declares the byte count it
built; a mismatch is a hard error. Length is computed in pure JS (the workflow
sandbox has no Buffer/TextEncoder; verified to match Buffer.byteLength on
multi-byte input) and memoized so command and count describe one build.
- Corrected an overstatement I wrote: the floor guarantees cluster SPAWN, not
lens coverage — within `breakage` the gate may still prune removed-behavior /
cross-file-trace, leaving that unit at lenses:['correctness'] for every voice.
Disclosed via gated-out, but "breakage ran" != "deletions were reviewed".
- Stale docs: README gate bullet (Claude-only pruning, skippable security),
knowledge isolation paragraph (per-voice now), miscounted list intro.
- Recorded the two repeatedly-refound-but-declined items (prose oversize gate,
per-unit grok probe) as accepted residuals with reasons, so later rounds see
them as decided rather than overlooked.
Not applied: widening the floor to the methodological lenses — that is the same
cost/coverage tradeoff the user already ruled on once, and the pruned lenses are
reported.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YWqcx14tMsuT7ptmywsMEQ
Owner
Author
|
@claude review |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
codexandgrokno longer run one broad multi-lens review each — every voice now fans out over the same gated lens clusters as the Claude finders (one call per cluster; per lens under--max). The gate prunes calls for everyone: a fully-gated-out cluster spawns nothing for any voice.[lens]tag becomes authoritative — the voice is its cluster, instead of a broad prompt self-tagging.LENS_BRIEFbecomes genuinely single-source: the SKILL.md external-prompt HDR drops its hand-mirrored lens list.MANDATORY_LENSESfloor (security,adversarial,correctness) is the code-level backstop, and the report's coverage line now partitions the full lens set.live-backends × unitsexternal calls — ≤2×4 by default, ≤2×11 under--max— logged at fan-out, never silently capped.Changes
Fan-out topology (
workflows/swarm-review.js)unitsFor()is the shared granularity ladder for both sides; externals reusefinderUnitswhenever a gate ran, so the two cannot drift.unitBrief()feeds both the Claude finder prompt and the externallensInstr().MANDATORY_LENSESfloor;gate.run/gate.skipare rewritten so together they partitionCANDIDATE_LENSES(a mismatch logs a warning rather than misreporting).backendErrorscarry the unit + its lenses — every backend is multi-voice now, so "codex errored" alone hid which cluster lost coverage.LENS_BRIEFguard rejects apostrophes, newlines, tabs and control chars (the transport command must stay one line).Adapter (
scripts/agents.sh)run --lens-instr <s>: prepends the workflow-supplied cluster briefs to the fenced-diff prompt in deterministic shell — the same "assembly is never an LLM step" contract the diff fencing follows, and backend-agnostic, so a future voice inherits per-cluster prompts for free.max_bytesinstead of a hardcoded string.Skill + tests + docs
--maxdocumented as splitting every voice per lens (≤22 external calls), not just Claude; the external HDR is now lens- and kind-agnostic.test_lens_sync.py: the SKILL mirror check flips to a negative assertion, the--lens-instrwiring is pinned to the actual argument construction, and the oversize headroom is checked againstmax_byteswith the worst-case instruction derived from the source template (a Python copy would have gone stale silently).swarm-review-pipelineknowledge entry.Readiness
plugin.json+marketplace.jsonin syncswarm-review-pipeline.md)check-structure.py— 0 errorsTest plan
Verified by two real ensemble runs against this branch's own diff:
External fan-out: 8 call(s) — codex + grok × 4 cluster(es); 0 backend errors, so the ~1 KB shell-quoted--lens-instrround-tripped through every transport agent; nounspecifiedlens tags inrawPerLens; 8 cross-family consensus clusters.claude: false, 8 voices): confirmed the no-gate path falls back to full-widthCANDIDATE_LENSES; 0 backend errors again (16/16 calls clean across both runs).gate.runnot rewritten, so a floored-in lens ran but appeared in neither report column).--lens-instr, missing flag value.unitBrief()→6036 B > 4096 B).MANDATORY_LENSEScost tradeoff — flooringcorrectnesspinsbreakage+threatalways-on, leaving the gate onlydesign/consistencyto prune.🤖 Generated with Claude Code
https://claude.ai/code/session_01YWqcx14tMsuT7ptmywsMEQ