Skip to content

Consolidate duplicate review and recovery guidance #98

Description

@raghubetina

Implementation status — 2026-09-27

PR #101 implements the bounded guidance cleanup at 46a6903887185b98f245571efd20a8f4fa6ff324. Both hosted Node jobs pass for that commit; deterministic package and release-compatibility checks pass. Three small local Claude instruction checks preserved approval reuse, ambiguous Compile stopping, and the optional changelog checkpoint. They use synthetic analysis and a controlled CLI, not a real Compilation or release.

The orchestrator merged PR #101 at 82628f77aabcd0c0722e147918492eedf0e1a329; its tree 37b5d862608ae9e48265d1e2c6e1e2c548724ee5 is the exact tested and independently approved tree. This completes the bounded source cleanup. It remains unpublished, and the catalog is unchanged. The separate Plan 0.23/PWA-key retirement companion is also complete in #102, merged through PR #103 at c69014e4751972a313b700bd678449b5b32130d7; its receipt verifies the matching landed Service contract. These source changes do not change package publication, the installed plugin, or the catalog.

Approved PWA guidance reduction — 2026-09-27

firstdraft/firstdraft#876 records the owner’s decision to drop broader PWA guidance for now, retaining only the useful head start of home-screen bookmark assets. In the Skill cleanup, omit PWA planning/interview/checklist and standalone qualification guidance; retain short asset-editing instructions or a link to the emitted assets guide. Follow-up owner direction also permits retiring the unnecessary application.pwa authoring choice; coordinate structural references and packaged schema/CLI pins with the Service contract change. The old option is retired in the now-merged source contract recorded by #102; this does not claim a deployed or published update. Do not replace it with another interview preference. Revisit wider PWA choices only after their product purpose and Compiler value are established. The guidance reduction is merged in #101; the structural contract change is complete in #102.

Approved decision — 2026-09-27

Approved; implementation is in PR #101 above. Consolidate duplicated review/recovery guidance while preserving a self-contained normal workflow in SKILL.md, references for additional detail, and brief reminders at consequential action boundaries. Preserve the existing authorization, interview, candidate/gap identity and recovery semantics. Implement after the public-reference repair in #99 / PR #100 so the two documentation changes do not overlap. The status above records the completed implementation; publication remains separate.

For tests, retain package/structural/link/CLI-contract checks and ordinary tests of executable helpers. Use small live-agent evaluations for consequential instruction changes and regressions, with observed behavior reported separately from static checks. Remove natural-language substring assertions that merely freeze wording or duplicated prose; exact technical commands, flags and serialized keys can still warrant contract checks. Do not replace those assertions with a broad evaluation framework, mandatory benchmark matrix, or a new publication gate.

The owner approved this disposition after the Skill-testing prior-art review. Related to Service #663, #824 and #517. This does not reopen #81's broad interview policy or change #82's list-content question.

Current evidence

At Skills 5e28b667ebeb76846ad4857e6d195ed01a5a8fa1, main SKILL.md:198–230 contains the detailed semantic read-back rule and links to modeling-guide.md#prepare-the-pre-compile-semantic-read-back, whose lines 356–398 repeat the review, digest, gap, authorization and no-meaning-loss obligations. Main lines 293–302 also repeat recovery rules from the diagnostics reference.

For example, both locations require displaying the matching valid AnalysisRun's GapSet digest and every ordered record. test/interview-evaluation-foundation.test.mjs pins long overlapping exact-substring lists in both files (257–291 and 321–349). A harmless move to one authoritative location can fail the suite even when the route and behavior remain intact.

Why it arose and what to preserve

Commit 2d6edee825e6d40fdb0786e1d6e21b346a0d7f8d made local output primary while retaining existing authorization and ambiguous-effect recovery; later sample-data work added first-preview guidance. Those behavioral boundaries remain important. The existence of prose assertions is not evidence that the complete policy must be duplicated.

Keep the main normal action sequence, prominent implementation-notes lifecycle, capability/version check, identity preservation, review checkpoint, support boundary and recovery stop. Keep the complete normal workflow in the main Skill; use existing direct references for additional checklists and phase-specific detail without repeating the same policy. Preserve authored meaning, exact candidate/gap identity, local-output defaults and explicit GitHub selection.

The Agent Skills specification supports references loaded for a task. The current main Skill is already under its 500-line guideline; word count alone is not a defect and is not the basis for this proposal.

Tradeoff and verification

Consolidation reduces drift and wasted maintenance, but hiding a crucial rule behind a poor link can make it less likely to be encountered. Keep the normal workflow self-contained, direct routes to additional detail, and concise checkpoints; do not silently weaken the policy. A full Skill rewrite would add unjustified scope.

Retain package/link/CLI-contract checks. Reuse the small approved-candidate/no-extra-confirmation and ambiguous-Compile-stop scenarios affected by the move. A wording edit should not fail solely because wording changed; a broken route or changed decision should. No new requirements parser, broad benchmark, agent-continuation certification, or release gate is needed.

Prior-art refinement

The Augment, empirical-paper, Fizzy and Discourse research supports direct task routing, not a universal size rule. Teach the full policy once while keeping brief approval/recovery reminders where an external action occurs. Do not remove self-contained shell helpers merely because their declarations repeat across independent tool shells.

Existing interview-evaluation tests validate scenario definitions and prose; passing them does not execute a model or prove behavior. Retain useful structural/command contracts, directly inspect linked headings, and report any small agent smoke separately. No broad evaluation matrix or publication is authorized by this research.

Independent Opus research recommended keeping the complete normal-path policy in SKILL.md and trimming reference restatements; this is the approved starting point. A thin main Skill is not a requirement. Preserve the normal workflow without unnecessary reference hopping. The separate private-authority access problem is tracked in #99.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions