Skip to content

[ab-advisor] Experiment campaign for smoke-project: A/B test prompt_style #37302

Description

@github-actions

🧪 Experiment Campaign: smoke-project

Workflow file: .github/workflows/smoke-project.md
Selected dimension: prompt_style
Triggered by: ab-testing-advisor on 2026-06-06


Background

smoke-project validates GitHub Projects API operations end-to-end: it creates a draft issue, adds PR #14477 and issue #14478 to an org project board, updates all three items to a new status, then posts a structured markdown status update. The current prompt is exceptionally dense — it spells out every update_project parameter inline (field names, values, temporary_id chains, ordering constraints, and a templated status-report body). Testing prompt_style variants will reveal whether this verbosity is load-bearing for correctness or whether a more concise or step-by-step form achieves the same pass rate with fewer input tokens.

Hypothesis

  • H0 (null): Variant prompt style does not change run_success_rate compared to the current detailed baseline.
  • H1 (alternative): The concise and/or step_by_step variants achieve ≥90% success rate (matching detailed) while reducing prompt token count by ≥30%, demonstrating that inline parameter verbosity is not required for correct execution.

Experiment Configuration

Add the following experiments: block to the workflow frontmatter:

experiments:
  prompt_style_test:
    variants: [detailed, concise, step_by_step]
    description: "Test whether reducing prompt verbosity preserves project-operation success rate"
    hypothesis: "H0: no change in run_success_rate. H1: concise/step_by_step achieve ≥90% success rate with ≥30% fewer prompt tokens"
    metric: run_success_rate
    secondary_metrics: [run_duration_ms, prompt_token_count, ops_completed_count]
    analysis_type: proportion_test
    guardrail_metrics:
      - name: empty_output_rate
        direction: min
        threshold: 0.05
      - name: missing_ops_rate
        direction: min
        threshold: 0.10
    min_samples: 20
    weight: [34, 33, 33]
    start_date: "2026-06-06"
    tags: [smoke-test, prompt-engineering, verbosity]
    notify:
      issue: <this_issue_number>

Variant descriptions:

  • detailed (baseline): Current form — all tool-call parameters, field names, and values spelled out inline per operation with explicit ordering constraints.
  • concise: High-level operation list referencing project URL; relies on the model's understanding of the update_project tool schema rather than restating every parameter.
  • step_by_step: Terse numbered one-liners per operation (no inline parameter blocks), with ordering enforced by numbering alone.

Workflow Changes Required

Replace the ## Test Requirements body section with a three-way conditional. Always compare against a specific variant value — use {{#if experiments.prompt_style_test == "<variant>" }} syntax:

-Default status field for any created items: "Todo".
-Do the following operations EXACTLY in this order.
-Do not re-create draft items but use their returned temporary-ids for the update operations.
-
-## Test Requirements
-
-1. **Add items**: Create items in the project using different content types:
-   a. update_project: content_type=draft_issue, draft_title="Test *draft issue* for `smoke-project`",
-      draft_body="Test draft issue for smoke test validation", temporary_id=draft-1,
-      fields={"Status":"Todo","Priority":"High"}
-   b. update_project: content_type=pull_request, content_number=14477,
-      fields={"Status":"Todo","Priority":"High"}
-   c. update_project: content_type=issue, content_number=14478,
-      fields={"Status":"Todo","Priority":"High"}
-2. **Update items**: ...(all inline parameter blocks)...
-3. **Project Status Update**: ...
+Default status field for any created items: "Todo".
+Do not re-create draft items but use their returned temporary-ids for the update operations.
+
+## Test Requirements
+
+{{#if experiments.prompt_style_test == "concise" }}
+Run the following project operations against `https://github.com/orgs/github/projects/24068`:
+
+1. Add a draft issue titled "Test *draft issue* for `smoke-project`" (temporary_id: draft-1, Status: Todo, Priority: High)
+2. Add PR #14477 (Status: Todo, Priority: High)
+3. Add issue #14478 (Status: Todo, Priority: High)
+4. Update the draft issue from step 1 to Status: In Progress
+5. Update PR #14477 to Status: In Progress
+6. Update issue #14478 to Status: In Progress
+7. Post a project status update with a markdown checklist of all operations performed.
+{{else if experiments.prompt_style_test == "step_by_step" }}
+Execute each step in order:
+
+1. `update_project` — draft_issue, title="Test *draft issue* for `smoke-project`", temporary_id=draft-1, Status=Todo, Priority=High
+2. `update_project` — pull_request #14477, Status=Todo, Priority=High
+3. `update_project` — issue #14478, Status=Todo, Priority=High
+4. `update_project` — draft_issue using temporary-id from step 1, Status=In Progress
+5. `update_project` — pull_request #14477, Status=In Progress
+6. `update_project` — issue #14478, Status=In Progress
+7. `create_project_status_update` — post a markdown pass/fail checklist covering all 6 operations above
+{{else}}
+(current detailed prompt body — all inline parameter blocks retained as-is)
+{{/if}}

Success Metrics

Metric Type Target
run_success_rate Primary ≥90% for all variants
run_duration_ms Secondary Monitor for regression
prompt_token_count Secondary ≥30% reduction for concise/step_by_step
ops_completed_count Secondary All 6 ops + status update = 7
empty_output_rate Guardrail Must stay < 5%
missing_ops_rate Guardrail Must stay < 10%

Statistical Design

  • Variants: detailed (baseline), concise, step_by_step
  • Assignment: Round-robin via gh-aw experiments runtime (cache-based)
  • Minimum runs per variant: 20 (for 80% power to detect a 15-point success-rate drop)
  • Expected run frequency: 1–3 per day (slash command, PR labeling, workflow_dispatch)
  • Expected experiment duration: ~15–30 days to reach minimum samples
  • Analysis approach: Two-proportion z-test (each variant vs detailed baseline); Bonferroni correction for the two comparisons (α = 0.025 per test)

Implementation Steps

  • Add experiments: section to frontmatter (YAML block above)
  • Add conditional blocks to workflow prompt body using {{#if experiments.prompt_style_test == "concise" }} (value-comparison form — never use the internal __GH_AW_EXPERIMENTS__ env-var syntax)
  • Run gh aw compile smoke-project to regenerate lock file
  • Monitor experiment artifact uploaded per run to /tmp/gh-aw/agent/experiments/state.json
  • After ≥20 runs per variant, analyze via workflow run artifacts
  • Document findings and promote winning variant
Infrastructure Status — All Fields Implemented ✅

✅ All three infrastructure fields are fully implemented — no sub-issue needed.

The field-presence-checker agent confirmed that analysis_type, tags, and notify are all end-to-end operational:

Field Compiler (Go) Picker (CJS)
analysis_type L194 — stored in cfg.AnalysisType L239, L263 — rendered in step-summary
tags L197-198 — parsed via parseStringSliceAny L240, L261-262 — rendered in step-summary
notify L200-213 — parsed into ExperimentNotify{Discussion,Issue} L241-246, L264-265 — rendered in step-summary

All three qualify as fully implemented and surfaced in run artifacts. No infrastructure sub-issue is required.

References

Generated by 🧪 Daily A/B Testing Advisor · 267.9 AIC · ⌖ 20.4 AIC · ◷

  • expires on Jun 20, 2026, 11:08 AM UTC

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions