Skip to content

feat(automations): Run automations in their own mode with a declared result - #2052

Merged
dcramer merged 31 commits into
mainfrom
feat/automation-run-mode
Oct 9, 2026
Merged

dcramer merged 31 commits into
mainfrom
feat/automation-run-mode

Conversation

@sentry-junior

@sentry-junior sentry-junior Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Scheduled automation and Event automation runs used the chat Turn contract. The runtime delivered the final assistant text, the model had to emit [[NO_REPLY]] to stay silent, and the interactive rules told an unattended run to ask questions. This is the root cause of most of #2014, #1741, and part of #554.

Automation runs now have their own mode, and each run ends with one declared result. Only a declared message is posted.

What changes

  • One switch. buildDispatchRoutingContext sets dispatch.declaresResult for an automation Source. The prompt, the tools, the result check, and Slack Delivery read that field. Watches and other plugin dispatches keep the chat Turn contract.
  • Automation prompt. <automation-run> rules replace the task, conversation, and Slack action rules. The stored instruction is the job. Nobody can answer a question during the run. The run checks every condition in the instruction before an action with side effects.
  • Declared result. finishAutomationRun is registered only for these runs. It must be the only tool call in its message, and the run stops after it.
    • send_message: the runtime posts the declared message to the stored outcomes. The tool does not offer this result when the automation has no message outcome.
    • no_action: the dispatch completes and nothing is posted.
    • misconfigured: the dispatch is recorded as blocked with the declared reason.
    • A message that is only [[NO_REPLY]] is saved as no_action, so instructions written before this change stay silent.
  • Who receives the message. The run context lists each stored outcome and says if it is the creator's direct message or a channel.
  • Who is mentioned. The task input line for the creator now says that "me" and "my" mean the creator, to mention them in the message, and not to mention them otherwise.
  • A run without a result. The model gets one reminder. A second stop without a result fails the dispatch.
  • A result is final. A resumed slice returns a saved result and does not call the model again.

Behavior changes to accept

  • A failed run posts nothing to its outcomes. Before, the channel got a failure reply. The failure shows on the dispatch, in the execution history, and as the last run status on the dashboard.
  • The model can end a run as misconfigured. For a Scheduled automation, the heartbeat already blocks the automation when its dispatch is blocked. For an Event automation, only that dispatch is blocked, and the automation runs again on the next event.
  • A person who replies in the thread of a posted message continues the run's Conversation with a normal chat Turn. That Turn sees the history of the run.

Not in this PR

Review focus

  • automation-result.ts owns the result contract. finishedRunReply decides what a finished run posts, for first runs in turn.ts and resumed runs in resume.ts.
  • The reminder loop and the early return for a saved result are in agent/index.ts. Unit tests cover their rules. No end-to-end test covers the reminder, because a real model cannot be made to skip the tool in a reliable way.
  • The agent test fixture now takes a person's reply in the thread of an automation's first post. One eval uses it to check the follow-up reply. The Slack mock now keeps a top-level post as the root of its thread.
  • The creator line changes the same text as fix(automations): Require the creator mention in a reply for "me" #2033.

The new prompt and tool change the model requests of the automation evals, so CI records them again. The chat prompt and the chat tools do not change.

Refs #2051

via David Cramer.

--

View Junior Session [Sentry]

…result

Scheduled automation and Event automation runs used the chat Turn contract.
The runtime delivered the final assistant text, and the model had to emit
[[NO_REPLY]] to stay silent. Interactive rules told unattended runs to ask
questions.

Automation runs now use their own prompt mode and end with one
finishAutomationRun call:

- send_message sends the declared message to the stored outcomes.
- no_action completes the dispatch and posts nothing.
- blocked records a blocked dispatch with the declared reason. A blocked
  scheduled automation stops until someone resumes it.

Final assistant text is never delivered. A run that stops without a result
gets one reminder, then fails. The prompt lists the stored outcome
destinations. Watches and other plugin dispatches keep the chat Turn
contract.

Refs #2051

Co-Authored-By: David Cramer <david@sentry.io>
@vercel

vercel Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
junior-docs Ready Ready Preview Oct 9, 2026 6:05pm UTC

Request Review

Give each automation steering rule one home. The tool description owns
result selection, the system prompt owns run rules, and task input only
names the closing tool. Drop the duplicate dispatch.delivery line and
share turn diagnostics between chat and automation results.

Merge the thread-binding and multi-outcome tests into one delivery test.
Move the missing-result reminder out of the routing test into its own
table that also covers the failed second stop.

Refs #2051

Co-Authored-By: David Cramer <david@sentry.io>
Add integration evals that run due and event automations through the
agent test fixture:

- A due automation with no message outcome does its work, posts
  nothing, and ends with no_action.
- An event automation whose condition does not match the event posts
  nothing and ends with no_action.

Add a behavioral eval for a due automation that needs a service Junior
cannot reach. It should end as blocked instead of posting a status
message.

Remove static prompt string assertions from the dispatch integration
test and the prompt unit test, as policies/testing.md requires. Drop the
task-input unit case that the event-automations snapshot already covers.

Refs #2051

Co-Authored-By: David Cramer <david@sentry.io>
@sentry-evals

sentry-evals Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Sentry Evals

junior-behavioral

Run Status Passed Failed Errored Cost Tokens Duration
Head Complete 69 4 0 $1.06 0 257.43s

Base commit 949d1b9 has no eval run to compare against.

junior-guardian

Run Status Passed Failed Errored Cost Tokens Duration
Head Complete 53 0 5 $0.01 0 1.38s

Base commit 949d1b9 has no eval run to compare against.

junior-integration

Run Status Passed Failed Errored Cost Tokens Duration
Head Complete 91 0 0 $0.51 0 183.00s

Base commit 949d1b9 has no eval run to compare against.

junior-router

Run Status Passed Failed Errored Cost Tokens Duration
Head Complete 12 0 0 $0.00 0 0.27s

Base commit 949d1b9 has no eval run to compare against.

… flow

Map the production failures from #2014, #554, and #1741 to realistic
automation runs:

- A due reminder for the creator's direct message is the reminder
  itself, not a failure note or third-person text.
- Channel reminders that name nobody mention nobody.
- An event automation whose condition does not match posts nothing,
  also when its task spells out the old [[NO_REPLY]] marker.
- Creator-bound Sentry work without a connected account does not ask
  the channel to connect it or wait for authorization.

The Sentry case replaces the PagerDuty case, which tested the same rule
with a made-up service. insertScheduledAutomation now takes sendTo
destinations, and slackDirectMessage() returns a direct message channel.
The scheduler eval README maps each failure to its case.

Refs #2051

Co-Authored-By: David Cramer <david@sentry.io>
@dcramer
dcramer marked this pull request as ready for review October 6, 2026 23:28
@github-actions github-actions Bot added the risk: high PR risk score: high label Oct 6, 2026

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/junior/src/chat/services/turn-result.ts Outdated
A run that yielded right after finishAutomationRun lost its result on
resume. The resumed slice read only its own new messages, so it either
got a missing-result reminder or failed and posted the error reply.

Read the declared result from the whole run history. An Automation run
owns its dispatch Conversation, so that history is only this run. The
recovery test now pauses right after the declared result.

Refs #2051

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/junior/src/chat/services/turn-result.ts
A declared Automation result decides the run outcome. A provider error
on a later slice no longer discards a result saved before a yield.
sentry-junior Bot and others added 2 commits October 7, 2026 01:27
A resumed slice re-prompted the model after an earlier slice had saved
finishAutomationRun. The model could act again, and a second result
replaced the first. A saved result now ends the run without a model call.

Co-Authored-By: David Cramer <david@sentry.io>
…red messages

The model could end a run as blocked for any problem, and that suspended
a Scheduled automation. The result is now misconfigured. Its tool
description limits it to an automation that cannot work until its creator
changes it, and says that it suspends the automation.

A declared message now gets the same cleanup as a chat reply. A message
that is only the old no-reply marker becomes no_action, so it is never
posted.

Co-Authored-By: David Cramer <david@sentry.io>
A misconfigured run blocked only Scheduled automations. Event automations
kept matching events. Add a blocked status and a stored status reason for
Event automations. markDispatchBlocked now blocks the Event automation, and
resume clears the reason. Show the blocked status and reason in the
Automations API, dashboard, and event automation tools.

Refs #2051

Co-Authored-By: David Cramer <david@sentry.io>
The model can still send the old no-reply marker as a message. The run
saves that as no_action and posts nothing, so assert on the saved result
instead of the call input.

Refs #2051

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/junior/src/chat/event-automations/store.ts Outdated
sentry-junior Bot and others added 2 commits October 7, 2026 02:17
The dispatch layer now sets declaresResult on the run. Prompt mode, tools,
the result check, and Slack Delivery read that field instead of the Source.
Turn and resume share one helper for what to post after a run.

A failed Automation run no longer posts the internal error reply to its
outcomes. The failure stays on the dispatch and execution record.

Outcome recipients in the prompt now say whether each one is the creator's
direct message or a channel.

Co-Authored-By: David Cramer <david@sentry.io>
When a dispatch first becomes blocked, its Scheduled automation or Event
automation creator gets one direct message with the reason and a resume
link. A redelivered block does not send it again. The notice is
best-effort because the reason also shows on the dashboard and in the
automation tools.

Declared Automation results are now a span attribute on first and
resumed runs.

Co-Authored-By: David Cramer <david@sentry.io>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/junior/src/chat/agent-dispatch/store.ts Outdated
…n blocks

A run can block after its creator pauses the Event automation. Before, the
block only applied to active automations, so the reason was lost and resume
made the automation active. Now a paused automation keeps the reason and
stays paused, and resume returns it to blocked.

Move the blocking scenario to its own integration test file so
event-automations.test.ts stays under the file-length limit.
dcramer and others added 4 commits October 9, 2026 08:09
Resolve conflicts in resume.ts and update-event-automation.ts on main's
structure. Regenerate the status_reason migration as 0049, because main
took 0048. Give slackDirectMessage() a stable fixture id.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep this change to the Automation run mode and its declared result.
The blocked status for Event automations, its migration, the dashboard
and chat pause and resume, and the creator notice move to a follow-up.

A declared misconfigured result still records a blocked dispatch. The
tool description no longer says that it suspends the automation, because
that is not true for an Event automation yet.

Restore the dispatch tests to the cases on main, with a declared result
as the model output. The agent test fixture owns new agent cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Unit tests own the fixed rules of a declared result: the reminder, a run
without a result, a saved result on a later slice, what each result
posts, and the tool input checks.

The agent test fixture now takes a person's reply in the thread of an
automation's first post. One eval checks that the reply gets a normal
chat reply. Another checks that a run which cannot work ends as
misconfigured without a post.

The Slack mock keeps a top-level post as the root of its thread, and it
reads a copy of a users.info request so the shared handler can answer
for a person it does not know.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
In Automation run mode, a reminder that says "me" lost its creator
mention in 3 of 8 samples, and a reminder that names nobody got a
mention in 1 of 8. The task input line now says that "me" and "my"
mean the creator, to mention them in the message, and not to mention
them when the instructions do not say "me" or "my". The three eval
cases then passed 10 of 10 samples each.

Correct the Automation runs section of the chat README: a reply in the
thread of a posted message continues the run's Conversation as a normal
chat Turn.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The behavioral evals recorded these model responses in https://github.com/getsentry/junior/actions/runs/37959302276.
The evals of this commit replay them in strict mode.

Eval-Suite: behavioral
The integration evals recorded these model responses in https://github.com/getsentry/junior/actions/runs/37959302398.
The evals of this commit replay them in strict mode.

Eval-Suite: integration
Comment thread packages/junior/src/chat/agent/tools.ts Outdated
Comment thread packages/junior/src/chat/providers/slack/turn.ts Outdated
].join(" ");

/** Create the tool that ends one Automation run with a declared result. */
export function createFinishAutomationRunTool(options: {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this just be an expected output format instead of a tool?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

alright i cant say i love this... we're gonna suppress assistant text in automations

but mayb eits also fine. these apis are a bit limiting already and we can revisit later

Rename runGetsDelivery to deliversFinalText. Add a TODO for the fallback
that lets a dispatch record without outcomes send to its Destination.
Automations have stored outcomes since #1788, and older dispatch records
have passed the 7-day state TTL.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sted

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The behavioral evals recorded these model responses in https://github.com/getsentry/junior/actions/runs/37964318944.
The evals of this commit replay them in strict mode.

Eval-Suite: behavioral
Main added recordings for 151 model requests that this branch also
recorded. Keep the copies from main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cursor

cursor Bot commented Oct 9, 2026

Copy link
Copy Markdown

This PR is too large for Bugbot to review. It changes 27,981 lines and 3,956,572 characters. Split the change into smaller pull requests to get a review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The behavioral evals recorded these model responses in https://github.com/getsentry/junior/actions/runs/37969606898.
The evals of this commit replay them in strict mode.

Eval-Suite: behavioral
@dcramer
dcramer merged commit eb97769 into main Oct 9, 2026
52 of 54 checks passed
@dcramer
dcramer deleted the feat/automation-run-mode branch October 9, 2026 18:20
dcramer added a commit that referenced this pull request Oct 10, 2026
…2095)

A run that cannot get the authorization or credentials that it needs
records a blocked dispatch. Its Event automation stayed active and ran
again on each matching event, with the same result. A blocked dispatch
now sets the Event automation to `blocked` with the reason, and a
blocked automation does not match new events until its creator resumes
it. Scheduled automations already work this way.

**What changes**

- `markDispatchBlocked` blocks the Event automation of a dispatch before
it marks the dispatch, so a retry after a failed write still blocks it.
- The reason is stored in a new `status_reason` column (migration
`0049_event_automation_status_reason`).
- Resume clears the reason and makes the automation active. Pause keeps
the reason. An automation that was paused when its run blocked stays
paused, and resume returns it to `blocked`.
- The dashboard, the Automations API, and the event automation tools
show `blocked` and the reason. A blocked automation counts as "needs
attention".
- `updateEventAutomation` takes `status: "active" | "paused"`, so a
person can pause or resume an Event automation from chat, as
`slackScheduleUpdateAutomation` does for Scheduled automations.

**Review focus**

- This PR tells nobody about the block. The creator sees it on the
dashboard and in the automation tools. A direct message to the creator
is a follow-up on top of this branch.
- A model can end a run as blocked only after #2052. Without #2052, the
blocked dispatches are the missing-authorization and credential failures
in `persistBlockedDispatchTurn`.
- The schema of `updateEventAutomation` is in the tool list of each chat
turn, so CI records the chat eval model requests again.

This is the second of three parts of the first version of #2052.

Refs #2051

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
dcramer added a commit that referenced this pull request Oct 10, 2026
#2096)

A blocked Scheduled automation or Event automation stops in silence. The
reason shows only on the dashboard and in the automation tools. The
creator now gets one direct message with the reason and a link to resume
the automation.

**What changes**

- `automations/blocked-notice.ts` sends the notice. It goes only after
the automation is stored as blocked, so the dashboard matches the
message:
- Event automations: when `blockEventAutomation` stores a new block. A
redelivered block sends nothing.
- Scheduled automations: when the heartbeat stores the block
(`advanceScheduledAutomationAfterRun` now returns whether it did). If
the creator changed the schedule during the run, the automation stays
active and no notice goes out.
- The notice is best-effort. The automation is already blocked when it
is sent, and a failed send logs `automation.blocked_notice.failed`.
- A missing-authorization reason now names the fix. On creator
credentials it reads "This run needs a connected `<provider>` account."
A run on system credentials cannot use a connected account, so its
reason says to switch to creator credentials and connect one.
- The notice and the creator direct-message outcome now share
`openSlackDirectMessage`.

**The result guidance of `finishAutomationRun` changes too**

For a job that needs a provider Junior does not have, an Event
automation run ended as `no_action` in 7 of 8 samples. Its reason said
that the result "could not be verified", and the `no_action` guidance
named exactly that. The automation then never became blocked, and no
notice went out. The guidance now says that `no_action` is for a run
that worked with nothing to post, or for a temporary failure, and that a
missing provider, tool, account, or access is `misconfigured`. The same
case then ended as `misconfigured` in 8 of 8 samples. One integration
eval covers the chain: the run ends as `misconfigured`, the only post is
the blocked notice, and it is in the creator's direct message.

This changes the tool description of every automation run, so CI records
the automation evals again.

**Review focus**

- The notice uses the same Slack token resolution as other background
posts. A binding of dispatch work to one workspace installation belongs
in a separate change.
- The two reason texts have no test of their own. The blocking and
heartbeat tests cover when a notice is sent and when it is not.
- The heartbeat also blocks a Scheduled automation before dispatch, when
it cannot build the dispatch metadata or create the dispatch. That path
sends the notice too.

This is the last of three parts of the first version of #2052. The other
two, #2052 and #2095, are merged.

Refs #2051

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

This branch was successfully deployed

1 active deployment
Preview – junior-docs — 08903cca Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

risk: high PR risk score: high

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant