Repository navigation
feat(automations): Run automations in their own mode with a declared result - #2052
Conversation
…result Scheduled automation and Event automation runs used the chat Turn contract. The runtime delivered the final assistant text, and the model had to emit [[NO_REPLY]] to stay silent. Interactive rules told unattended runs to ask questions. Automation runs now use their own prompt mode and end with one finishAutomationRun call: - send_message sends the declared message to the stored outcomes. - no_action completes the dispatch and posts nothing. - blocked records a blocked dispatch with the declared reason. A blocked scheduled automation stops until someone resumes it. Final assistant text is never delivered. A run that stops without a result gets one reminder, then fails. The prompt lists the stored outcome destinations. Watches and other plugin dispatches keep the chat Turn contract. Refs #2051 Co-Authored-By: David Cramer <david@sentry.io>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Give each automation steering rule one home. The tool description owns result selection, the system prompt owns run rules, and task input only names the closing tool. Drop the duplicate dispatch.delivery line and share turn diagnostics between chat and automation results. Merge the thread-binding and multi-outcome tests into one delivery test. Move the missing-result reminder out of the routing test into its own table that also covers the failed second stop. Refs #2051 Co-Authored-By: David Cramer <david@sentry.io>
Add integration evals that run due and event automations through the agent test fixture: - A due automation with no message outcome does its work, posts nothing, and ends with no_action. - An event automation whose condition does not match the event posts nothing and ends with no_action. Add a behavioral eval for a due automation that needs a service Junior cannot reach. It should end as blocked instead of posting a status message. Remove static prompt string assertions from the dispatch integration test and the prompt unit test, as policies/testing.md requires. Drop the task-input unit case that the event-automations snapshot already covers. Refs #2051 Co-Authored-By: David Cramer <david@sentry.io>
Sentry Evalsjunior-behavioral
Base commit junior-guardian
Base commit junior-integration
Base commit junior-router
Base commit |
… flow Map the production failures from #2014, #554, and #1741 to realistic automation runs: - A due reminder for the creator's direct message is the reminder itself, not a failure note or third-person text. - Channel reminders that name nobody mention nobody. - An event automation whose condition does not match posts nothing, also when its task spells out the old [[NO_REPLY]] marker. - Creator-bound Sentry work without a connected account does not ask the channel to connect it or wait for authorization. The Sentry case replaces the PagerDuty case, which tested the same rule with a made-up service. insertScheduledAutomation now takes sendTo destinations, and slackDirectMessage() returns a direct message channel. The scheduler eval README maps each failure to its case. Refs #2051 Co-Authored-By: David Cramer <david@sentry.io>
A run that yielded right after finishAutomationRun lost its result on resume. The resumed slice read only its own new messages, so it either got a missing-result reminder or failed and posted the error reply. Read the declared result from the whole run history. An Automation run owns its dispatch Conversation, so that history is only this run. The recovery test now pauses right after the declared result. Refs #2051
A declared Automation result decides the run outcome. A provider error on a later slice no longer discards a result saved before a yield.
A resumed slice re-prompted the model after an earlier slice had saved finishAutomationRun. The model could act again, and a second result replaced the first. A saved result now ends the run without a model call. Co-Authored-By: David Cramer <david@sentry.io>
…red messages The model could end a run as blocked for any problem, and that suspended a Scheduled automation. The result is now misconfigured. Its tool description limits it to an automation that cannot work until its creator changes it, and says that it suspends the automation. A declared message now gets the same cleanup as a chat reply. A message that is only the old no-reply marker becomes no_action, so it is never posted. Co-Authored-By: David Cramer <david@sentry.io>
A misconfigured run blocked only Scheduled automations. Event automations kept matching events. Add a blocked status and a stored status reason for Event automations. markDispatchBlocked now blocks the Event automation, and resume clears the reason. Show the blocked status and reason in the Automations API, dashboard, and event automation tools. Refs #2051 Co-Authored-By: David Cramer <david@sentry.io>
The model can still send the old no-reply marker as a message. The run saves that as no_action and posts nothing, so assert on the saved result instead of the call input. Refs #2051
The dispatch layer now sets declaresResult on the run. Prompt mode, tools, the result check, and Slack Delivery read that field instead of the Source. Turn and resume share one helper for what to post after a run. A failed Automation run no longer posts the internal error reply to its outcomes. The failure stays on the dispatch and execution record. Outcome recipients in the prompt now say whether each one is the creator's direct message or a channel. Co-Authored-By: David Cramer <david@sentry.io>
When a dispatch first becomes blocked, its Scheduled automation or Event automation creator gets one direct message with the reason and a resume link. A redelivered block does not send it again. The notice is best-effort because the reason also shows on the dashboard and in the automation tools. Declared Automation results are now a span attribute on first and resumed runs. Co-Authored-By: David Cramer <david@sentry.io>
…n blocks A run can block after its creator pauses the Event automation. Before, the block only applied to active automations, so the reason was lost and resume made the automation active. Now a paused automation keeps the reason and stays paused, and resume returns it to blocked. Move the blocking scenario to its own integration test file so event-automations.test.ts stays under the file-length limit.
Resolve conflicts in resume.ts and update-event-automation.ts on main's structure. Regenerate the status_reason migration as 0049, because main took 0048. Give slackDirectMessage() a stable fixture id. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep this change to the Automation run mode and its declared result. The blocked status for Event automations, its migration, the dashboard and chat pause and resume, and the creator notice move to a follow-up. A declared misconfigured result still records a blocked dispatch. The tool description no longer says that it suspends the automation, because that is not true for an Event automation yet. Restore the dispatch tests to the cases on main, with a declared result as the model output. The agent test fixture owns new agent cases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Unit tests own the fixed rules of a declared result: the reminder, a run without a result, a saved result on a later slice, what each result posts, and the tool input checks. The agent test fixture now takes a person's reply in the thread of an automation's first post. One eval checks that the reply gets a normal chat reply. Another checks that a run which cannot work ends as misconfigured without a post. The Slack mock keeps a top-level post as the root of its thread, and it reads a copy of a users.info request so the shared handler can answer for a person it does not know. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
In Automation run mode, a reminder that says "me" lost its creator mention in 3 of 8 samples, and a reminder that names nobody got a mention in 1 of 8. The task input line now says that "me" and "my" mean the creator, to mention them in the message, and not to mention them when the instructions do not say "me" or "my". The three eval cases then passed 10 of 10 samples each. Correct the Automation runs section of the chat README: a reply in the thread of a posted message continues the run's Conversation as a normal chat Turn. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The behavioral evals recorded these model responses in https://github.com/getsentry/junior/actions/runs/37959302276. The evals of this commit replay them in strict mode. Eval-Suite: behavioral
The integration evals recorded these model responses in https://github.com/getsentry/junior/actions/runs/37959302398. The evals of this commit replay them in strict mode. Eval-Suite: integration
| ].join(" "); | ||
|
|
||
| /** Create the tool that ends one Automation run with a declared result. */ | ||
| export function createFinishAutomationRunTool(options: { |
There was a problem hiding this comment.
shouldn't this just be an expected output format instead of a tool?
There was a problem hiding this comment.
alright i cant say i love this... we're gonna suppress assistant text in automations
but mayb eits also fine. these apis are a bit limiting already and we can revisit later
Rename runGetsDelivery to deliversFinalText. Add a TODO for the fallback that lets a dispatch record without outcomes send to its Destination. Automations have stored outcomes since #1788, and older dispatch records have passed the 7-day state TTL. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sted Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The behavioral evals recorded these model responses in https://github.com/getsentry/junior/actions/runs/37964318944. The evals of this commit replay them in strict mode. Eval-Suite: behavioral
Main added recordings for 151 model requests that this branch also recorded. Keep the copies from main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
This PR is too large for Bugbot to review. It changes 27,981 lines and 3,956,572 characters. Split the change into smaller pull requests to get a review. |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The behavioral evals recorded these model responses in https://github.com/getsentry/junior/actions/runs/37969606898. The evals of this commit replay them in strict mode. Eval-Suite: behavioral
…2095) A run that cannot get the authorization or credentials that it needs records a blocked dispatch. Its Event automation stayed active and ran again on each matching event, with the same result. A blocked dispatch now sets the Event automation to `blocked` with the reason, and a blocked automation does not match new events until its creator resumes it. Scheduled automations already work this way. **What changes** - `markDispatchBlocked` blocks the Event automation of a dispatch before it marks the dispatch, so a retry after a failed write still blocks it. - The reason is stored in a new `status_reason` column (migration `0049_event_automation_status_reason`). - Resume clears the reason and makes the automation active. Pause keeps the reason. An automation that was paused when its run blocked stays paused, and resume returns it to `blocked`. - The dashboard, the Automations API, and the event automation tools show `blocked` and the reason. A blocked automation counts as "needs attention". - `updateEventAutomation` takes `status: "active" | "paused"`, so a person can pause or resume an Event automation from chat, as `slackScheduleUpdateAutomation` does for Scheduled automations. **Review focus** - This PR tells nobody about the block. The creator sees it on the dashboard and in the automation tools. A direct message to the creator is a follow-up on top of this branch. - A model can end a run as blocked only after #2052. Without #2052, the blocked dispatches are the missing-authorization and credential failures in `persistBlockedDispatchTurn`. - The schema of `updateEventAutomation` is in the tool list of each chat turn, so CI records the chat eval model requests again. This is the second of three parts of the first version of #2052. Refs #2051 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
#2096) A blocked Scheduled automation or Event automation stops in silence. The reason shows only on the dashboard and in the automation tools. The creator now gets one direct message with the reason and a link to resume the automation. **What changes** - `automations/blocked-notice.ts` sends the notice. It goes only after the automation is stored as blocked, so the dashboard matches the message: - Event automations: when `blockEventAutomation` stores a new block. A redelivered block sends nothing. - Scheduled automations: when the heartbeat stores the block (`advanceScheduledAutomationAfterRun` now returns whether it did). If the creator changed the schedule during the run, the automation stays active and no notice goes out. - The notice is best-effort. The automation is already blocked when it is sent, and a failed send logs `automation.blocked_notice.failed`. - A missing-authorization reason now names the fix. On creator credentials it reads "This run needs a connected `<provider>` account." A run on system credentials cannot use a connected account, so its reason says to switch to creator credentials and connect one. - The notice and the creator direct-message outcome now share `openSlackDirectMessage`. **The result guidance of `finishAutomationRun` changes too** For a job that needs a provider Junior does not have, an Event automation run ended as `no_action` in 7 of 8 samples. Its reason said that the result "could not be verified", and the `no_action` guidance named exactly that. The automation then never became blocked, and no notice went out. The guidance now says that `no_action` is for a run that worked with nothing to post, or for a temporary failure, and that a missing provider, tool, account, or access is `misconfigured`. The same case then ended as `misconfigured` in 8 of 8 samples. One integration eval covers the chain: the run ends as `misconfigured`, the only post is the blocked notice, and it is in the creator's direct message. This changes the tool description of every automation run, so CI records the automation evals again. **Review focus** - The notice uses the same Slack token resolution as other background posts. A binding of dispatch work to one workspace installation belongs in a separate change. - The two reason texts have no test of their own. The blocking and heartbeat tests cover when a notice is sent and when it is not. - The heartbeat also blocks a Scheduled automation before dispatch, when it cannot build the dispatch metadata or create the dispatch. That path sends the notice too. This is the last of three parts of the first version of #2052. The other two, #2052 and #2095, are merged. Refs #2051 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Scheduled automation and Event automation runs used the chat Turn contract. The runtime delivered the final assistant text, the model had to emit
[[NO_REPLY]]to stay silent, and the interactive rules told an unattended run to ask questions. This is the root cause of most of #2014, #1741, and part of #554.Automation runs now have their own mode, and each run ends with one declared result. Only a declared message is posted.
What changes
buildDispatchRoutingContextsetsdispatch.declaresResultfor an automation Source. The prompt, the tools, the result check, and Slack Delivery read that field. Watches and other plugin dispatches keep the chat Turn contract.<automation-run>rules replace the task, conversation, and Slack action rules. The stored instruction is the job. Nobody can answer a question during the run. The run checks every condition in the instruction before an action with side effects.finishAutomationRunis registered only for these runs. It must be the only tool call in its message, and the run stops after it.send_message: the runtime posts the declared message to the stored outcomes. The tool does not offer this result when the automation has no message outcome.no_action: the dispatch completes and nothing is posted.misconfigured: the dispatch is recorded as blocked with the declared reason.[[NO_REPLY]]is saved asno_action, so instructions written before this change stay silent.Behavior changes to accept
misconfigured. For a Scheduled automation, the heartbeat already blocks the automation when its dispatch is blocked. For an Event automation, only that dispatch is blocked, and the automation runs again on the next event.Not in this PR
blockedstatus for Event automations (feat(automations): Block an Event automation when its run is blocked #2095), and a direct message that tells the creator about a block (feat(automations): Tell the creator when an automation becomes blocked #2096). These were in the first version of this PR. They do not depend on this PR.sendFilesstill post to the automation's Destination, not to its outcomes.Review focus
automation-result.tsowns the result contract.finishedRunReplydecides what a finished run posts, for first runs inturn.tsand resumed runs inresume.ts.agent/index.ts. Unit tests cover their rules. No end-to-end test covers the reminder, because a real model cannot be made to skip the tool in a reliable way.The new prompt and tool change the model requests of the automation evals, so CI records them again. The chat prompt and the chat tools do not change.
Refs #2051
via David Cramer.
--
View Junior Session [Sentry]