-
Notifications
You must be signed in to change notification settings - Fork 0
Document stopping and restarting agent work #60
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,127 @@ | ||
| --- | ||
| layout: ../../layouts/Doc.astro | ||
| title: Stopping and restarting | ||
| eyebrow: Docs | ||
| description: Recover agent work after a restart, crash, or account change without requiring a perfect shutdown. | ||
| --- | ||
|
|
||
| You should be able to restart your machine or agent app when you need to. | ||
| Preparation helps, but recovery must work without warning. The aim is to | ||
| preserve completed work and identify what still needs verification before | ||
| continuing. | ||
|
|
||
| ## Use the time you have | ||
|
|
||
| | Time available | What to do | What to expect | | ||
| | --- | --- | --- | | ||
| | None | Restart and recover afterward. | The last operation may have an uncertain result. | | ||
| | A little | Stop new work and allow one shared grace period, normally 60 seconds. | Some tasks may not acknowledge in time. | | ||
| | Planned maintenance | Choose a longer deadline for handoffs and known sensitive operations. | More complete evidence, with no guarantee against a crash. | | ||
|
|
||
| The grace period covers the whole affected machine, including dispatch and | ||
| waiting. Each task receives the same absolute deadline. It does not get another | ||
| minute when its message arrives. At the deadline, the coordinator reports what | ||
| is known and returns control to you. Missing replies mean preparation is | ||
| incomplete; they do not mean the machine is verified ready. | ||
|
|
||
| A deadline limits preparation. It does not authorize an agent to kill a running | ||
| deployment or restart the app automatically. If an operation is still running, | ||
| the report should identify it specifically so you can decide. Prompt-based | ||
| deadlines also cannot preempt every stalled host tool call. | ||
|
|
||
| ## Recover from three sources together | ||
|
|
||
| 1. **Checkpoints:** the task objective, current repository and branch, completed | ||
| work, next action, owners, and existing approvals and limits. | ||
| 2. **Newer logs:** task messages and tool results after the checkpoint. If there | ||
| is no checkpoint, start with recent history and expand only as needed. | ||
| 3. **Live state:** files, commits, running processes, ownership, and the actual | ||
| result of any remote operation. | ||
|
|
||
| A checkpoint can be stale. A log can stop between starting a push and recording | ||
| its result. A remote command can finish after the app disappears. Check the | ||
| remote branch or PR before repeating the action, and verify that the previous | ||
| worker is no longer writing before replacing it. | ||
|
|
||
| Recover tasks independently. An ambiguous deployment can wait for inspection | ||
| while another task resumes verified unfinished work. Completed work stays | ||
| complete; successful tests need repeating only when interruption or subsequent | ||
| changes invalidate their evidence. | ||
|
|
||
| Paste this into the existing task after reopening it: | ||
|
|
||
| ```text | ||
| Recover from interruption; preparation may be missing or incomplete. | ||
| Combine any checkpoint with newer relevant task logs and live Git, process, | ||
| remote-operation and ownership evidence. Verify uncertain effects before | ||
| repeating them. Preserve completed work, goals, approval scope and deliberate | ||
| pauses. Resume verified unfinished work within existing windows and budgets. | ||
| Use the installed recovery workflow and report evidence and next action. | ||
| ``` | ||
|
|
||
| Recovery does not require a new goal, new control tower, or replacement task. | ||
| One existing task can coordinate recovery on a machine that has no tower. | ||
|
|
||
| ## Verify that every selected task recovers | ||
|
|
||
| “Ready to restart” only means a task has prepared to stop. The coordinator still | ||
| owns the return trip: send an explicit recovery message after the restart and | ||
| verify each task's response and next action. A queued message or an idle task | ||
| does not prove recovery. | ||
|
|
||
| Give each restart an identifier and keep recovery acknowledgments in the existing | ||
| fleet record. Match the recipient task ID to the identity inside its handoff and | ||
| the handoff filename. Correct a mismatched path; never append one task's evidence | ||
| to another task's checkpoint. An acknowledgment for an older restart cannot | ||
| close the current one. | ||
|
|
||
| Distinguish a hold introduced for this restart from a pause that already existed. | ||
| Authorized recovery clears the temporary restart hold after verification; older | ||
| pauses and expired work limits stay in effect. A delayed preparation message | ||
| must not put an already-recovered task back to sleep. | ||
|
|
||
| Track each task as resumed, already complete, intentionally still paused, or | ||
| needing attention. For every missing acknowledgment, retain an owner and exact | ||
| next action. Do not call the fleet recovered while a selected task has silently | ||
| remained paused. Automatic recovery requires a verified continuation mechanism; | ||
| otherwise the handoff must name who sends the recovery message. | ||
|
|
||
| ## Record progress before shutdown | ||
|
|
||
| Use existing task transcripts and workflow records throughout ordinary work. | ||
| Update compact recovery facts when a target starts, ownership changes, a | ||
| meaningful step finishes, or work enters an external wait. For consequential | ||
| operations, record intent before execution and the result afterward. | ||
|
|
||
| The source pack's [restart and recovery procedure](https://github.com/shakacode/agent-workflows/blob/main/docs/agent-runner-restarts.md) | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This links to There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This links to Failure scenario: this PR merges before (or without) a matching |
||
| defines the installed skills and available recording support. Install or update | ||
| the pack and check its revision before relying on a newer helper. The procedure | ||
| distinguishes commands recorded through a wrapper from native tool calls that | ||
| still rely on their transcripts; it does not promise a universal logging hook. | ||
|
|
||
| Keep recovery records in approved private local storage outside disposable | ||
| worktrees. Network or shared-drive outages should not prevent local recovery. | ||
| Mirror or back up separately for disk-loss protection. Do not commit raw task | ||
| transcripts, credentials, or private operation details to public repositories. | ||
|
|
||
| Codex documents its [session transcript and app log locations](https://developers.openai.com/codex/app/troubleshooting/). | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This page adds several |
||
| Those records are useful evidence, but neither logs nor a local recorder can | ||
| guarantee that the last write survives every crash or that an external action | ||
| happens exactly once. | ||
|
|
||
| ## Usage limits and account changes | ||
|
|
||
| Reaching a usage limit is different from quitting the app. OpenAI says an | ||
| active turn can continue after reaching usage limits, subject to fair use. | ||
| That does not guarantee completion of every goal, queued turn or future worker. | ||
| See the current [OpenAI usage guidance](https://developers.openai.com/codex/pricing/#what-happens-when-you-hit-usage-limits). Recovery instructions should | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This links to an external page with a Failure scenario: OpenAI reorganizes |
||
| not buy credits, change models, or extend a task's budget without authorization. | ||
|
|
||
| Treat switching accounts as an access and recovery boundary. Verify the intended | ||
| account, workspace, repository access, connectors and actual runtime permissions | ||
| before continuing. Do not assume cloud tasks or authenticated connections transfer | ||
| between accounts. See [OpenAI authentication guidance](https://developers.openai.com/codex/auth/). | ||
|
|
||
| The workflow provides preparation and evidence-based recovery. It cannot keep | ||
| a local agent turn alive after its host exits or make an uncertain external | ||
| operation safe to replay without checking its result. | ||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
At narrow viewport widths or with increased browser text size, this three-column table can exceed the document width because each column contributes its longest unbreakable word to the table's intrinsic minimum width. The global stylesheet provides no responsive or overflow treatment for
.doc-body table, while.docreserves 48px for horizontal padding, so the table can force page-level horizontal scrolling. Add a responsive table treatment, such as contained horizontal overflow or cell word wrapping.Useful? React with 👍 / 👎.