Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions src/layouts/Base.astro
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,7 @@ const ogImage = new URL('/og.png', Astro.site);
<a href="/docs/terminology/">Terminology</a>
<a href="/docs/quickstart/">Quickstart</a>
<a href="/docs/architecture/">Architecture</a>
<a href="/docs/restarts/">Stopping and restarting</a>
<a href="/docs/distributions/">Distributions</a>
<a href="/methodology/">Methodology</a>
<a href="/case-studies/">Case studies</a>
Expand Down
1 change: 1 addition & 0 deletions src/pages/docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@ pack: its versioned skills, workflow rules, and technical docs define what agent
objective.
- [Quickstart](/docs/quickstart/) — install the pack and run your first skill in minutes.
- [Architecture](/docs/architecture/) — the pack, the seam, and the coordination backend.
- [Stopping and restarting](/docs/restarts/) — optional preparation and recovery after interruption.
- [Astra and model routing](/docs/astra/) — the advisory pilot, portable fallbacks, and evaluation guidance.
- [Git-native distributions and customization](/docs/distributions/) — the target model for
trusted Upstream Releases, explicit forks, and contribution back.
Expand Down
127 changes: 127 additions & 0 deletions src/pages/docs/restarts.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
---
layout: ../../layouts/Doc.astro
title: Stopping and restarting
eyebrow: Docs
description: Recover agent work after a restart, crash, or account change without requiring a perfect shutdown.
---

You should be able to restart your machine or agent app when you need to.
Preparation helps, but recovery must work without warning. The aim is to
preserve completed work and identify what still needs verification before
continuing.

## Use the time you have

| Time available | What to do | What to expect |
| --- | --- | --- |
| None | Restart and recover afterward. | The last operation may have an uncertain result. |
| A little | Stop new work and allow one shared grace period, normally 60 seconds. | Some tasks may not acknowledge in time. |
| Planned maintenance | Choose a longer deadline for handoffs and known sensitive operations. | More complete evidence, with no guarantee against a crash. |
Comment on lines +15 to +19

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the new comparison table responsive

At narrow viewport widths or with increased browser text size, this three-column table can exceed the document width because each column contributes its longest unbreakable word to the table's intrinsic minimum width. The global stylesheet provides no responsive or overflow treatment for .doc-body table, while .doc reserves 48px for horizontal padding, so the table can force page-level horizontal scrolling. Add a responsive table treatment, such as contained horizontal overflow or cell word wrapping.

Useful? React with 👍 / 👎.


The grace period covers the whole affected machine, including dispatch and
waiting. Each task receives the same absolute deadline. It does not get another
minute when its message arrives. At the deadline, the coordinator reports what
is known and returns control to you. Missing replies mean preparation is
incomplete; they do not mean the machine is verified ready.

A deadline limits preparation. It does not authorize an agent to kill a running
deployment or restart the app automatically. If an operation is still running,
the report should identify it specifically so you can decide. Prompt-based
deadlines also cannot preempt every stalled host tool call.

## Recover from three sources together

1. **Checkpoints:** the task objective, current repository and branch, completed
work, next action, owners, and existing approvals and limits.
2. **Newer logs:** task messages and tool results after the checkpoint. If there
is no checkpoint, start with recent history and expand only as needed.
3. **Live state:** files, commits, running processes, ownership, and the actual
result of any remote operation.

A checkpoint can be stale. A log can stop between starting a push and recording
its result. A remote command can finish after the app disappears. Check the
remote branch or PR before repeating the action, and verify that the previous
worker is no longer writing before replacing it.

Recover tasks independently. An ambiguous deployment can wait for inspection
while another task resumes verified unfinished work. Completed work stays
complete; successful tests need repeating only when interruption or subsequent
changes invalidate their evidence.

Paste this into the existing task after reopening it:

```text
Recover from interruption; preparation may be missing or incomplete.
Combine any checkpoint with newer relevant task logs and live Git, process,
remote-operation and ownership evidence. Verify uncertain effects before
repeating them. Preserve completed work, goals, approval scope and deliberate
pauses. Resume verified unfinished work within existing windows and budgets.
Use the installed recovery workflow and report evidence and next action.
```

Recovery does not require a new goal, new control tower, or replacement task.
One existing task can coordinate recovery on a machine that has no tower.

## Verify that every selected task recovers

“Ready to restart” only means a task has prepared to stop. The coordinator still
owns the return trip: send an explicit recovery message after the restart and
verify each task's response and next action. A queued message or an idle task
does not prove recovery.

Give each restart an identifier and keep recovery acknowledgments in the existing
fleet record. Match the recipient task ID to the identity inside its handoff and
the handoff filename. Correct a mismatched path; never append one task's evidence
to another task's checkpoint. An acknowledgment for an older restart cannot
close the current one.

Distinguish a hold introduced for this restart from a pause that already existed.
Authorized recovery clears the temporary restart hold after verification; older
pauses and expired work limits stay in effect. A delayed preparation message
must not put an already-recovered task back to sleep.

Track each task as resumed, already complete, intentionally still paused, or
needing attention. For every missing acknowledgment, retain an owner and exact
next action. Do not call the fleet recovered while a selected task has silently
remained paused. Automatic recovery requires a verified continuation mechanism;
otherwise the handoff must name who sends the recovery message.

## Record progress before shutdown

Use existing task transcripts and workflow records throughout ordinary work.
Update compact recovery facts when a target starts, ownership changes, a
meaningful step finishes, or work enters an external wait. For consequential
operations, record intent before execution and the result afterward.

The source pack's [restart and recovery procedure](https://github.com/shakacode/agent-workflows/blob/main/docs/agent-runner-restarts.md)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This links to agent-workflows/blob/main/docs/agent-runner-restarts.md, but the PR description says "the source-pack implementation is a companion change" — i.e. this file may not exist in main of agent-workflows yet. If that companion PR hasn't merged first, this ships a dead link. Also worth noting: scripts/check-links.mjs explicitly only validates internal links ("this script never touches the network, so it cannot and does not check external links" — see its header comment), so npm test passing does not prove this URL resolves. Worth confirming the companion doc is live before/at merge time.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This links to github.com/shakacode/agent-workflows/blob/main/docs/agent-runner-restarts.md, a filename not referenced anywhere else in this repo. Since this docs site and the agent-workflows source pack are separate repos, if the companion doc hasn't landed on main of agent-workflows yet (or lands under a different filename), this is a dead link at merge time and stays that way, since check-links.mjs only validates internal links and skips all http(s): targets.

Failure scenario: this PR merges before (or without) a matching agent-runner-restarts.md file existing in shakacode/agent-workflows, so the "restart and recovery procedure" link 404s for every reader with no build/CI failure to catch it.

defines the installed skills and available recording support. Install or update
the pack and check its revision before relying on a newer helper. The procedure
distinguishes commands recorded through a wrapper from native tool calls that
still rely on their transcripts; it does not promise a universal logging hook.

Keep recovery records in approved private local storage outside disposable
worktrees. Network or shared-drive outages should not prevent local recovery.
Mirror or back up separately for disk-loss protection. Do not commit raw task
transcripts, credentials, or private operation details to public repositories.

Codex documents its [session transcript and app log locations](https://developers.openai.com/codex/app/troubleshooting/).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This page adds several developers.openai.com/codex/... links (here, and lines 91 and 98). Like the agent-workflows link above, these are external URLs that scripts/check-links.mjs cannot and does not check (it's documented as offline/internal-only). If any of these paths are wrong or move, the site's npm test will still pass silently. Worth a manual click-through before merge, and periodically afterward, since nothing in CI will catch drift here.

Those records are useful evidence, but neither logs nor a local recorder can
guarantee that the last write survives every crash or that an external action
happens exactly once.

## Usage limits and account changes

Reaching a usage limit is different from quitting the app. OpenAI says an
active turn can continue after reaching usage limits, subject to fair use.
That does not guarantee completion of every goal, queued turn or future worker.
See the current [OpenAI usage guidance](https://developers.openai.com/codex/pricing/#what-happens-when-you-hit-usage-limits). Recovery instructions should

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This links to an external page with a #what-happens-when-you-hit-usage-limits fragment. scripts/check-links.mjs (per its own header comments) only checks internal links and explicitly never touches the network, so if OpenAI restructures/renames that page or heading, this link (and the other three new external links in this file: lines 72, 83, 99) will 404 or land on the wrong anchor silently, with no CI signal.

Failure scenario: OpenAI reorganizes developers.openai.com/codex/pricing (product docs churn is common), the fragment or page moves, and the doc keeps citing stale/dead evidence for a specific behavioral claim about usage limits — readers get a broken citation with no automated check ever catching it.

not buy credits, change models, or extend a task's budget without authorization.

Treat switching accounts as an access and recovery boundary. Verify the intended
account, workspace, repository access, connectors and actual runtime permissions
before continuing. Do not assume cloud tasks or authenticated connections transfer
between accounts. See [OpenAI authentication guidance](https://developers.openai.com/codex/auth/).

The workflow provides preparation and evidence-based recovery. It cannot keep
a local agent turn alive after its host exits or make an uncertain external
operation safe to replay without checking its result.
Loading