# Install & use: `resilient-agent-handoff` (v3)

A self-contained installer and user guide for the `resilient-agent-handoff` Claude Code skill. Share this one file: whoever
receives it can install the skill on any machine, in any project, without cloning anything.

> **Claude Code — if you were asked to run, install or update from this file:** carry out the steps in
> **section 9** (everything between `~~~PASTE START~~~` and `~~~PASTE END~~~`) exactly as if the user had
> pasted them. That means: check existing versions (never downgrade), write the embedded `SKILL.md`, commit the
> tracked mirror, then run the skill's AUTO-UPDATE + AUTO-SETUP. Skip AUTO-SETUP only if the user said
> "install only". Sections 1-8 are documentation for humans; don't just summarize them.

## Contents
1. What it does
2. Requirements
3. Install or update (one-line prompt, or paste)
4. Manual install (shell)
5. What happens after install: auto-setup
6. Usage
7. Updating: auto-update
8. Uninstall
9. Paste block with the full skill

## 1. What it does

A playbook that keeps multi-agent work, or one long session, from losing progress when a turn dies
(rate limit, crash, context limit), and stops concurrent agents from colliding. It has five practices:

1. Commit every verified step immediately, with explicit paths.
2. Keep a `PROGRESS_<task>.md` resumable-state file per task, written BEFORE each step.
3. Use a mkdir-locked `AGENT_CLAIMS.md` claim ledger, checking liveness before reaping a claim.
4. Run long work as detached OS processes.
5. Keep an append-only achievement log.

It also includes a "recreate the system" prompt for resuming after an interruption.

## 2. Requirements

- **Claude Code:** CLI, desktop app, IDE extension or web.
- **git:** the skill activates inside a git repository.
- **A POSIX `sh`:** Git for Windows (Git Bash) on Windows; built in on macOS and Linux.
- **No other dependencies.**

## 3. Install or update (recommended)

**Option A: one-line prompt (easiest).** Open Claude Code in the project where you want the skill active,
with this file in that folder (or know its path), and type:

```
run and install INSTALL_RESILIENT_AGENT_HANDOFF.md
```

Other phrasings work the same way:
- `install the skill from C:/path/to/INSTALL_RESILIENT_AGENT_HANDOFF.md`
- `update resilient-agent-handoff from INSTALL_RESILIENT_AGENT_HANDOFF.md`
- `install INSTALL_RESILIENT_AGENT_HANDOFF.md - install only, don't run AUTO-SETUP` (installs the skill without changing the current repo)

Claude reads this file and carries out section 9. Claude Code may ask you to confirm first, because the file
installs instructions that future sessions will follow; your prompt naming the file is that approval.

**Option A from the web: don't have the file yet?** It's published at
<https://www.siriusart.bg/handoff-toolkit/INSTALL_RESILIENT_AGENT_HANDOFF.md>. Type:

```
download https://www.siriusart.bg/handoff-toolkit/INSTALL_RESILIENT_AGENT_HANDOFF.md with curl, then run and install it
```

Claude runs `curl -fsSLO https://www.siriusart.bg/handoff-toolkit/INSTALL_RESILIENT_AGENT_HANDOFF.md`, which downloads the exact file (a web-fetch tool may return a summary
instead), and then carries out section 9 from it.

**Option B: paste.** Paste the **whole block in section 9**, from `~~~PASTE START~~~` to
`~~~PASTE END~~~`, into Claude Code. This is useful when the session can't see the file, for example
when you received it by email or chat.

Either way, Claude will:

- **Install safely:** check for existing copies and never downgrade.
- **Install the skill:** globally, for every project on this machine.
- **Mirror it:** commit a tracked copy to `docs/handoff/resilient-agent-handoff-skill.md`, so it travels via git.
- **Activate it:** run AUTO-UPDATE + AUTO-SETUP in the current repo right away.

Running either option again later is safe: an older installed copy gets updated, and a newer one is kept.

## 4. Manual install (shell)

Without Claude, from the folder that contains this file, run in Git Bash, macOS or Linux:

```sh
n=resilient-agent-handoff
home="${USERPROFILE:+$USERPROFILE/.claude}"; home="${home:-$HOME/.claude}"
mkdir -p "$home/skills/$n"
awk '/^--- BEGIN SKILL.md CONTENT/{f=1;next} /^--- END SKILL.md CONTENT/{exit} f' INSTALL_RESILIENT_AGENT_HANDOFF.md > "$home/skills/$n/SKILL.md"
```

Then start Claude Code in a project and say *"run the resilient-agent-handoff skill's auto-setup"*.
This overwrites without a version check. The next invocation's AUTO-UPDATE restores a newer version if one exists in a repo mirror.

## 5. What happens after install: auto-setup

AUTO-UPDATE + AUTO-SETUP run **first on every invocation** and right after install. They're idempotent: they
only add or upgrade what's missing, and everything they create in a project is listed here:

| Where | What |
|---|---|
| `.claude/scripts/agent-claim.sh` | Claim-ledger helper (POSIX `sh`: Git Bash, macOS, Linux) |
| `AGENT_CLAIMS.md` | The ledger, created on the first claim. Machine-local: added to `.git/info/exclude`. |
| `.gitattributes` | `.claude/scripts/*.sh text eol=lf` |
| `CLAUDE.md` (root) | "Multi-agent resilience rules", between `<!-- resilient-agent-handoff rules: begin vN -->` / `end` anchors |

## 6. Usage

| Say to Claude | What happens |
|---|---|
| "run several agents in parallel on this repo" | Agents follow the claim / commit / PROGRESS discipline |
| "resume after the rate limit / crash" | The recreate-the-system procedure: orient, verify liveness, resume |
| "launch this long job so it survives" | Detached process, plus claim renewal if needed |
| "claim src/app.js" / "claim the GPU" | Uses the claim helper |

**Claim helper:**
```
sh .claude/scripts/agent-claim.sh claim   src/app.js  agent-a "refactor auth"   # exit 3 if someone else holds it
sh .claude/scripts/agent-claim.sh claim   DEVICE:gpu0 trainer "36-run sweep"
sh .claude/scripts/agent-claim.sh renew   DEVICE:gpu0 trainer                   # long jobs: every <90 min
sh .claude/scripts/agent-claim.sh list                                          # active claims + age
sh .claude/scripts/agent-claim.sh stale 90                                      # older than 90 min - VERIFY before releasing
sh .claude/scripts/agent-claim.sh release src/app.js  agent-a
```

## 7. Updating: auto-update

- **Version marker:** every copy of the skill carries a `<!-- skill-version: N -->` line.
- **What an invocation updates:** each time the skill runs, it brings every existing copy on the machine up to
  the highest version found. That covers the global install, a project install, the repo's
  `docs/handoff/resilient-agent-handoff-skill.md`, and a root-level `resilient-agent-handoff-skill.md`.
- **Generated files:** AUTO-SETUP rewrites its generated scripts and versioned `CLAUDE.md` sections, so they upgrade too.
- **Safety:** it never downgrades and never creates copies in new places.

**To roll out a new version:**
- **On one machine:** `run and install` the new installer file (or paste it), or put the new `SKILL.md` in any
  of those locations and invoke the skill.
- **Across machines:** commit the new `docs/handoff/resilient-agent-handoff-skill.md`. Every machine that pulls it updates the
  next time the skill runs.

## 8. Uninstall

```sh
rm -rf "<CLAUDE_HOME>/skills/resilient-agent-handoff"
# In a project: rm .claude/scripts/agent-claim.sh; delete the lines between the resilient-agent-handoff rules
# begin/end anchors in CLAUDE.md; remove AGENT_CLAIMS.md* lines from .git/info/exclude.
```

## 9. Paste block (full skill, v2)

~~~PASTE START~~~
Install or update the `resilient-agent-handoff` Claude Code skill (v3) on this machine. Do this:

1. Determine the Claude home directory: Windows Git-Bash "$USERPROFILE/.claude"; macOS/Linux "$HOME/.claude".
   If you are inside a git repository, check the live state with `git status` / `git log -1` (don't rely on
   a snapshot taken when the session started).
2. Find EVERY existing copy of the skill and read its `<!-- skill-version: N -->` line (no line = version 0):
   <CLAUDE_HOME>/skills/resilient-agent-handoff/SKILL.md, and inside a git repository also <repo-root>/.claude/skills/resilient-agent-handoff/SKILL.md,
   <repo-root>/docs/handoff/resilient-agent-handoff-skill.md, <repo-root>/handoff/resilient-agent-handoff-skill.md and <repo-root>/resilient-agent-handoff-skill.md.
   - ANY copy has a HIGHER version than v2: write nothing at all (not the skill, not the mirror). Tell me which
     copy is newer, then run the skill's AUTO-UPDATE + AUTO-SETUP from that newer copy instead. Never downgrade.
   - An installed SKILL.md (global or project) exists with the same or a lower version: overwrite THAT copy in
     place (keep its scope).
   - No installed SKILL.md exists: install it globally at <CLAUDE_HOME>/skills/resilient-agent-handoff/SKILL.md.
3. Write everything between the BEGIN/END markers below, verbatim and with LF line endings, to that SKILL.md
   (create the directory if needed).
4. If you are inside a git repository, write the same content to the tracked mirror
   <repo-root>/docs/handoff/resilient-agent-handoff-skill.md (use handoff/ instead if the repo has no docs/ dir but has handoff/).
   Commit it ONLY if `git status --porcelain -- <that path>` shows a change, with an explicit path
   (`git commit -m "..." -- <that path>`). Never `git add -A` / `git commit -a`: other work may be uncommitted.
   Don't commit this installer file itself.
5. Immediately run the skill's "AUTO-UPDATE + AUTO-SETUP" section (unless I said "install only"). Its scripts are
   the ```sh blocks in that section: write each to a temp file in your scratchpad and run it with `sh`. It is
   idempotent: it updates any other, older copies on this machine and activates the skill in the current repository.
6. Report the path(s) written, the version, what AUTO-SETUP detected and did, and whether the skill shows in your
   available-skills listing. A freshly written skill can take a moment (or a new session) to appear there; if it
   doesn't yet, confirm the file exists and its YAML frontmatter parses instead of reporting a failure.

--- BEGIN SKILL.md CONTENT (write everything between these lines, verbatim) ---
---
name: resilient-agent-handoff
description: Project-AGNOSTIC multi-agent resilience playbook (v3, self-activating) - survive interruptions (rate limits, crashes, context limits) with zero lost work and no collisions between concurrent agents. Commit every verified step with explicit paths; keep PROGRESS_<task>.md resumable-state files written BEFORE each step; coordinate agents through a mkdir-locked AGENT_CLAIMS.md ledger (bundled agent-claim.sh helper, liveness check before reaping); run long jobs as detached OS processes; keep an append-only achievement log; plus the recreate-the-system multiprompt for resuming after an interruption. AUTO-UPDATE + AUTO-SETUP run first on every invocation and right after install or re-install - they bring every installed copy of the skill up to the latest version, install/upgrade the claim helper, git-exclude the local ledger, and add or upgrade the standing rules in the project's CLAUDE.md -- v3 checks first whether the project already describes its own claim-ledger convention and refuses to silently write or overwrite a competing one if so. Invoke on "install the resilient-agent-handoff skill", "run several agents in parallel on this repo", "resume after a rate limit / crash / interruption", "claim a file or GPU", "launch a long-running job", "make this work interruption-proof", or when starting multi-agent work.
---
<!-- skill-version: 3 -->

# resilient-agent-handoff (v3) — portable, self-activating, self-updating multi-agent resilience playbook

**Project-agnostic.** Nothing here is specific to any one repo, language, or toolchain. Copy this
single file into `.claude/skills/resilient-agent-handoff/SKILL.md` in ANY project (or
`<CLAUDE_HOME>/skills/resilient-agent-handoff/SKILL.md` for every project on the machine) to get the
same interruption-resilient, multi-agent working discipline that was proven — twice, in one session —
to survive real Claude API rate-limit outages with zero data loss.

This skill composes with (does not replace) `session-handoff`: that skill exports/imports a
knowledge *snapshot* (state, decisions, next actions). This skill is about the *mechanics of
survival* while multiple agents (or one agent across many turns) do real work: how to not lose
progress when a turn dies unexpectedly, how to not collide with a sibling agent, and how to hand
the exact running system — not just a description of it — to whoever picks it up next.

## AUTO-UPDATE + AUTO-SETUP (run FIRST on every invocation, and immediately after this skill is installed or re-installed)

Idempotent — safe to re-run. Whoever installs, re-installs or starts this skill ends up on the
**latest version** with the practices below wired into the current repo, without any manual step.

0. **AUTO-UPDATE.** Run the update script below as `sh update.sh resilient-agent-handoff` (write it to a temp
   file in your scratchpad). It brings every existing copy of this skill — global install, project
   install, the repo's tracked `docs/handoff/resilient-agent-handoff-skill.md` mirror, a root-level
   copy — up to the highest `skill-version` among them. It never downgrades and never creates copies in
   new places.
   - If it reports a newer version than the file you are following, re-read that `SKILL.md` and follow
     it instead.
   - Commit any updated tracked mirror with explicit paths.
   - When installing over an existing copy, never replace a higher `skill-version` with a lower one.
1. **Not inside a git repo?** The practices still apply, but there is nothing to install. Report that
   and continue with whatever the user asked.
2. **Run the setup script below** from the repo root (write it to a temp file in your scratchpad, then
   `sh` it). It:
   - installs the claim-ledger helper `.claude/scripts/agent-claim.sh` (§3), always rewriting it so an
     older helper is upgraded, pinned to LF line endings;
   - adds the ledger `AGENT_CLAIMS.md` and its lock dir to `.git/info/exclude` — claims are live,
     machine-local coordination state, not project history;
   - manages the "Multi-agent resilience rules" section of the root `CLAUDE.md` (creating the file if
     missing) — "Applying this to a fresh project" step 4, automated, so future sessions follow the
     rules without invoking this skill by name. The section sits between
     `<!-- resilient-agent-handoff rules: begin vN -->` / `<!-- resilient-agent-handoff rules: end -->`
     anchors:
     - no anchors → the section is appended;
     - an older `vN` → the section between the anchors is replaced with this version's;
     - anchors not found exactly once each → the script stops with an error instead of guessing (§2).
3. **Adapt to the project** (only on the run that first adds the rules):
   - if `CLAUDE.md` was just created, add a one-line project description at the top (from the
     README/package manifest);
   - if the project already has its own progress-file, claim-ledger, or achievement-log convention,
     edit the appended rules to point at those instead (see "Applying this to a fresh project" step 1).
4. **Commit** what setup created or changed, with explicit paths (§1):
   `git add .claude/scripts/agent-claim.sh .gitattributes CLAUDE.md && git commit -m "..." -- .claude/scripts/agent-claim.sh .gitattributes CLAUDE.md`.
5. **Report** the update result and the setup script's one-line summary.

```sh
#!/bin/sh
# skill AUTO-UPDATE (shared by session-handoff, achievements-system, resilient-agent-handoff) - idempotent.
# Usage: sh update.sh <skill-name>. Brings every EXISTING copy of the skill on this machine/repo up to
# the highest skill-version found among them. Never downgrades, never creates copies in new places.
name=${1:?usage: update.sh <skill-name>}
if [ -n "${USERPROFILE:-}" ]; then home="$(printf '%s' "$USERPROFILE" | sed 's|\\|/|g')/.claude"; else home="$HOME/.claude"; fi
root=$(git rev-parse --show-toplevel 2>/dev/null || true)
ver() { v=$(tr -d '\r' < "$1" 2>/dev/null | sed -n 's/^<!-- skill-version: *\([0-9][0-9]*\) *-->$/\1/p' | head -1); echo "${v:-0}"; }
tmp=$(mktemp)
printf '%s\n' "$home/skills/$name/SKILL.md" > "$tmp"
if [ -n "$root" ]; then
  printf '%s\n' "$root/.claude/skills/$name/SKILL.md" "$root/docs/handoff/$name-skill.md" \
    "$root/handoff/$name-skill.md" "$root/$name-skill.md" >> "$tmp"
fi
best=; bestv=-1
while IFS= read -r f; do
  [ -f "$f" ] || continue
  v=$(ver "$f")
  if [ "$v" -gt "$bestv" ]; then best=$f; bestv=$v; fi
done < "$tmp"
if [ -z "$best" ]; then rm -f "$tmp"; echo "$name: no installed copy found"; exit 0; fi
updated=0
while IFS= read -r f; do
  { [ -f "$f" ] && [ "$f" != "$best" ]; } || continue
  if [ "$(ver "$f")" -lt "$bestv" ]; then
    tr -d '\r' < "$best" > "$f.new" && mv -f "$f.new" "$f" && updated=$((updated + 1)) && echo "$name: updated $f -> v$bestv"
  fi
done < "$tmp"
rm -f "$tmp"
echo "$name: latest is v$bestv at $best ($updated older copies updated)"
```

```sh
#!/bin/sh
# resilient-agent-handoff AUTO-SETUP (v3) — idempotent. Run from anywhere inside the repo.
set -e
root=$(git rev-parse --show-toplevel 2>/dev/null) || { echo "resilient-agent-handoff: not a git repo - practices apply, nothing to install"; exit 0; }
cd "$root"

mkdir -p .claude/scripts
cat > .claude/scripts/agent-claim.sh <<'CLAIM'
#!/bin/sh
# agent-claim.sh (resilient-agent-handoff v2) - mkdir-locked, append-only claim ledger. Never auto-reaps.
#   claim   <resource> <agent> [why] [notes]   claim a path or DEVICE:x (exit 3 if another agent holds it)
#   renew   <resource> <agent> [why]           refresh your own claim's timestamp (long-running jobs)
#   release <resource> <agent> [notes]         release your claim
#   list                                       active claims, with age in minutes
#   stale   [minutes]                          active claims older than N min (default 90) - VERIFY LIVENESS before releasing
# Ledger: $AGENT_CLAIMS_LEDGER or <repo-root>/AGENT_CLAIMS.md. Line: STATUS | ISO-8601 UTC | resource | agent | why | notes
set -eu
root=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
LEDGER=${AGENT_CLAIMS_LEDGER:-$root/AGENT_CLAIMS.md}
LOCK="$LEDGER.lock"

usage() { sed -n '2,8p' "$0" | sed 's/^# \{0,1\}//' >&2; exit 2; }
clean() { printf '%s' "$1" | tr '|\r\n' '/  '; }
to_epoch() { date -u -d "$1" +%s 2>/dev/null || date -u -j -f '%Y-%m-%dT%H:%M:%SZ' "$1" +%s; }

lock() {
  n=0
  until mkdir "$LOCK" 2>/dev/null; do
    n=$((n + 1))
    if [ "$n" -ge 150 ]; then
      echo "agent-claim: ledger locked for 30s ($LOCK). If no agent is writing right now: rmdir \"$LOCK\"" >&2
      exit 1
    fi
    sleep 0.2
  done
  trap 'rmdir "$LOCK" 2>/dev/null' EXIT
  trap 'exit 130' INT TERM
}

append() {
  [ -s "$LEDGER" ] || printf '# AGENT_CLAIMS ledger (resilient-agent-handoff) - append-only, write via agent-claim.sh\n# STATUS | ISO-8601 UTC | resource | agent | what/why | notes\n' > "$LEDGER"
  printf '%s | %s | %s | %s | %s | %s\n' "$1" "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$2" "$3" "$(clean "${4:-}")" "$(clean "${5:-}")" >> "$LEDGER"
}

active() {
  [ -f "$LEDGER" ] || return 0
  awk -F' [|] ' '
    /^#/ || NF < 4 { next }
    { k = $3 SUBSEP $4; if (!(k in st)) order[++n] = k; last[k] = $0; st[k] = $1 }
    END { for (i = 1; i <= n; i++) if (st[order[i]] == "CLAIMED") print last[order[i]] }' "$LEDGER"
}

cmd=${1:-}
[ $# -gt 0 ] && shift
case "$cmd" in
  claim|renew)
    [ $# -ge 2 ] || usage
    r=$(clean "$1"); a=$(clean "$2")
    lock
    held=$(active | awk -F' [|] ' -v r="$r" -v a="$a" '$3 == r && $4 != a')
    mine=$(active | awk -F' [|] ' -v r="$r" -v a="$a" '$3 == r && $4 == a')
    if [ "$cmd" = claim ] && [ -n "$held" ]; then
      echo "agent-claim: '$r' is already claimed:" >&2; echo "$held" >&2; exit 3
    fi
    if [ "$cmd" = renew ] && [ -z "$mine" ]; then
      echo "agent-claim: no active claim on '$r' by '$a' to renew" >&2; exit 4
    fi
    why=${3:-}
    if [ "$cmd" = renew ] && [ -z "$why" ]; then why=$(printf '%s\n' "$mine" | awk -F' [|] ' '{ print $5 }'); fi
    append CLAIMED "$r" "$a" "$why" "${4:-$cmd}"
    echo "CLAIMED $r by $a" ;;
  release)
    [ $# -ge 2 ] || usage
    r=$(clean "$1"); a=$(clean "$2")
    lock
    append RELEASED "$r" "$a" "" "${3:-}"
    echo "RELEASED $r by $a" ;;
  list)
    now=$(date -u +%s)
    active | while IFS= read -r line; do
      ts=$(printf '%s\n' "$line" | awk -F' [|] ' '{ print $2 }')
      e=$(to_epoch "$ts" 2>/dev/null || echo "$now")
      printf '%5sm  %s\n' $(( (now - e) / 60 )) "$line"
    done ;;
  stale)
    max=${1:-90}
    sh "$0" list | awk -v m="$max" '{ a = $1; sub(/m$/, "", a); if (a + 0 > m + 0) print }' ;;
  *) usage ;;
esac
CLAIM
chmod +x .claude/scripts/agent-claim.sh

# Pin to LF, or core.autocrlf=true checkouts turn the script CRLF and sh fails on '\r'
line='.claude/scripts/*.sh text eol=lf'
grep -qxF "$line" .gitattributes 2>/dev/null || printf '%s\n' "$line" >> .gitattributes

# The ledger is live, machine-local state: keep it out of git without touching the shared .gitignore
exclude=$(git rev-parse --git-path info/exclude)
mkdir -p "$(dirname "$exclude")"
for p in AGENT_CLAIMS.md AGENT_CLAIMS.md.lock/; do
  grep -qxF "$p" "$exclude" 2>/dev/null || printf '%s\n' "$p" >> "$exclude"
done

# CLAUDE.md standing rules, between versioned anchors so later versions can replace them in place.
# v3 adds a real pre-write check (see "existing_convention" below): v2 blindly appended/replaced
# this section, which on its very first real-world use silently created a SECOND, competing
# claim-ledger convention next to a project's own pre-existing one (found and hand-fixed after the
# fact). The check below makes that a load-bearing gate instead of a note the installing agent has
# to remember on their own.
RULES_VERSION=3
BEGIN='<!-- resilient-agent-handoff rules: begin'
END='<!-- resilient-agent-handoff rules: end -->'
blk=$(mktemp)
{ printf '%s v%s -->\n' "$BEGIN" "$RULES_VERSION"; cat <<'RULES'; printf '%s\n' "$END"; } > "$blk"
## Multi-agent resilience rules (resilient-agent-handoff)

Standing rules so an interrupted session (rate limit, crash, context limit) loses no work and
concurrent agents don't collide. Full rationale: the `resilient-agent-handoff` skill.

1. **Commit every verified step immediately, with explicit paths:**
   `git add <new-file> && git commit -m "..." -- <every path in this commit>`. Never a bare
   `git commit`, `git commit -a` or `git add -A` while another agent may be staging files.
2. **One `PROGRESS_<task>.md` per unit of in-flight work** with DONE (real commit hashes),
   IN PROGRESS, NEXT, HOW TO VERIFY (an exact command). Update it BEFORE starting each step, by
   point-edit or append — never a wide block replace.
3. **Claim shared files/resources before touching them** when other agents may be active:
   `sh .claude/scripts/agent-claim.sh claim <path|DEVICE:name> <agent> "<why>"`, then `release`
   when done; `list` / `stale` to inspect. The ledger is `AGENT_CLAIMS.md` (local, git-excluded).
   Never reap a stale claim by age alone — verify the process behind it is really gone, one claim
   at a time, never with a bulk cleanup.
4. **Run anything longer than a few minutes as a detached OS process** (`Start-Process` without
   `-Wait` on Windows; `nohup <cmd> >log 2>&1 & disown` or `setsid` on POSIX), confirm it is not a
   child of the agent's process, and `renew` its claim if it outlives the 90-minute staleness window.
5. **Record achievements append-only the moment they are verified** — in this project's achievement
   log if it has one (see its rule in this file), otherwise `docs/achievements/`.

After an interruption, resume with the skill's "recreate-the-system multiprompt": orient from
`git log`, the PROGRESS files and the ledger, verify liveness, then resume — never redo committed work.
RULES
nb=$(grep -cF "$BEGIN" CLAUDE.md 2>/dev/null || true); ne=$(grep -cF "$END" CLAUDE.md 2>/dev/null || true)
nb=${nb:-0}; ne=${ne:-0}
# Best-effort, not perfect: a real check beats none. Looks for signs CLAUDE.md already describes
# its OWN claim/lock/ledger convention, outside our own anchors, before we ever write anything.
existing_convention=""
if [ -f CLAUDE.md ]; then
  hit=$(grep -inE 'claim_file\.sh|FILES\.MD|claim.?ledger|file.?claim|resource.?lock' CLAUDE.md 2>/dev/null | grep -vF "$BEGIN" | head -1 || true)
  [ -n "$hit" ] && existing_convention="$hit"
fi
if [ "$nb" -eq 0 ] && [ "$ne" -eq 0 ]; then
  if grep -qF '## Multi-agent resilience rules (resilient-agent-handoff)' CLAUDE.md 2>/dev/null; then
    rules="LEGACY un-anchored section found - replace it by hand with the current block"
  elif [ -n "$existing_convention" ]; then
    rules="NOT auto-added: CLAUDE.md already appears to describe its own claim-ledger convention (found: ${existing_convention#*:}). Read that convention, then either point bullet 3 of the block at $blk to it, or add the block verbatim if it's genuinely unrelated - don't leave two competing conventions side by side."
  else
    if [ -f CLAUDE.md ]; then rules=added; else printf '# CLAUDE.md\n' > CLAUDE.md; rules="added (new CLAUDE.md - add a one-line project description)"; fi
    { printf '\n'; cat "$blk"; } >> CLAUDE.md
  fi
elif [ "$nb" -eq 1 ] && [ "$ne" -eq 1 ]; then
  if grep -qF "$BEGIN v$RULES_VERSION -->" CLAUDE.md; then
    rules="present (v$RULES_VERSION)"
  else
    # Don't blindly replace a section that's been hand-adapted (e.g. bullet 3 repointed at the
    # project's own ledger instead of AGENT_CLAIMS.md/agent-claim.sh) - only auto-replace when the
    # OLD section still matches the untouched stock default (both tell-tale identifiers present).
    old=$(awk -v b="$BEGIN" -v e="$END" 'index($0,b)==1{f=1;next} f&&index($0,e)==1{exit} f' CLAUDE.md)
    if printf '%s' "$old" | grep -qF 'AGENT_CLAIMS.md' && printf '%s' "$old" | grep -qF 'agent-claim.sh'; then
      awk -v f="$blk" -v b="$BEGIN" -v e="$END" '
        index($0, b) == 1 { while ((getline l < f) > 0) print l; skip = 1; next }
        skip && index($0, e) == 1 { skip = 0; next }
        !skip' CLAUDE.md > CLAUDE.md.new && mv -f CLAUDE.md.new CLAUDE.md
      rules="updated to v$RULES_VERSION"
    else
      rules="NOT auto-updated: the existing section looks hand-adapted (doesn't reference the stock AGENT_CLAIMS.md/agent-claim.sh), left alone rather than silently overwritten. New stock template for comparison, if useful: $blk"
    fi
  fi
else
  rm -f "$blk"
  echo "resilient-agent-handoff: rules anchors found begin=$nb end=$ne times in CLAUDE.md (expected exactly 1 each) - fix by hand" >&2
  exit 1
fi
rm -f "$blk" 2>/dev/null || true

sh .claude/scripts/agent-claim.sh list >/dev/null
echo "resilient-agent-handoff active: helper .claude/scripts/agent-claim.sh, ledger AGENT_CLAIMS.md (git-excluded), CLAUDE.md rules $rules"
```

## Why this exists (the proof, not just the theory)

In one real session on an unrelated project, this exact set of practices was tested twice:

1. A rate-limit outage killed 7 concurrent agents simultaneously. Recovery cost near-zero because
   every agent had written its own resumable state (§2) *before* its last action, and one agent
   had even pre-written a reverse-appliable patch of its uncommitted work before dying.
2. A second outage killed the orchestrating agent mid-analysis — but the actual 36-run compute job
   it had launched was a **detached OS process** (§4), so it finished anyway, unattended, and its
   results were still on disk when a fresh agent picked the work back up. Only the *orchestration*
   was lost, never the *work*.

Two real near-misses were also caught by these same practices: an over-eager cleanup operation
that nearly reaped another agent's still-live claim (caught by real liveness verification, §3.3),
and a partial file write that left zero-byte corruption in a shared ledger (caught by noticing the
file, not trusted implicitly).

## The five practices

### 1. Commit every verified step immediately, with explicit paths

Never batch multiple verified pieces of work into one deferred commit, and never use a bare
`git commit` / `git commit -a` / `git add -A` when any other agent might be concurrently staging
files — that combination has a real, documented failure mode: the git index is shared state, and
a bare commit can silently swallow another process's staged-but-uncommitted files under your
message. The safe form, even for a brand-new file the index doesn't know about yet:

```
git add <new-file> && git commit -q -m "..." -- <every explicit path this commit should contain>
```

The `-- <paths>` on the commit itself is what prevents the collision; the `add` merely makes a new
path known to git. Commit the moment something is verified — not at the end of a task, not when
convenient. A verified step that stays uncommitted for more than ~30 minutes or survives past one
logical unit of work exists only in the working tree, and a working tree is the first thing an
interruption takes with it.

### 2. A resumable-state file per unit of work, written BEFORE each step, not after

For every distinct piece of in-flight work (one per agent, or one per logical task if a single
agent is doing several), keep a small state file — call it `PROGRESS_<task>.md` or similar — with
four fields, always current:

- **DONE** — what's verified and committed, with the real commit hash.
- **IN PROGRESS** — the exact file and function/step currently being worked, and what's left in it.
- **NEXT** — what starts next, and what it depends on.
- **HOW TO VERIFY** — the exact command that proves the current state (a test, a gate script, a
  diff check) — not a description of what verification *would* look like.

**Write or update this file BEFORE starting a new step, not after finishing one.** If the step
gets interrupted, the file still describes the state truthfully — "IN PROGRESS: editing function
X, have not yet changed the call site" — rather than lying by omission about work that started but
never got recorded.

**Update by point-edit or append, never by "replace everything between two markers."** A wide,
loosely-anchored replace can silently delete sections another concurrent edit inserted between
your last read and your write. If you must replace a block, put an anchor comment around it that
occurs exactly once, and fail loudly (don't silently proceed) if that anchor isn't found exactly
once when you go to edit it again.

### 3. A lightweight, dependency-free file/resource claim ledger

When more than one agent might touch the same repo concurrently, coordinate with a plain-text
ledger any shell can write to — no database, no special tool required. A single file at the repo
root (e.g. `AGENT_CLAIMS.md`) with one line per event:

```
<STATUS> | <ISO-8601 timestamp> | <path-or-resource> | <agent-name> | <what/why> | <notes>
```

Where `<STATUS>` is a claimed/released marker (any two consistent tokens work) and
`<path-or-resource>` can be a real file path OR a pseudo-path for a scarce non-file resource this
project actually contends over (a GPU, all CPU cores, a specific port) — prefix those distinctly,
e.g. `DEVICE:gpu0`, so they're visually distinct from real paths.

AUTO-SETUP installs a ready implementation of exactly this ledger as `.claude/scripts/agent-claim.sh`
(`CLAIMED`/`RELEASED` tokens, UTC timestamps, `claim`/`renew`/`release`/`list`/`stale`). It follows
3.1–3.3 below and deliberately has **no reap command**: releasing someone else's claim is a manual,
per-claim decision made after checking liveness.

**3.1 — Locking the ledger itself.** The ledger is shared, so appends to it need a lock. `mkdir`
is atomic on essentially every real filesystem (NTFS, ext4, APFS, FAT32) — use a lock directory,
not a lock file (`touch`+`test -f` is NOT atomic and will race):

```sh
# acquire (retry briefly on failure — someone else is writing right now)
until mkdir "$LEDGER.lock" 2>/dev/null; do sleep 0.2; done
# ... append your line to $LEDGER ...
rmdir "$LEDGER.lock"
```

**3.2 — Staleness.** Every claim carries its own timestamp; a claim older than some threshold
(90 minutes is a reasonable default — long enough that a slow agent doesn't get evicted mid-task,
short enough that a dead agent doesn't block the project for hours) is presumptively reapable. But:

**3.3 — Never reap by elapsed time alone for anything that might be a real, still-running
process** (a GPU job, a long compute). Verify actual liveness first — a process list check, a GPU
utilization query, or equivalent for whatever the resource actually is — *before* treating a
timestamp as proof of death. This is not a hypothetical: an agent that skipped this step once
today reaped three other agents' genuinely-still-running claims in one sweep, because it used an
unscoped "clean up everything under this name" operation instead of checking each claim
individually. It was caught (the affected agents self-detected and restored their own claims), but
the lesson generalizes: **an unscoped bulk-cleanup operation is more dangerous than N individual,
verified ones. Prefer the individual, verified path even when it's more typing.**

**3.4 — The ledger only records intent, not enforcement.** Nothing stops a careless process from
ignoring it. The discipline is what protects the project; the ledger is what makes a violation
*visible* after the fact, which is most of the value — "no one can tell what happened" is a worse
failure mode than "something went wrong but we can see exactly what."

### 4. Long-running work survives as a detached OS process, not as a Claude Code turn

Anything that will take longer than a few minutes — a build, a benchmark sweep, a training run —
should be **launched as a process independent of the orchestrating session**, so that session's
own death (a rate limit, a context limit, a crash) costs only the orchestration layer, not the
actual work. The concrete requirement: the process's parent chain must not include the agent's own
process. On Windows, `Start-Process` in PowerShell with no `-Wait` accomplishes this; on POSIX,
`nohup <cmd> >log 2>&1 &` followed by `disown`, or a `setsid`-style double-fork, does the same.
**Verify detachment, don't assume it** — walk the launched process's actual parent-process chain
and confirm the orchestrating agent's process isn't in it.

If that detached process holds a claim from §3 for longer than the staleness window, give it its
own **claim-renewal** loop (a tiny script that re-touches the claim's timestamp every N minutes,
running under the same detached umbrella) — otherwise a perfectly-alive long job can have its
claim expire and get legitimately reaped by someone else mid-run. With the bundled helper:
`while kill -0 <job-pid> 2>/dev/null; do sh .claude/scripts/agent-claim.sh renew <resource> <agent>; sleep 1800; done`.

### 5. A permanent, append-only achievement log — separate from the resumable-state files

`PROGRESS_*.md` files (§2) are a *rolling* snapshot — each gets overwritten as work continues.
Separately, keep a **permanent record** of every substantive achievement (a real feature shipped,
a real bug found and fixed, a milestone verified) as its own timestamped, never-overwritten file —
e.g. `docs/achievements/YYYY-MM-DD_HHMMSS_short-slug.md` — containing:

1. What was achieved (the concrete deliverable).
2. How it was done (the real method, specific enough to redo).
3. What was found along the way (anything non-obvious).
4. What broke and how it was actually fixed (the real failure mode, not glossed over).
5. What was verified, with real numbers (not "should work" — what was actually run/measured).

Write this **the moment the achievement is verified**, not deferred to a later commit or the end
of a session. A long-running session may cross several genuine achievement moments before any
other kind of commit happens — each gets its own record, written and committed on its own.

## Applying this to a fresh project

AUTO-SETUP (above) performs steps 2 and 4 automatically; steps 1 and 3 remain judgment calls.

1. Check whether the project already has an equivalent of §2 (progress files), §3 (a claim
   ledger), and §5 (an achievement log). If it does — even under different names — **use its
   existing convention**, don't introduce a competing one.
2. If it has none: create `docs/achievements/` (or wherever this project's docs already live) and
   start using per-task `PROGRESS_*.md` files and the `AGENT_CLAIMS.md` ledger from §3 going
   forward. Don't retrofit history — start the discipline from the current moment.
3. For anything long-running, apply §4 regardless of whether the project has any of the above —
   it costs nothing and only pays off, never hurts.
4. If a `CLAUDE.md` or equivalent project-instructions file exists, these practices belong there
   as standing rules, not just in this skill file — copy the relevant sections in, in the
   project's own voice, so future sessions see them without needing to invoke this skill by name.

## The generic recreate-the-system multiprompt

Fill in the bracketed placeholders for the actual project, then use this after any interruption —
paste it into a fresh session. Unlike a plain "here's what happened" summary, this is meant to be
**executed**, not just read.

```
You are resuming in-progress work on [PROJECT NAME] after a session interruption. Your job is to
RECREATE the running system, not just describe it.

STEP 0 - make sure the resilient-agent-handoff skill is installed and run its AUTO-SETUP (idempotent).

STEP 1 - orient:
  git log --oneline -30
  cat [PROGRESS/HANDOFF FILE(S), if any]
  cat [ACHIEVEMENT LOG DIR]/*.md | tail  # the most recent few
  sh .claude/scripts/agent-claim.sh list   # or cat [CLAIM LEDGER FILE], if one exists

STEP 2 - before reaping or assuming anything is dead, verify real liveness:
  [process-list command for this OS] | grep -i "[known long-running process names]"
  [GPU-utilization query, if relevant to this project]
  Only treat a claim/lock as abandoned after confirming no real process backs it — elapsed time
  alone is not proof of death (see this skill's section 3.3).

STEP 3 - for each unit of in-flight work found in STEP 1, read its own state file first to check
whether it already finished. If not, resume it as the same logical task (reclaim its scope, don't
start over from scratch unless its state file says the work was lost). Apply the same discipline
going forward: commit every verified step immediately with explicit paths (section 1), keep the
state file current before each new step (section 2), write an achievement record the moment
something verifies (section 5), and launch anything long-running as a detached process with its
own claim-renewal if needed (section 4).

STEP 4 - refresh whatever knowledge-bundle/handoff document this project uses (or start one, using
the session-handoff skill's structure, if none exists) so the NEXT interruption costs as little as
this one did.

STEP 5 - report back: current state of the repo (commit hash or equivalent), which in-flight items
were already done vs. just resumed, and the single most important next action.

Do not re-do already-committed/verified work. Do not report progress you have not actually
verified against real file/process/repo state.
```

## Notes

- This skill is deliberately tool-agnostic: every mechanism above is expressible in plain shell
  against any real filesystem and any git repo. The bundled `agent-claim.sh` is a minimal POSIX `sh`
  implementation (Git Bash on Windows, macOS, Linux), not a dependency — a project with its own
  ledger tooling keeps using that.
- If the destination project already has a more elaborate version of any of these five practices
  (a dedicated locking tool, a structured achievement-tracking system, a CI-integrated progress
  dashboard), **defer to it** — this skill describes the minimum viable version of each practice,
  not a maximum. The goal is the underlying resilience property, not this specific implementation.
- Reuse this skill for any project, the same way `session-handoff` is reused — the specifics
  (which files, which claim names, which processes) are always discovered fresh from the project
  at hand, never hardcoded here.
--- END SKILL.md CONTENT ---
~~~PASTE END~~~

---
*Source:* <https://www.siriusart.bg/handoff-toolkit/INSTALL_RESILIENT_AGENT_HANDOFF.md> · skill only: <https://www.siriusart.bg/handoff-toolkit/resilient-agent-handoff-skill.md> · guide: <https://www.siriusart.bg/handoff-toolkit/>
*Keep in sync:* this file embeds `resilient-agent-handoff` v2. The skill's AUTO-UPDATE refreshes installed copies and repo
mirrors, but not this installer. Regenerate it from the latest `SKILL.md` when the skill's version changes.
