From 1505bf0b7e58d3cd1950a582a22b6dd0178b8b32 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Aur=C3=A9lien=20Sibiril?= <81782+aureliensibiril@users.noreply.github.com> Date: Mon, 27 Apr 2026 10:03:52 +0200 Subject: [PATCH] Refine agent cancellation guideline MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Lead with the observable contract (ctx.Done = graceful suspend, return is *SuspendedError, framework shields downstream calls) and mention agent.ErrSuspendForCheckpoint as the recommended cancel cause for graceful-stop intent. Drop the leak of the WithoutCancel mechanism — readers need the contract, not the strategy. Signed-off-by: Aurélien Sibiril <81782+aureliensibiril@users.noreply.github.com> --- contrib/claude/agent.md | 31 +++++++++++++++++++------------ 1 file changed, 19 insertions(+), 12 deletions(-) diff --git a/contrib/claude/agent.md b/contrib/claude/agent.md index 7685154ad..a1e26fae2 100644 --- a/contrib/claude/agent.md +++ b/contrib/claude/agent.md @@ -43,25 +43,32 @@ type Tool interface { ## Cancellation semantics -`ctx.Done()` is a **graceful-suspend signal**, not a hard abort. When the -caller cancels `ctx`, `Run`/`RunStreamed`/`Resume`/`Restore` finish their -in-flight LLM call and tool, persist a checkpoint via the configured -`Checkpointer`, and return `*SuspendedError`. Internally `coreLoop` -shadows the incoming ctx with `context.WithoutCancel(ctx)` and uses the -shadow for every downstream call so the cancel never kills work -in-progress; only the at-boundary check observes the original ctx. +`ctx.Done()` is a **graceful-suspend signal**, not a hard abort. When +the caller cancels `ctx`, `Run`/`RunStreamed`/`Resume`/`Restore` let +the in-flight LLM call and tool finish, persist a checkpoint via the +configured `Checkpointer`, and return `*SuspendedError`. The framework +shields downstream calls (LLM, tools, hooks, guardrails, save) from +the cancellation so they complete naturally; the cancel is only +observed at the next safe boundary. + +Use `agent.ErrSuspendForCheckpoint` as the cancel cause when the +intent is graceful suspend — supervisors that distinguish a +graceful-stop request from infrastructure-level causes (lease loss, +heartbeat failure) inspect `context.Cause(ctx)` to dispatch. Implications: -- A `context.WithTimeout` becomes a "max wall-clock budget then suspend" - — strictly better than today's "deadline kills work outright." +- A `context.WithTimeout` becomes a "max wall-clock budget then + suspend" — strictly better than the alternative where the deadline + kills work outright. - There is no in-process hard-abort path. Callers that genuinely need to kill a run terminate the process; stale recovery handles the row. -- Tools that need their own deadline must derive it themselves - (`context.WithTimeout(ctx, ...)` *inside* the tool body). +- Tools receive a non-cancellable ctx; if a tool needs a hard deadline + it must derive its own with `context.WithTimeout(ctx, ...)` inside + the tool body. The supervisor (`pkg/probo/agent_run_handler.go`) maps a SIGTERM-driven -shutdown broadcast onto a per-run `cancelRun(ErrSuspendForCheckpoint)`, +shutdown broadcast onto a per-run `cancelRun(agent.ErrSuspendForCheckpoint)`, so the same contract drives both the public Go API and the worker infrastructure path.