An agent can fail between changing the world and recording the result. A remote service may accept a request just before the agent's process dies. When the process returns, it has to decide whether the operation happened. Repeating it may create a second effect; skipping it may leave the task unfinished. The problem is older than AI, but agent workflows make it harder to hide.
Earendil's experimental Pi Durable framework, released alongside Pi 1.0 on October 1, puts recovery into the machinery around a model. Its documentation describes checkpointed tasks and persistent transcripts. Interrupted tool calls rerun only when declared safe to replay; otherwise the model is informed of the interruption. A request identifier prevents duplicate submission within the harness. Earendil also provides an example in which an external payment operation uses an idempotency key. These are documented design choices, not a universal guarantee about every integration.
That last detail carries the story. A durable conversation cannot, by itself, make an external action exactly once. The system performing the action needs a way to recognize a retry, return an existing receipt, or expose enough state for reconciliation. Without that cooperation, recovery can restore the agent's memory while leaving the actual outcome uncertain.
Consider a hypothetical maintenance agent opening a ticket. The ticket service creates it, but the network connection drops before the receipt reaches the agent. If the agent merely repeats the request, a duplicate may appear. A persistent log of its intention is useful, but intention does not establish completion. The workflow needs a stable operation identifier or a way to search for the result it already caused.
Our reading is that agent reliability should be evaluated at these interruption points. A successful uninterrupted demonstration says little about the moment after a write is accepted and before acknowledgment is stored. Tests that interrupt model requests, tool execution, and result recording can show whether the system preserves the difference between an unfinished plan and an uncertain outcome.
The replay decision also has a human dimension. Declaring a tool safe to retry is a promise about its behavior, not a convenience setting. A search often tolerates repetition. A tool that changes permissions or sends a message may not. Some operations can be designed to return the same result for the same key. Others need an explicit reconciliation step. Treating them all as ordinary retries would erase the distinction that recovery depends on.
Persistent agents raise another issue: what survives when an operator believes the work has stopped? Earendil documents a distinction between foreground and background tasks. In any application using that distinction, the interface must make the continuing work understandable. A person should be able to tell whether they stopped the current exchange, cancelled a job, or ended every activity owned by the agent.
The framework remains experimental, and a documented recovery model is a starting point for testing rather than proof of production behavior. Its importance lies in moving attention toward the surrounding system. Models can produce competent plans while the application loses receipts, repeats actions, or cannot explain which work remains active.
Long-running agents need more than a longer memory. They need a reliable account of what was intended, what was attempted, what was confirmed, and what is still unknown after failure. Recovery becomes useful when it preserves those distinctions instead of making the agent confidently continue from the wrong one.
Recovery behavior is described by the developers and was not independently tested. Harness submission deduplication is not presented as an exactly-once guarantee for arbitrary external effects.

