The immortal life of Pi (Running the Pi coding agent on Temporal)

AUTHORS
Moe Abadi
PUBLISHED
Oct 08, 2026
DURATION
10 MIN
  • AI/ML
  • Durable Execution
  • Architecture

We built pi-temporal at Temporal to let the Pi coding agent recover when a machine fails mid-command. It runs on a fleet of Temporal Workers, processes that execute work across multiple machines. You keep your Pi sessions and tools, including the Pi Terminal UI. If a machine fails during a turn, another Worker picks up the turn without automatically repeating the interrupted tool call.

The tiger is chaos, the saving grace that is the boat is Temporal, and Pi is just trying to finish the coding task (note the sad fate of worker-1).

That handover applies to tasks run by a Worker and started with /background or the CLI. A turn in the terminal UI runs in your pi process. If that process dies, reopen pi to recover the interrupted turn. Here’s how we got there.

The Earendil team released Pi 1.0, alongside Pi Durable, their framework for agents that outlive a terminal session. Reading it, we kept nodding. Tasks checkpoint before they move on. Submissions carry a request ID, so a duplicate returns the existing submission instead of starting another one. An “effect sandwich” records the intent before running the external effect, then commits the outcome. Memos record a value once, so a retry reuses it instead of recomputing it.

If you work at Temporal, that list reads like our own docs. One Hacker News commenter put it better than we could: tools like this are “the natural evolution of playing around building temporal like things.” Agent harnesses keep arriving at Durable Execution by solving the failures they encounter.

There’s one difference that matters to us: who takes over when the host fails. Pi Durable’s built-in persistent backends have one process owning storage at a time, with no cross-process locking. Its default SQLite configuration survives process crashes, but the README warns that the newest commits “may be lost on power or host failure.” Recovering on another machine requires access to the surviving storage and a new process to reopen it. Pi Durable supports custom storage backends. Automatic failover across our Worker fleet is what we wanted Temporal to provide.

So we tried an experiment: keep Pi’s model of the world and let Temporal supply the execution engine.

The split between Temporal and Pi#

Pi keeps the conversation in its session file, and you keep using its Terminal UI. Temporal drives execution. A turn spans the work from a user prompt to the agent’s final response. It can contain multiple steps, each consisting of one model call followed by any tool calls the model requests. A session is a Workflow, and each unit of work in a turn is an Activity:

runModelCall appends the model’s answer to the transcript. runToolCall runs one tool call and reports its result. sealStep records the step’s tool results in the order the model requested them and reports whether the turn is done.

The Workflow loops around these Activities. It counts steps and schedules the next Activity within the configured budgets. It doesn’t hold the conversation state. Each Activity reads the transcript to determine its work, and sealStep tells the Workflow whether to loop again.

That is the main rule: the transcript is the conversation state. Each Activity reconstructs the turn from the session file before doing its work. The outcome is recorded before the turn advances, so recovery does not depend on a Worker’s memory. We assume the process will die.

Another Worker can pick up a turn because it has access to the conversation and project files. The session file lives in a shared directory, so every Worker reads the same record. The files edited by the tools travel too, shipped the same way Git ships history.

Those files remain ordinary files in each Worker’s project directory. After a step, the Worker snapshots that directory and writes a Git bundle into <session>.jsonl.tree/, beside the session file in shared storage. Temporal does not store the project files, and the project’s own .git is untouched.

Before starting the next step, a Worker applies any bundles it missed and checks out the latest tree. A machine joining the session for the first time gets both the conversation and the saved project state from shared storage.

Temporal drives the execution loop. Pi’s session file holds the conversation, and Git bundles preserve project snapshots.

The chaos demo#

The repo has a chaos demo. One command starts a Temporal dev server with three Workers in Docker and submits a multi-step coding task. Every 15 to 40 seconds, the demo kills whichever Worker is doing the work. The task is submitted once and never resubmitted.

21:25:18  running: runToolCall attempt 1 on worker-1
21:25:31  chaos: killed worker-1 while running: runToolCall attempt 1
21:25:36  chaos: worker-1 is back, as a fresh container
21:25:49  running: runModelCall attempt 1 on worker-2
          | tool result: The outcome of this tool call is unknown.
            The session stopped after the call started but before the
            result was recorded. It can have taken effect. Check the
            current state before you try again.
          | tool: bash
          | tool result: ls: cannot access '/project/*.txt': No such file

The Temporal UI shows two interrupted tool calls as orange bars, each followed by a fresh step on another Worker.

In the runs we’ve tested, the turn completes and the requested file is on disk with the right contents.

Unknown outcomes#

When a Worker dies during a shell command, the session file may contain no result even though the command took effect. We cannot assume an arbitrary command has an idempotency key: a git push might have succeeded before its result was recorded, while a file write might never have started.

The executor records the interrupted call as “outcome unknown” instead of automatically running it again. It tells the model to check the current state before deciding whether to issue a fresh call. The model inspects the world. If the work is there, it moves on. If it is not, it redoes the step.

That is what the ls: cannot access line shows: the file was absent from the project directory the recovering Worker inspected, so the model redid the work. This check establishes absence at that moment; it does not establish exactly how far the interrupted command got. Checking contents requires reading the file back. The executor’s guarantee is narrower: it does not automatically repeat an interrupted call whose outcome is unknown, and it reports that uncertainty to the model.

The mechanism behind “unknown” is a claim on disk. Before invoking a tool, the dispatch writes a claim beside the session file, under <session>.jsonl.pending/<turn>/<step>/. When the command returns, its result is stored next to the claim.

A retry that finds the result returns it without running the tool again. If it finds only the claim, the tool may have started, but its outcome was not recorded. The retry reports that uncertainty. Temporal’s attempt counter cannot answer this question because an Activity attempt can fail before it reaches the tool.

Model calls work differently. A retry reads the transcript first. If the failed attempt’s answer is already there, it returns that answer without calling the model. If no answer was saved, it sends the prompt again. That can cost more tokens and produce a different answer, but generating a response does not itself execute a tool.

Temporal can therefore retry model-call Activities on another Worker. A retry that finds a saved answer needs no additional model request. In the demo’s Temporal UI, model calls reach attempt 2 on another machine, while interrupted tool calls close with an unknown outcome. Recovery proceeds through a fresh call after the model checks the current state.

There’s one more failure worth talking about: the attempt that did not die. Temporal can stop waiting for a Worker without proving that the Worker stopped. A timed-out attempt may still be running and may produce effects later.

We close the abandoned step and record that closure in shared storage. Every append to the session file carries the Worker’s lease, which the write guard checks before accepting the append. When an old attempt resumes and tries to publish, the guard rejects its append, even if its bytes match the accepted result.

Because a step’s tools share a working directory, a started step stays on one Worker. Its tool calls are pinned to the Worker that ran its model call. Losing that host closes the step; the turn continues on another Worker.

The abandoned step closes with an unknown outcome. The write guard rejects the stale attempt’s late append to the session file.

The rules that make recovery work#

Temporal lets the Workflow continue on another Worker after a crash, provided the required infrastructure remains available. A timeout does not prove the previous attempt stopped: overlapping attempts may still try to write to the session file or invoke a tool. Our integration follows these rules to handle that overlap:

  • The transcript is the only conversation state. A Worker reconstructs the turn from the session file. We test this by killing the sealStep Activity after its writes land, then retrying it in a fresh process. The hooks at the step boundary fire once.

  • sealStep is the only writer of a step’s tool results. Pi’s session file is a tree, so concurrent appends could create separate branches. Writing the results through sealStep preserves their order.

  • A claim is recorded before a tool is invoked. A competing attempt that finds the claim does not invoke the same tool call again.

  • A Worker that lost its lease cannot append to the session file. We tested this by freezing a Worker mid-call and letting another take over the session. When we resumed the first Worker, the write guard rejected its late append. The lease requires exclusive file creation and coherent reads in the shared directory. Its timing windows account for clock differences, but do not establish a maximum clock skew. Rejecting the append prevents the transcript write; the tool may still produce external effects. The claim on disk prevents automatic re-execution of the same tool call.

  • The operator configures token and elapsed-time budgets per turn and per session. The executor enforces those budgets, including a hard deadline for the turn.

Conclusion#

Pi Durable got the programming model right. We wanted the Pi coding agent to recover across machines while preserving its session format. Temporal supplies the execution loop, while shared storage makes the saved conversation and project state available to the next Worker.

Cloudflare also integrates Pi Durable with Durable Objects. Its Agents SDK provides a PiHarness integration, with lifecycle support that keeps work running across restarts. Our experiment keeps the existing Pi coding agent and its shell tools, with execution distributed across Temporal Workers on infrastructure we choose.

Everything here was a retrofit through a small fork of Pi. If you’re building a new agent, the experimental Temporal Agent Harness offers a starting point built around Workflows. It includes recovery within a turn and tool approvals that can pause for a human decision. Its event stream exposes the agent’s execution. What we learned here feeds into it.

You can try pi-temporal like this:

git clone https://github.com/temporalio/pi-temporal
cd pi-temporal
./install.sh

Run scripts/run-pi.sh to start Pi with Temporal driving execution.

Run demo/run.sh to try the chaos demo. It starts three Workers and interrupts them while the agent works through the task.

PI-TEMPORAL

Ready to break something?

Clone pi-temporal, run the demo, and watch three Workers get killed mid-task.

Build invincible applications

It sounds like magic, we promise it's not.