Skip to content
PacSpace
Talk to us

Safety and idempotency

Retries that never record twice, keys and fingerprints the agent can make again, what to retry and what not to, and the two states a write can be left in.

An agent retries. Networks drop, processes restart, a step runs twice. The rules on this page make every retry safe: the same action recorded once, however many times the write is sent.

The three rules

  1. Every write carries an idempotencyKey, and the key, and the body, can be made again from the action that produced them.
  2. Retries are bounded and only for errors that can pass on their own.
  3. The agent records an action after it happened, once, and reads the outcome from the history or the webhook; the write's answer says only that the entry was queued.

Idempotency

One key is one write. The API keeps the key with a fingerprint of the body and answers a repeat by the rule:

A repeat withThe API answers
The same key and the same body, after the first write finishedThe original answer, with "idempotent": true. Nothing is written twice.
The same key while the first write is still running202 IDEMPOTENCY_PENDING with retryAfterSeconds. The SDKs return it rather than throwing. Wait, then send the same request again.
The same key and a different body409 IDEMPOTENCY_CONFLICT. A different body is a different write; give it its own key.

A write without a key is keyed by a fingerprint of the request, so an exact resend is still safe; but a key of your own is what lets you retry from a fresh process after a crash, so always set one.

Making the key

Build it from what the agent already holds: the run and the step.

typescript
const idempotencyKey = `run-${runId}:call-${callId}`;

Three constraints. It must be possible to make again from what the agent has on the retry, so never from Date.now() and never from a value the agent cannot recompute. It must be unique per action, so two actions never share one. And it is never the record id alone, because every entry of the record needs its own key. The record id is run-${runId}; the entry's key is that plus the step.

Keeping the body stable

A retry sends the same body. Serialize deterministically, and do not stamp the body with the time of the attempt. If the agent needs to record something different, that is a new entry with a new key.

The fingerprints are part of the body. A fingerprint is blinded by default, so fingerprinting the same file again gives a different fingerprint, and a retry that fingerprints afresh is a different body: the answer is 409 IDEMPOTENCY_CONFLICT. Keep the ref from the first attempt, or fingerprint again with the file's blinding file, which gives the same ref:

typescript
import { createReadStream } from 'node:fs';
import { fingerprint, readBlindingFile, writeBlindingFile } from '@pacspace-io/sdk';

// The first call writes the blinding file beside the file; every later call uses it, so a retry sends the same fingerprint.
async function refOf(path: string) {
  const kept = await readBlindingFile(`${path}.pacspace.json`).catch(() => undefined);
  const { ref, blinding } = await fingerprint(createReadStream(path), { blinding: kept });
  if (!kept && blinding) await writeBlindingFile(path, blinding);
  return ref;
}

What to retry

AnswerRetry?Why
A timeout, a dropped connectionYes, same key, same bodyThe write may or may not have arrived; the key makes the resend safe.
202 IDEMPOTENCY_PENDINGYes, same key, after retryAfterSecondsThe first write is still running.
429 RATE_LIMIT_EXCEEDEDYes, same key, after retryAfterSecondsThe workspace is at its limit of 120 writes a minute.
408, 500, 502, 503, 504Yes, same keySomething on our side that passes.
400, 422 with a record codeNoThe body breaks a rule; the sentence names it. Fix the entry and send it with a new key.
401, 403NoThe key is wrong or the workspace is not a records workspace. Alert the operator.
402 PLAN_LIMIT_REACHED or PAID_ENTITLEMENT_REQUIREDNoThe plan is at its limit, or a paid plan's payment is not settled. Alert the operator; the same key sends the entry once the plan is fixed.
409 IDEMPOTENCY_CONFLICTNoThe body differs from the first use of this key. Check that the retry reused its fingerprints; a new entry takes a new key.
409 RECORD_LAYOUT_MISMATCHNoThe record was started under another record type. Use a new record id.

The SDKs retry timeouts, dropped connections, and 408, 429, 500, 502, 503, and 504 answers themselves: two retries by default, waiting the seconds a 429 or a 503 names, and a growing wait otherwise. A 202 IDEMPOTENCY_PENDING comes back to you as the write's answer, with code and retryAfterSeconds; wait that long and send again. Bound your own loop as well: a maximum number of attempts and a maximum elapsed time, then hand the failure to a person. Never retry forever, and never retry a 400.

The two states a write can be left in

Queued but not yet committed. The write answered QUEUED. Committed follows, and the harness learns it from record.committed at its webhook, which lists each entry's idempotency key in records[].referenceId, or from status: committed in the history. Do not treat QUEUED as done for anything that matters; treat committed as done. When an agent must wait for a person's approval before it acts, wait for the approval's entry to commit, so the record shows the approval first.

Queued and then failed. A queued entry can fail to commit, for example a record whose first entry was sent as closed. The harness learns it from record.failed, which carries the code and a plain sentence, or from status: failed in the history. Nothing was written and the earlier entries are unchanged. Fix what the sentence names and send the entry again with a new key.

Both are why the harness reads the history or listens for the webhook rather than trusting the write's answer. records.history is the first thing to read after a write, and it is safe to poll.

Webhook handlers

Answer 2xx as soon as the signature checks, and do the work after. Deliveries are at least once, so persist X-Event-ID and skip an id you have seen. A slow handler times out at ten seconds, is retried, and hands you the same event again. See Webhooks.

Sandbox and Production

A pk_test_ key writes to your Sandbox and a pk_live_ key to Production, on the same host; the key decides. Keep the two apart: one key for each environment, never shared, and never copied from one environment's settings into the other's.

Failure modes

  • A key built from the clock, so a retry becomes a second entry.
  • A retry that fingerprints the file afresh, so the body differs and the answer is IDEMPOTENCY_CONFLICT.
  • The record id used as the idempotency key, so the second entry of a record is refused as a repeat of the first.
  • Retrying a 400 or 401, which multiplies failures without fixing anything.
  • Treating QUEUED as committed.
  • A webhook handler that does the work before it answers, times out, and is delivered again.