Skip to content

Errors and replay ​

A workflow talks to systems that sometimes fail: a mail provider that times out, an AI model that is rate-limited, a third-party API that refuses a request. Mankomail runs each workflow step by step and saves the result of every step, so a failure stops at the step where it happened and never silently loses what came before.

This page follows a failure from start to finish: what the step does on its own, what you decide per node, what happens to the run, how you are told, and how you replay it once fixed. For the statuses of a run, see Runs.

Temporary and permanent failures ​

Every failure of a step falls into one of two families, and that family decides what happens next.

  • Temporary: a failure that can succeed if tried again later — a network error, a provider briefly unavailable, an attempt that ran out of time. The step is tried again, after a delay.
  • Permanent: a failure that trying again would not fix — an invalid parameter, a missing or revoked connection, an address blocked by the network guard, a response too large. The step is not tried again.

A failure whose nature is unknown is treated as temporary: retrying a few times costs less than abandoning a step because of a passing network error.

A provider's rate limit is not a failure: the step is postponed until the time the provider gives (Retry-After), without using up an attempt.

Retries, at three levels ​

How many times a step is tried, and how long it waits in between, comes from the first of these three levels that sets it:

LevelWhereDefault
The nodethe node's Settings tab, Retries section—
The node typethe catalogAI nodes: 3 attempts, 10 s, 5 min per attempt
Mankomailbuilt in3 attempts, 5 s, 2 min per attempt

In the node's Settings tab, the Retries section shows the default that applies ("Default for this node type: …" or "Default: …"). Set for this node opens two fields:

  • Number of attempts: first attempt included, so 1 means no retry. Between 1 and 10.
  • Delay before the first retry: between 1 second and 1 hour. Later delays double, up to 15 minutes, with a little randomness so that many runs do not retry at the same instant.

A sentence under the fields says what will happen, for example "If the step fails: retry after 5 s, 10 s, then give up." During a test run, every delay is cut to 2 seconds: a test in the editor does not wait for a provider to calm down.

Only temporary failures are retried. A permanent failure ends the step at its first attempt.

The timeout of an attempt ​

The Timeout section of the same tab sets Timeout of one attempt, between 5 seconds and 30 minutes. When an attempt runs past it, it is interrupted for real, even if the node does not stop by itself, and the step fails with the code node_timeout. An interrupted attempt counts as a temporary failure: "not finished in time" is not "will never finish".

If the failure persists ​

Once a step has used up its attempts, or fails permanently, the If the failure persists section decides what happens. Three choices:

ChoiceWhat happens
Stop the runThe run stops as failed. It shows up in Activity › To do and starts the error workflow, if there is one.
Continue without this stepThe next nodes run anyway, without data from this step: its output path stays empty.
Follow the Error outputAn Error output appears on the node: link it to the steps to run on failure (alert someone, set the email aside).

The retries always come before this choice: a node set to continue or to follow the Error output is still retried first, and only a final failure takes that path. A step caught by Continue or by the Error output counts as succeeded: the run carries on, and is not a failed run.

Uncertain effects ​

Some nodes act on a third-party service: HTTP request, notifications, and the nodes of the integrations (Airtable, Notion, Google, Microsoft, Yousign, MyNotary…). If an attempt of such a node is cut off mid-call — its timeout passed, or the server restarted — there is no telling whether the service received the request.

Mankomail does not retry it blindly. The step fails with the code step_effect_uncertain, which is a permanent failure: the If the failure persists choice applies, and you decide. Check on the service whether the action happened, then, if it did not, resume the run from this step (see Replay a run). The Settings tab of these nodes says so in a note.

The other nodes are retried safely after an interruption:

  • nodes with no external effect (AI, conditions, transformations, sub-workflow calls, flow control): nothing can have happened outside Mankomail;
  • sending, moving or flagging an email, writing to a table, emitting a signal: Mankomail records these effects under a key that stays the same across attempts, so an effect that already went through is recognised, not repeated.

What the error of a run tells you ​

A failed step shows its error in the run, on the canvas and in the details of the node. The details list:

FieldContent
CodeA stable code, translated on screen (node_timeout, step_effect_uncertain, http_blocked, mail.mailbox_required…). See Error codes.
Server messageThe technical detail, to paste into a ticket. Never a secret.
Node and Node typeThe step that failed.
KindPermanent failure, or Temporary failure, attempts used up.
AttemptFor example "attempt 3/3".
Earlier attemptsEach attempt that failed before the last one, with its error, and interrupted when it was cut off.
Version that ran and TimestampWhich version of the workflow ran, and when.

The step also keeps the parameters it actually received, with expressions resolved and secrets masked: you see what the node was given without replaying it. While a step waits for its next attempt, it stays Running and says Retry scheduled.

Copy the error details copies all of this. The details may contain an email address or subject: read them before sharing.

The error workflow ​

A workflow can name another workflow that starts when one of its live runs fails. Set it in the workflow's Settings tab, On failure section, field Error workflow ("None" by default). The chosen workflow must be one of yours, different from this one, and start with the Workflow failure trigger. It does not have to be published to be chosen, but it must be published, and not paused, to start.

The error workflow receives:

  • the same triggering email as the failed run, when there was one: {{ email.… }} works, and you can reply in the thread or move the email;
  • a report under data.failure: workflowId, workflowName, executionId, workflowVersionId, failedAt, error (code, message, kind, attempts), nodeId, nodeName, nodeType, mailboxId, messageId.

It never starts for a test run, for an iteration of a loop (the loop fails in its parent run, once), for a run that was itself started by an error workflow (no chain of alerts), nor for itself. The failure itself is always recorded in Activity › To do, error workflow or not: the error workflow is an extra, not the alert channel.

Where failures show up ​

  • Activity › To do lists the failures you have not seen yet, together with the approvals waiting for you. Each card names the workflow, the step that failed ("Step “Sort the quotes”") and a short summary of the error, with Open the run, Replay and Mark as seen. Marking as seen fixes nothing and erases nothing: the run stays failed, it merely leaves the banner.
  • The Runs tab of the editor lists the runs of the workflow. Filter them by status, by kind (Test runs, Live runs) and with Search an email, which searches the subject and the sender of the triggering email. Open a run to see it on the canvas.
  • The home dashboard shows the recent failures and the workflows that fail the most.

Replay a run ​

Replay runs a failed or cancelled run again, for real, with the same trigger data: the same email, the same webhook body. A test run is never replayed: it is run again from the editor. The confirmation asks two questions.

Where to start from

ChoiceWhat runsEffects
From the start (default)the whole workflowevery live effect already produced happens again: an email that left is sent again
From the failed steponly what had not gone through; steps that succeeded are kept as they wereno effect already produced happens again

On which version

ChoiceWhich graph
The published version (default)the one carrying your fixes: the usual choice after a repair
The original versionthe one that ran, to reproduce the failure

Before you confirm, the dialog lists the live effects this run already produced and, for each one, whether it will happen again or will not happen again. When the failed step is open on the canvas, Resume from this step is the same replay, started from the failed step.

From the failed step, steps are kept by node: on the published version, a node that still exists keeps its original output even if you changed its parameters; a new node runs. A step caught by Continue or by the Error output in the original run counts as succeeded and is kept; to run it again, replay from the start.

A replay is refused when the workflow is archived, when the mailbox of the original run is disconnected or gone, when the run is still in progress or succeeded, when the requested version is not available (no published version, or the original version was deleted), and when the sender of the email is now excluded from your scope. The new run is linked to the original: it shows Replay, with Open the original run.

Replay every failure ​

After a fix, Replay all failures in the Runs tab replays at once the failed live runs of the workflow that have not been replayed yet, oldest first, over 24 hours, 7 days or 30 days, with the same two choices. Runs that can no longer be replayed are left out and counted. Test runs are never included.

Cancel a run ​

Cancel run stops a run that is queued, running or waiting: no further step starts, its waits are closed, and the runs it called (sub-workflows, loop iterations) are cancelled too. Steps already done are not undone: an email that left stays sent, and a step calling a service right now may still go through. Keep running closes the dialog without cancelling.

Unpublishing, archiving or pausing a workflow does not cancel its runs in progress.

Through the API, POST /api/v1/workflows/{id}/executions/cancel cancels every run in progress of a workflow (up to 200 per call, call again while hasMore is true), and POST /api/v1/executions/cancel cancels a list of runs.

Limits of a workflow ​

The Runs section of the workflow's Settings tab holds two limits. They apply right away, without publishing again.

Simultaneous runs. How many live runs of this workflow can be in progress at once. The next ones wait their turn, Queued, in order of arrival: nothing is lost. Test runs, runs started by another workflow, loop iterations and runs that are waiting (for an approval, a delay or a signal) do not count. Empty: the instance default (WORKFLOW_MAX_CONCURRENCY, 10). The instance also caps all workflows together (EXECUTIONS_MAX_CONCURRENT, 100).

Maximum run duration. Only working time counts: waiting for an approval, a delay or a signal uses none of it. Beyond it, the run is stopped and fails with the code execution_timeout, up to two minutes late; the steps still in progress are closed, and the error workflow starts if there is one. At least one minute. Empty: the instance default (EXECUTION_TIMEOUT_MS, one hour), capped by EXECUTION_TIMEOUT_MAX_MS (24 hours). A step calling a service at that moment is not interrupted. See Environment variables.

What happens at most once ​

EffectWithin a run (retries, restarts)Replay from the startReplay from the failed step
Email: send, draft, move, flagat most oncedone againnot done again
Table writeat most oncedone againnot done again
Signalat most oncedone againnot done again
AI model callrepeated without consequence, except its costdone againnot done again
HTTP, notification, integrationsat least once when the node reports its failure; at most once after an interruption (step_effect_uncertain)done againnot done again
Sub-workflow, loop iterationstarted oncestarted againnot started again if the calling step had succeeded