English
Runs
A run is one pass of one workflow version, started by one event: an email that matched, a scheduled time, a webhook call, a call from another workflow, a change at an integration, a manual run or a test run. It records everything the run did, node by node, so that you can answer "what did the automation do with this email, and why?".
A run is made of steps: one step per node that the run reaches. Every step is written to the database before and after its node runs. Nothing is kept only in memory, which is what lets a run survive a restart, wait for days, and be inspected afterwards.
How a run starts
Each trigger creates runs in its own way, and each one protects against duplicates:
| Trigger | One run per | Duplicates |
|---|---|---|
| Email received | matching email and published version | The same email delivered twice (for example by a push notification and a periodic check) creates a single run. Publishing a new version allows an email already seen by the previous version to be processed again. |
| Schedule | planned occurrence | An occurrence fires once, even with several servers. |
| Webhook received | request received | Two calls are two events: nothing is merged. For an integration that sends a delivery identifier, a delivery sent again by the provider is recognized and absorbed. |
| Integration change (polling, for example Airtable) | changed item | An item already processed is not processed again after a crash. |
| Called by a workflow | call from the Call a workflow node | One child run per calling step, even if that step is replayed. |
| Manual run | click | No deduplication: running a workflow twice on the same email is a deliberate gesture and produces two runs. |
| Test from the editor | test | Each test run is a new, simulated run (see Test runs). |
An email that matches several workflows starts one run per workflow, or only some of them, depending on the "When several workflows match" setting of the email trigger (see Triggers).
A run uses the workflow version that was published when it was created, and keeps a reference to it. Editing the draft or publishing again never changes a run that has already started. Only the branches of the trigger that started the run are executed: a workflow that also listens to emails does not run its email branch because a scheduled time arrived.
Run statuses
| Status | Label | Meaning |
|---|---|---|
queued | Queued | Created, no step has started yet. |
running | Running | At least one step is ready or in progress. |
waiting | Waiting | Nothing is running: the remaining steps wait for an event (a date, an approval, a called workflow, loop iterations). |
succeeded | Succeeded | Every branch reached its end without a definitive failure. |
failed | Failed | At least one step failed definitively. |
cancelled | Cancelled | A signal ended the run while it was waiting. |
succeeded, failed and cancelled are final: nothing is scheduled after them.
Step statuses
| Status | Label | Meaning |
|---|---|---|
queued | Queued | The node is ready to run and waits for a worker. |
running | Running | The node is running, or failed temporarily and will be retried. |
waiting | Waiting | The node asked to wait (see Waiting and resuming). No process is busy with it. |
succeeded | Succeeded | The node finished and chose an output. |
failed | Failed | The node failed definitively. |
skipped | Skipped | The node was not on the path taken by this run. |
The run view also shows two variants of a succeeded step:
- Frozen (not executed): in a test run, the node has a pinned output. It did not run; its pinned output fed the next nodes.
- Passed through (disabled): the node is disabled. It did not run, called nothing, and the data passed straight through to the next nodes.
Branches and skipped steps
After each step, Mankomail decides which nodes can run next. The rule:
A node runs when all its incoming connections are settled and at least one of them is live. A connection is settled when the node it comes from has finished (succeeded, failed or skipped). It is live when that node succeeded and chose the output the connection starts from.
Consequences:
- A Condition (If) that takes
truemakes the nodes behindfalseskipped, and everything that only depends on them is skipped in turn. - A node fed by two exclusive branches (for example the
trueoutput of a condition and a category of Categorize) runs as soon as one of them arrives. It does not wait for both. - A node that returns no output (it ends its branch without error, such as Categorize set to stop when no category fits) settles its connections without making them live.
- Nodes in the for each body of a Loop appear as skipped in the run that contains the loop: they ran in the iterations, which are separate runs.
- When several nodes are ready at the same time, they run in parallel.
When a step fails
A step that fails definitively makes the run end as failed. Other branches already running go to their end first; steps still waiting for an event on another branch are then given up, so that nobody approves a sending that belongs to a failed run.
Each node can be set to continue or to follow an error branch instead of stopping the run. These settings, the error codes and how to replay a failed run are described in Errors and retries.
Automatic retries
Errors are classified in two kinds:
- Temporary (network failure, provider unavailable, timeout): the step is retried automatically.
- Permanent (invalid setting, missing mailbox, refused by the provider for good): the step fails immediately. Retrying would fail the same way.
An error that does not say which kind it is is treated as temporary.
By default, a step gets 5 attempts. The wait before each new attempt doubles, starting at 2 seconds, with a random variation of ±20% so that many runs hitting the same incident do not retry at the same instant:
| After the failure of attempt | Wait before the next one |
|---|---|
| 1 | about 2 s |
| 2 | about 4 s |
| 3 | about 8 s |
| 4 | about 16 s |
| 5 | none: the step fails definitively |
While it is being retried, the step stays Running and its attempt counter increases (shown as "attempt 2", "attempt 3"… in the run view). When a provider answers that a quota is reached, the step is postponed until the provider says it can be called again, without using up an attempt.
How retries combine with the failure settings of a node is described in Errors and retries.
External effects are never doubled
A step can be run more than once: after a retry, after a worker crash, or after a restart. Mankomail makes this harmless:
- A step that is already settled is never run again.
- Each step gets a stable idempotency key, derived from the run and the node, the same for every attempt.
- Emails leave through an outgoing queue keyed on that step: a replayed step finds its sending already recorded and does not send a second email.
- Nodes that call an external service pass the key to that service when it supports one. When the provider has no such mechanism, the node page says so and, when possible, offers a safeguard (for example the Deduplication column of Sheets — append a row). With the generic HTTP request node, prefer idempotent endpoints for writes.
A manual run or a replay is a new run with new keys: running the same email twice on purpose does send twice.
Waiting and resuming
Four nodes can make a step wait without blocking anything:
| Node | Waits for |
|---|---|
| Wait | a date or a delay, possibly cut short by a reply in the thread or a signal |
| Approval | a person's decision (see Review and approvals) |
| Call a workflow | the end of the called workflow |
| Loop | the end of all its iterations |
The node does not sleep. It records what it is waiting for, the step becomes Waiting, and no process is busy with it. When nothing else is running, the run becomes Waiting too. When the event arrives (the date is reached, someone approves, the child run ends), the step is resumed, takes the output that matches the event, and the run continues. A waiting run costs nothing while it waits, whether it waits for a minute or for a year.
A long wait is capped by the instance (730 days by default, WAIT_MAX_DAYS, see Environment variables).
Cancelling a waiting run. The Emit a signal node, with "What the signal does" set to Cancel the waiting runs, ends every run waiting on the same correlation key. They become Cancelled and nothing behind the wait runs: the reminder is never sent.
Durability
Every state change is written to the database before the work it describes is done:
- the steps to run are recorded, then the jobs that run them are queued;
- a worker that dies in the middle of a step loses its claim after a while, and another worker takes the step over;
- a step left ready by a crash is found again and run;
- waits rely on durable scheduled entries, not on timers: a redeployment or a restart loses nothing, even during a one-year wait;
- the data produced by each step is stored with the step.
Restarting or updating the server therefore never loses a run. The worst case is a step run a second time, which the idempotency key above makes harmless.
Step data
For each step, the run records:
- its status, its attempt number, its start and end times and its duration;
- the output it took (the port name), or none if the branch ended there;
- the data it produced: the values that later nodes read as
data.<node>.…(see Data and expressions). Most nodes also add asummary, a readable sentence such as what was found or sent, written in the language of the member who owns the workflow; - its effects: what the node did outside (an email sent, a file uploaded, a model called), each marked as real or simulated;
- the error if it failed: a code, a technical message and the node concerned.
Data up to 32 KiB is returned with the step. Larger data is stored separately: later nodes still read it in full, but the run view shows "This node’s output was too large to be returned with the step."
Pausing a workflow and stopping sends
Two switches stop activity without deleting anything:
- Pause on a published workflow: no new run starts; the ones in flight finish, and waiting steps keep waiting. Resume starts new runs again immediately. A draft cannot be paused, since it does not run.
- Stop sending in Admin → Sending: real sends (replies, forwards, new messages) are held back for the whole organisation, while the rest keeps running. On Resume sending, what was held goes out; nothing is lost, nothing is sent twice. An instance can also be stopped from its configuration (
SEND_ENABLED); the organisation switch cannot reopen it. See Governance.
Looking at runs
- Runs tab of a workflow, next to Editor: the list of its runs, newest first, with their status and their trigger (Scheduled, Webhook, Called, Manual…). Filter by status ("All statuses") and by kind ("Live and test runs", "Test runs", "Live"). Pick a run to inspect it on the canvas: each node shows its status. Open full screen opens the detailed view.
- Run page: the summary (Simulated, or "Real — effects have been sent"), the Requested / Started / Finished times, the Related email with Open thread, then the Steps, each with its status, attempt, effects and Show data. A loop step lists its Iterations ("12 iterations · 11 succeeded · 1 failed") and opens each one. Produced files lists the documents the run fetched or built. A replay links to the original run.
- Why? on an email in the webmail: the workflows that ran on this email and their runs.
- Activity page: To do lists the decisions your workflows are waiting on and the failed runs nobody has looked at yet; Waiting lists what is on hold; Received shows the runs of the last 24 hours. A banner reports the failed runs of the last 24 hours until you acknowledge them.
Retention
Finished real runs, with their steps and their data, are deleted 180 days after their creation. Test runs are deleted after 7 days by default, configurable with SIMULATED_EXECUTIONS_RETENTION_DAYS (1 to 365). Real runs that are still running or waiting are never deleted by retention.