Skip to content

Extract ​

Extracts the fields you declare from the email (number, date, amount…) into the working data. A missing field is null.

Extract reads the email and fills the fields you declare: an order number, a delivery date, an amount, a yes/no answer. Each field becomes a key of the working data, ready for a condition, a table write or a message downstream.

Choose Categorize when you need a routing decision, and Free instruction (AI) in JSON mode when you need a free-form structure (nested objects, lists). To extract from a PDF attachment, place Read attachments before this node and point Content to process at the extracted text.

At a glance ​

  • Type: ai.extract · version 1
  • Category: AI
  • Kind: Step — one stage of a run
  • Effect: No external effect (none) — nothing is written outside Mankomail; safe to replay
  • Needs a carrier email: No
  • Connection: None
  • Inputs: main
  • Outputs: main
  • Service ports: model (llm.model, optional)

Parameters ​

fields ​

Fields to extract — Each field becomes a key of the working data: {{ data.extract_1.order_number }}.

  • Type: List of items (collection)
  • Required: Yes
  • 1 to 15 items
  • Each item has:
    • name — Identifier. The key name in the working data. Short, ideally without spaces.
      • Type: Text (string)
      • Required: Yes
      • 80 characters at most
      • Example: order_number
      • Expressions: {{ }} not accepted
    • description — Description. What the model should look for. The single most useful instruction of this node. Write it in any language: the model matches it to the email by meaning, not by words.
      • Type: Long text (text)
      • Required: No
      • Default: "" (empty)
      • 500 characters at most
    • type — Type
      • Type: One choice (options)
      • Required: Yes
      • Default: string
      • Options:
        • string — Text
        • number — Number
        • boolean — Yes / no
        • date — Date (YYYY-MM-DD)

inputTemplate ​

Content to process — What the model sees of the email. Subject and text body by default.

  • Type: Long text (text)
  • Required: Yes
  • 10000 characters at most
  • Shown under “Advanced” in the editor
  • Expressions: {{ }} accepted
    Default value
    text
    {{ email.subject }}
    
    {{ email.bodyText }}

Outputs ​

  • main

Data produced ​

What this node adds to the run data, and how to read it in an expression. <step> stands for the step key: the node name turned into an identifier (see Data and expressions).

  • {{ data.<step>.<field> }} — string, number, boolean or null. One key per declared field, named exactly after the field Identifier. The value is converted to the field type (Text, Number, Yes / no, Date). null when the field was not found or could not be converted.
  • {{ data.<step>._missing }} — array of string. The identifiers of every field left at null, in declaration order. Empty when every field was found.

Example ​

Supplier invoices arrive by email. Add an Extract node named Invoice with these fields:

fields:
  - name: number
    description: The invoice number, as printed.
    type: string
  - name: amount
    description: The total amount including tax.
    type: number
  - name: due_date
    description: The payment due date.
    type: date
  - name: is_reminder
    description: Whether this email is a payment reminder rather than a first invoice.
    type: boolean

For an email that says "Please find invoice F-2026-118 for €1,250.50, due on 15 November 2026", the step data reads:

json
{
  "number": "F-2026-118",
  "amount": 1250.5,
  "due_date": "2026-11-15",
  "is_reminder": false,
  "_missing": []
}

Downstream, {{ data.invoice.amount }} gives 1250.5. A Condition (If) on {{ data.invoice._missing }} lets you send incomplete invoices to a human.

How values are converted ​

Every field is declared to the model as nullable and required: the model is told that null is a correct answer for a field that is absent or uncertain, rather than guessing. The node then converts each value strictly:

  • Text: kept as returned.
  • Number: spaces and the €, $ and £ signs are removed; 1 250,50 and 1,250.50 both become 1250.5. A value that still is not a number becomes null.
  • Yes / no: true, yes, oui, 1 give true; false, no, non, 0 give false. Anything else becomes null.
  • Date: the model is asked for YYYY-MM-DD. A date returned in another format (for example 12/03/2026) is kept as is, not rejected.

Every field left at null is listed in _missing.

How the model is chosen ​

The node has a model service port. Leave it empty and the call uses the default model the administrator set for the Extract use on the Connections page (Artificial intelligence section) (or the instance default). Link a provider node such as Anthropic (Claude), OpenAI (GPT) or Ollama (local) to the model port to run this node on that provider.

Prompt and safety ​

The instructions given to the model are fixed by the node. The email content, as rendered by Content to process, never enters the system message: it is sent in a separate user message, inside a delimited block, with an explicit instruction to treat it as data and ignore any instruction it contains. The field list and descriptions are sent on the user side too, since descriptions are templatable. Descriptions can be written in any language; the model matches them to the email by meaning.

Tips ​

  • Field identifiers. Keep them short, without spaces or accents (order_number): expression paths only accept letters, digits, _ and $, so {{ data.<step>.order number }} cannot be written.
  • Reserved name. A field named _missing is refused, as are duplicate or empty identifiers. At most 15 fields are read.
  • Descriptions matter. The description is what the model looks for. Say what to take and what to ignore ("the total including tax, not the subtotal").
  • Long emails. The content sent to the model is cut at 20,000 characters, and the model is told the text was truncated.
  • Test runs. In test runs the model is really called, so you see real values. If the instance has no usable AI provider, the output is fabricated (placeholder values) and the run detail says the AI was skipped. Unlike the other AI nodes, Extract does not write a simulated key: its data only holds your fields and _missing.
  • Errors. No usable field gives node_invalid_param; empty content to process gives node_nothing_to_do. A model answer that is not a JSON object gives llm_invalid_output, which is retried. See Error handling.