Skip to content

Read attachments ​

Extracts the text of the email attachments (PDF, text, HTML…) into the working data: {{ data.extract_text_1.attachments.0.text }}. Place it before an AI node to hand it a document’s content.

Read attachments extracts the text of attachments and puts it in the working data, where the next nodes can use it. It is the node to place before an AI node when the useful content is in a document rather than in the email body: an invoice PDF, a contract, an exported CSV.

The attachments are those of the run: the attachments of the triggering email first, then the files added by earlier steps (a document fetched from a drive, a signed PDF). A run without an email can therefore read files added by its own steps.

At a glance ​

  • Type: attachment.extract_text · version 1
  • Category: Data
  • Kind: Step — one stage of a run
  • Effect: No external effect (none) — nothing is written outside Mankomail; safe to replay
  • Needs a carrier email: No
  • Connection: None
  • Inputs: main
  • Outputs: main

Parameters ​

selection ​

Attachments to use

  • Type: One choice (options)
  • Required: Yes
  • Default: all
  • Options:
    • all — All attachments: Those of the triggering email, then the files added by earlier steps (a fetched deed, a signed PDF), in that order.
    • first — The first one only: The first attachment, when only one matters.
    • byMime — By file type: By MIME type: application/pdf, or a whole family with image/*.
    • byName — By file name: By name pattern: *.pdf, invoice-*.

mime ​

File type — Exact MIME type (application/pdf) or a whole family (image/*). Left empty, no attachment is kept.

  • Type: Text (string)
  • Required: No
  • Default: "" (empty)
  • 200 characters at most
  • Example: application/pdf
  • Shown when: selection is byMime
  • Expressions: {{ }} accepted

namePattern ​

File name — Simple pattern: * matches anything, ? one character (*.pdf, invoice-*). Case is ignored.

  • Type: Text (string)
  • Required: No
  • Default: "" (empty)
  • 200 characters at most
  • Example: *.pdf
  • Shown when: selection is byName
  • Expressions: {{ }} accepted

includeInline ​

Include inline images — Signature logos and body images are ignored by default — they are not attachments in the member’s sense.

  • Type: Yes / no (boolean)
  • Required: No
  • Default: false
  • Shown under “Advanced” in the editor

maxChars ​

Characters read per attachment — Beyond this the text is cut and flagged as truncated. A 200-page contract has no business sitting in the execution data.

  • Type: Number (number)
  • Required: No
  • Default: 20000
  • Whole number, from 1000 to 100000
  • Shown under “Advanced” in the editor

onMissing ​

If no attachment can be read — Applies when no attachment matches the filter, and when an attachment’s content no longer exists (purged).

  • Type: One choice (options)
  • Required: Yes
  • Default: skip
  • Options:
    • skip — Carry on without it: The workflow carries on with an empty list. The step succeeds and the attachment lands in "skipped".
    • fail — Fail the step: When reading the attachment is the whole point: a visible failure beats a workflow running on nothing.

onUnsupported ​

If the format cannot be read — Formats with no extractor (archives, images without OCR, CAD files) and oversized attachments.

  • Type: One choice (options)
  • Required: Yes
  • Default: skip
  • Options:
    • skip — Carry on without it: The other attachments are read as usual. The step succeeds and the attachment lands in "skipped".
    • fail — Fail the step: When reading the attachment is the whole point: a visible failure beats a workflow running on nothing.

Outputs ​

  • main

Data produced ​

What this node adds to the run data, and how to read it in an expression. <step> stands for the step key: the node name turned into an identifier (see Data and expressions).

  • {{ data.<step>.attachments }} — array of { position, filename, mime, size, text, truncated, pages?, extractor }. The attachments read, in order. Empty when nothing could be read.
  • {{ data.<step>.attachments.0.text }} — string. The text of the first attachment read, cut at Characters read per attachment.
  • {{ data.<step>.attachments.0.filename }} — string. The file name of the first attachment read. mime and size (in bytes) give its type and size, and position its position in the run's attachment list.
  • {{ data.<step>.attachments.0.truncated }} — boolean. true when the text was cut at Characters read per attachment.
  • {{ data.<step>.attachments.0.pages }} — number. For a PDF, the number of pages actually read. Absent for other formats.
  • {{ data.<step>.attachments.0.extractor }} — string. The reader that handled the file: text, html or pdf.
  • {{ data.<step>.count }} — number. The number of attachments read (the length of attachments).
  • {{ data.<step>.skipped }} — array of { position, filename, reason }. The selected attachments that were set aside, with the reason: unsupported_type, too_large, too_many, not_found or unavailable.
  • {{ data.<step>.summary }} — string. A one-line description for the run detail (files read, pages, truncation, files ignored), in the language of the member running the workflow. summaryKey and summaryParams hold the same sentence as a translation key and its parameters.

Example ​

Supplier invoices arrive as PDF attachments. Add a Read attachments node named Files:

selection: byMime
mime: "application/pdf"
maxChars: 20000
onMissing: skip
onUnsupported: skip

Then an Extract node whose Content to process is {{ data.files.attachments.0.text }}. For an email carrying one PDF and a signature logo, the step data reads:

json
{
  "attachments": [
    {
      "position": 1,
      "filename": "F-2026-118.pdf",
      "mime": "application/pdf",
      "size": 84211,
      "text": "INVOICE F-2026-118 …",
      "truncated": false,
      "pages": 2,
      "extractor": "pdf"
    }
  ],
  "count": 1,
  "skipped": [],
  "summary": "…"
}

The logo is an inline image: it is ignored without appearing in skipped.

Supported formats and limits ​

  • Text: text/plain, text/csv, text/tab-separated-values and any other text/* type.
  • HTML: text/html and application/xhtml+xml, read as text.
  • PDF: application/pdf and application/x-pdf. At most the first 50 pages are read; the following pages are ignored without setting truncated, so compare pages with the length of the document when it matters.

Any other type (archives, images, office documents, CAD files) is set aside as unsupported_type without its content being loaded. There is no OCR. A file that cannot be parsed (an encrypted or damaged PDF, for instance) or whose reading exceeds the time budget is also set aside as unsupported_type.

Files above 10 MB are set aside as too_large before being loaded. Each extraction has a 20-second time budget. Each text is cut at Characters read per attachment (20,000 by default, between 1,000 and 100,000) and flagged as truncated. At most 10 attachments are read per step; the next ones are set aside as too_many.

When an attachment cannot be read ​

By default an unreadable attachment does not fail the workflow: an email often carries a useful PDF next to something that cannot be read. Each set-aside attachment is listed in skipped with its reason, and the step succeeds.

  • If no attachment can be read governs the case where nothing matches the selection, and attachments whose content no longer exists (not_found, unavailable).
  • If the format cannot be read governs unsupported_type and too_large.

Set either to Fail the step when reading the document is the whole point of the step: a visible failure beats a workflow running on empty text.

Tips ​

  • Selection is closed by default. By file type with an empty type, or By file name with an empty pattern, selects nothing. Types accept application/pdf, image/* or image/; names accept a simple pattern with * and ? (invoice-*.pdf), case-insensitive.
  • Inline images. Signature logos and images embedded in the body are ignored unless Include inline images is ticked.
  • Extracted text is third-party content. It is written by whoever sent the file. When you insert it into an AI node's instructions, the model is told that inserted values are data; delimiter tags used by the AI nodes are neutralised in the extracted text. Prefer Content to process over the instructions when the node has one.
  • Test runs. Reading is not an effect: in test runs the attachments are really read, so the next nodes see the real text.
  • Errors. With If no attachment can be read set to Fail the step and nothing selected, the step fails with node_nothing_to_do; with an attachment set aside under a Fail the step setting, it fails with the matching code: attachment.unsupported_type, attachment.too_large, attachment.not_found or attachment.unavailable. A failure of the attachment service itself is relayed as is, and retried when it is transient.