Data Processor
The Data Processor step runs a data processor job as part of a workflow — custom Python, Node.js, or Spark code, against files in the Intely File System.
Typical use: retrieve a large file from a vendor, run a Spark job to aggregate it, then load
the aggregated result onward.
Configuration
| Field | Notes |
|---|---|
| Job | The data processor job to run |
| Synchronous | Whether the integration waits for the job to finish. Leave it on when a later step needs the job's output; turn it off for work the workflow does not depend on |
| Input files | Each of the job's input slots is bound to a file or directory — a fixed path in the Intely File System, or a file produced by an earlier step |
| Arguments | Values for the job's arguments. A value argument takes a value; a switch argument is set as a boolean |
Compute resources are allocated by the platform — there is nothing to size or select here.
Output
The job's output files are available to later steps, and are also written into IFS at
system/temp/job-execution-<executionId>/. The execution log is written alongside them.
Notes
- Data processor runs are the slowest step type in most workflows. Where the job doesn't need
to block the rest of the workflow, turn Synchronous off — or consider whether the work
belongs in a separate scheduled integration instead. - Running asynchronously means there is no output for later steps to use. The job's own
execution history is where its outcome lives.
Updated about 1 month ago
Did this page help you?
