Skip to main content

How Workflows Run

This page explains what happens behind a run: where a draft run executes, how a Loop's results are collected, how a loop's items are spread over threads and pods, how data is cached between nodes, and how nodes record metrics you can read and compare afterwards.

Your worker pool for draft runs​

A draft run is a run of a workflow that is not deployed: a Run from the builder, or an execute call through the REST API, STAN or the SDK while the workflow is still a draft. Draft runs execute on your own worker pool. A deployed workflow never uses a pool; its runs execute on the workflow's own deployment (see Deploying Workflows).

How the pool behaves​

  • One pool per user. Your pool runs only your draft runs, as you. Other users in your organization have their own pools.
  • It starts when you open the builder. Opening a batch workflow in the builder starts your pool with 2 pre-warmed workers, so a Run starts at once instead of waiting for a new worker.
  • Each run uses a worker once. A run takes a free worker, and the worker is removed when the run ends. While you are still in the builder, the pool is topped back up so the next Run also finds a worker waiting.
  • A run that needs more gets a worker made for it. A pre-warmed worker has up to 1 vCPU and 1 GB of memory and reaches only the platform's own services. A run whose nodes need more (set on a node's Resources tab), or that connects to data sources, add-ons or other services, gets a worker started for it with exactly that size and access. The run waits for that worker to start, which can take up to a minute when new compute has to be added.
  • It stops when you leave. The builder keeps your pool alive while you stay in the workflow section (it checks in every minute). About 5 minutes after you leave the workflow section or close the browser tab, the pool stops. A run already in progress finishes; it does not restart the pool.

A draft run started when your pool is stopped (for example through the REST API) still runs: a worker is started for that run.

What the builder shows​

The pool's state is shown as a badge at the top right of the builder, next to Back to Workflow List:

BadgeMeaning
Connecting...The builder is joining your pool
Starting pool...Your pool's workers are starting
N workers readyN workers are running for you, including any started for a run
Stopping pool... / Pool stoppedYour pool is shutting down, or has stopped
Workers unavailableThe workers could not be started; hover the badge for the reason
Pool unavailableThe builder could not reach the pool service; hover the badge for the reason

Streaming workflows show Starting streaming worker... / Streaming worker ready instead: a streaming test session runs on its own test worker (see Streaming Workflows).

Limits​

  • A workflow runs one execution at a time unless its concurrency limit is raised (up to 10). A second Run while one is pending or running is refused with "already has 1 in-flight execution(s)"; wait for it to finish or cancel it.
  • Pool workers are billed to you, as the user whose pool they belong to. FinOps shows them under your user.

How to check it worked​

Open the builder and wait for N workers ready, then click Run: the canvas shows each node's state as the run moves through the graph, and the run appears in Workflow Monitor with its trace.

Closing a loop with a Loop Accumulator​

A Loop runs the nodes on its body once per item. The Loop Accumulator node (node id merge, shown in the palette as Loop Accumulator (Sink)) closes the loop: it collects the output of every iteration into one array, and the rest of the workflow continues from it.

How to wire a Loop and its accumulator, and the Map variant, is described in Creating Workflows: Loop bodies. In short:

  1. Draw the body from the Loop's Loop Body handle.
  2. Connect the body's last node to a Loop Accumulator, and connect the Loop's Completed handle to the same accumulator.
  3. Continue the workflow from the accumulator's output. The next node reads the collected array at data.data (and the item count at data.count), for example an input mapping items = data.data.

The accumulator collects every item in both kinds of loop run: one item at a time, and spread over pods with Pod Scaling (see below). The Loop's own Completed output carries only the counts (completed, totalIterations, failedItems), never the items, so always read results from the accumulator.

What happens without one​

  • Deploy refuses the workflow: loop "..." has no terminating Merge node, with the suggestion to add one after the body.
  • A draft run fails before any node runs: 'loop' requires an aggregator node to collect results.
  • If an accumulator is present but nothing from the body is wired into it, deploy refuses the workflow because the accumulator would collect nothing.

How to check it worked​

Run the workflow and click the Loop Accumulator (or the node after it) on the canvas: its output holds one entry per item, in item order, and count equals the number of items. In the run's trace, the body nodes have one span per item.

Threads and pods​

A loop's items can run one at a time, several at once in one pod (threads), or spread over several pods (pods). You choose on the node's Scaling tab; the settings and their ranges are listed in Creating Workflows: Scaling a Loop or Map.

ThreadsPods
NodeMap (always), and each worker pod of a scaled LoopLoop with Pod Scaling (Distributed) on
Where items runInside one pod, several at the same timeOver the workflow's own pod plus worker pods started for the loop
Best forI/O-bound work: API calls, model calls, database queriesCPU- or memory-heavy work, long work per item, thousands of items

How a scaled Loop runs:

  1. The loop's items are put on a shared work queue.
  2. The number of pods is the item count divided by Target Items per Pod, rounded up, at most Max Pods and never more pods than items. The workflow's own pod is one of them, so the run starts that many minus one worker pods.
  3. The workflow's own pod starts on the items at once; worker pods join as soon as they are running and take items from the queue. A small loop can finish on the workflow's own pod before the worker pods are up.
  4. An item that fails on a transient error, or whose worker pod stops, is run again, up to Retry Attempts.
  5. When every item is done, the Loop Accumulator collects all of their results, as for a loop that runs one item at a time.

Pod Scaling works in draft runs and in deployed runs alike. A scaled loop inside another scaled loop is allowed up to 5 levels deep; deeper nesting fails the run. While and goal loops always run one iteration at a time, because each iteration depends on the previous one.

How to check it worked​

Open the run's trace in Workflow Monitor. A scaled Loop has a span of type DISTRIBUTED named Distributed: (loop name), recording the total item count and the Max Pods, Target Items per Pod and threads it ran with. Each body node span records the item index and the pod that processed it.

Caching between nodes​

Nodes hand data to each other through a cache with three tiers. Every item written to the cache goes to object storage, and smaller items are also kept closer to the node for fast reads:

TierWhereWhat it holdsLifetime
L1: MemoryThe worker pod's memoryItems up to 100 KB. The tier uses up to 10% of the pod's free memory (between 50 MB and 500 MB)5 minutes
L2: DiskThe worker pod's local diskItems up to 10 MB (an L1 item is written here too). Up to 500 MB in total1 hour
L3: Object storageThe platform's storage bucketEvery item, of any size. Shared by every pod of the runUntil the run's cache is cleaned up (below)

A node reading an item checks memory first, then disk, then object storage. An item read from disk again and again is moved up to memory. When a pod runs low on memory, new items are kept in memory only if they are small, and the oldest are dropped first; nothing is lost, because every item is also in object storage.

Because L3 is shared, every pod of a run sees the same data: the worker pods of a scaled Loop read what the workflow's own pod wrote, and a resumed run reads what the first attempt wrote.

Cleanup: a run's cache is deleted 24 hours after the run ends, except for the most recent finished run of each workflow, which keeps its cache so it can still be resumed. A run's final result is never removed by this cleanup.

There is nothing to configure: the tier is chosen for each item by its size. The memory tier grows with the pod's free memory, so a run whose nodes have more memory (their Resources tab, or the deploy dialog for a deployed workflow) keeps more in memory.

Node metrics​

Many built-in nodes record metrics as they run: numbers such as a judge's score, the rows a query affected, or the tokens an LLM call used. Metrics are stored with the run, and you can read them on the run's trace and compare them across runs.

Which nodes log metrics​

Node typeExample metrics
Evaluation nodes (LLM as Judge, Faithfulness Checker, RAG Metrics, Relevance Grader, Answer Quality, Pairwise Comparator)evaluation_score, evaluation_passed, faithfulness_score, relevance_score, quality_score
LLM, Model Registryprompt_tokens, completion_tokens, total_tokens, response_time_ms
Database sources that write (Snowflake, BigQuery, SQL Server, Oracle and others)rows_affected
Transform nodes (Filter, Sort, Limit, Dedupe, Aggregate, Set Fields, Value Map, PDF Redactor and others)matched_count, items_sorted, items_limited, duplicates_removed, pdf_pages_redacted
Split, Stop and Erroritems_split, workflow_stopped
CodeWhatever its code records with metric(name, value, unit) (see below)

A node inside a loop records its metrics once per item.

Log your own metrics from a Code node​

A Code node can record any number you want to see per run. Call metric(name, value, unit) in its Python:

redacted = 0
for item in items:
for field in ("ssn", "phone"):
if item.get(field):
item[field] = "[REDACTED]"
redacted += 1
metric("values redacted", redacted, "count")
result = items
  • name is the metric's name, a non-empty string.
  • value is a number (an integer or a decimal; True/False are not accepted).
  • unit is optional text such as count, score, percent or ms.

Each call becomes a span-level metric with the Code node as its source, and the trace-level average, minimum and maximum include it. In Run Once for Each Item mode every item's calls are recorded. A wrong argument stops the node with an error that points at the line. Metrics are recorded when the code finishes; code that fails records none.

Metric levels​

Each metric has a level:

LevelWhat it is
SpanRecorded by one node run. Its Source is the node's name
TraceComputed when the run ends: for every metric name, the average, minimum and maximum over all of the run's span metrics (avg_<name>, min_<name>, max_<name>). Its source is shown as Aggregated
ExecutionA summary metric for the whole run

Read and compare metrics​

  1. Open Workflows > Monitor, click the workflow, and open a run's trace.
  2. The Execution Metrics card lists every metric of the run with its Metric Name, Value, Unit, Level and Source. Filter by level (All Levels, Span, Trace, Execution), search by metric or source, and sort by name, value, level or source.
  3. To compare runs, select two or more runs in the workflow's execution history and click Compare. The comparison lines up each metric across the selected runs.

A run whose nodes recorded no metrics shows No metrics recorded for this execution.

How to check it worked​

Run a workflow that contains one of the nodes above, for example a Limit node: its trace's Execution Metrics card shows items_limited at span level with the node as its source, and avg_items_limited, min_items_limited and max_items_limited at trace level.

For a Code node, add metric("values redacted", 4, "count") to its code and run the workflow: the trace's Execution Metrics card shows values redacted at span level with the Code node as its source, and avg_values redacted, min_values redacted and max_values redacted at trace level.

Next Steps​