Skip to main content

Testing Workflows

Test your workflows before deployment to ensure they work correctly and handle edge cases. Workflow testing creates real executions with results tracked per node.

Test Run Overview​

Test runs execute your workflow through the same pipeline as production executions. There is no separate "test environment" -- testing creates a real execution record and processes each node. Runs of an undeployed workflow are tagged as development runs (and are not billed as production executions); runs of a deployed workflow count as production.

Starting a Test Run (Batch Workflows)​

  1. Save the workflow (the Run button is disabled until the workflow is saved and has nodes)
  2. Click the Run button in the workflow builder toolbar. On a draft, the Run draft dialog opens (see Giving the trigger test data)
  3. The builder validates node configuration first; errors open the validation sidebar and block the run
  4. The workflow is auto-saved and the execution is submitted for processing
  5. Watch the run bar above the canvas and each node's status on the canvas
  6. Review execution traces and output data

While the run is in progress the Run button reads Running... and stays disabled until the run ends, also after a page reload. The run bar under the toolbar says what the run is doing, from the run's own state:

Run barMeaning
Saving and submitting the run...The workflow is being saved and the run submitted
Starting worker: ...The run is waiting for its worker pod, with the pod's state as Kubernetes reports it, for example waiting for a node (Unschedulable) while a node is added, ContainerCreating (workflow-worker), or not Ready (ContainersNotReady). A cold start can take a minute or more; this is not an error
Running: 3 of 5 nodes doneThe workflow is running; failed nodes and an active loop's progress are added
Workflow completedThe run ended successfully; the bar closes on its own
A warning or errorThe run ended any other way (failed, partly succeeded, completed with gaps, stopped); the bar stays up with the reason until you close it

An error is only shown when the run's state says there is one: a worker that cannot start (an image that cannot be pulled, a node that cannot be provisioned) ends the run failed with Kubernetes' reason.

Giving the trigger test data​

A draft's trigger has nothing to capture while you build: no request has arrived, no email, no feed item, no chat message. So on a draft, Run opens the Run draft dialog instead of starting the run at once:

  1. If the workflow has more than one trigger, choose the Trigger the run starts from.
  2. Edit the trigger's output, in JSON. The editor starts from the example output the trigger declares (for a REST API trigger, method, body, headers and the rest), so the fields the next nodes map are already there. Reset to example puts the example back.
  3. Click Run.

The trigger does not run. Its output in this run is the data you entered, and the nodes after it receive it through their input mappings exactly as they would from a real event. For example, with a REST API trigger and a ReAct Agent whose Task is mapped to data.body.task, enter {"method": "POST", "body": {"task": "In one sentence, what causes ocean tides?"}}. In the run's trace, the trigger's span is completed with the data you entered as its output and is marked as test input.

Text that is not valid JSON is refused in the dialog, before anything runs. Test trigger data is accepted only on a draft run, and only for trigger nodes of the workflow being run. Retry from start on such a run runs the saved draft again with the same data; once the workflow is deployed, its runs take no test data.

A deployed workflow's builder Run (and a Run by someone who can view but not edit the workflow) starts the workflow with empty trigger inputs. To run a deployed workflow with a realistic payload, trigger it with data:

  • REST API: POST /api/v1/workflows/{workflowId}/execute with a triggerInputs object in the body. Both draft and deployed workflows can be executed this way.
  • Webhook trigger: call the workflow's webhook endpoint (/api/v1/webhooks/{workflowId}) with the payload your external system would send.

Starting a Test Session (Streaming Workflows)​

For streaming workflows, the toolbar Run button opens the live test panel instead:

  1. Save the workflow, then click Run ("Run with live mic / text"): the test panel opens on the right
  2. Click Start Test Session to connect
  3. Interact with the workflow using your microphone or by typing text; the conversation transcript streams into the panel
  4. The live session doubles as the execution for tracing: clicking a node on the canvas during or after the session shows that session's spans

A builder test runs what you last saved, on your own test worker, never on the workflow's deployment. Save after each change before running again. See Streaming Workflows.

What Happens During a Test​

  1. Execution record created with status pending

  2. STRONGLY_SERVICES generated dynamically based on the workflow's node dependencies (add-ons, data sources, AI models)

  3. Execution submitted for processing; a submission failure marks the execution failed and queues an automatic retry

  4. Nodes processed through the workflow graph

  5. Traces recorded for each node execution

  6. Execution status updated with an end timestamp:

    • completed: every node ran and every loop delivered all its items
    • partial_success (the warning Partial success): the run finished but not everything succeeded: a node failed under continueOnError, or a loop, map or distributed loop delivered some of its items but not all. The run's error message names each failed node with its error and each loop with how many of its items failed, for example Score failed: ValueError: no rows; Process files: 124 of 150 items failed
    • failed: a node failed without continueOnError, or a loop had items and delivered none of them
    • completed_with_gaps (the warning Completed with gaps): the run reached the end but some nodes produced no output

    A loop's failed items still reach its aggregator as failed-item records (the workflow's failedItemPolicy), and continueOnError still lets the run go on after a failed node or failed items; neither makes the run report completed.

Real-Time Execution Visualization​

During a run, and afterwards for the workflow's latest run, the canvas shows each node's state in that run:

AppearanceMeaning
PulsingRunning
Red outlineFailed
Info button on hoverThe node ran in this run: click it to open its details in the side bar

A distributed node shows its items' progress (completed, failed, running and pending, and its active workers) when you hover it. The run bar above the canvas says how many nodes are done.

Execution Flow​

  • Nodes update their visual state as execution progresses through the graph
  • Parallel branches (via parallel-branch node) show multiple nodes executing simultaneously
  • Conditional branches (switch-case, conditional) show which path was taken

Node Tracing​

Every node execution produces a span that provides detailed tracing for each step.

Span Types​

TypeDescription
WORKFLOWRoot span covering the entire execution
NODEIndividual node execution
LLMAI Gateway / LLM call within a node
TOOLTool call (MCP or native)
RETRIEVALRAG or knowledge base retrieval
AGENTAgent reasoning loop
CHAINChain of operations
EMBEDDINGEmbedding generation
PARSERParsing operation
DISTRIBUTEDDistributed (fan-out) work item
SYNTHETICSynthetic span filled in for gaps in a trace
SESSIONStreaming: root span for a streaming session
TURNStreaming: one span per conversational turn
OPERATIONStreaming: STT/TTS API call within a node

Span Data​

Each span records:

FieldDescription
execution_idParent execution ID
span_typeOne of the types above
parent_span_idID of parent span (for tree structure)
start_timeWhen the span started
end_timeWhen the span completed
statusrunning, completed, failed
inputsData received by the node
outputsData produced by the node
errorError message if failed
metadataAdditional context (model used, token count, etc.)

Viewing Spans​

Click any node on the canvas to open the node debug panel for the current execution:

  • Output data: The JSON output produced by the node
  • Input data: What the node received from upstream nodes
  • Timing: Start time, end time, and duration
  • Errors: Error message and details if the node failed
  • Nested spans: For agent nodes, sub-spans for each LLM call and tool invocation

For past runs, open Workflow Monitor, select the workflow's execution history, and open an execution's trace to see the full span tree.

Span Tree​

Spans form a hierarchical tree:

WORKFLOW (root)
├── NODE: webhook (trigger)
├── NODE: ai-gateway
│ └── LLM: model call
├── NODE: react-agent
│ ├── AGENT: reasoning loop
│ ├── LLM: tool selection
│ ├── TOOL: web-search call
│ └── LLM: final response
└── NODE: webhook-response

Inspecting Execution Results​

Execution Details​

Click on a completed execution to view:

  • Execution ID and status (pending, running, completed, completed_with_gaps, partial_success, failed, cancelled); Completed with gaps and Partial success show as warnings
  • Start time and end time
  • Duration calculated from start to end
  • Trigger inputs that initiated the execution
  • Error message if the execution failed

Node-Level Output​

Click any node in a completed execution to view:

  • The JSON output produced by that node
  • Input data received from upstream connections
  • Duration of that specific node
  • Error details if the node failed

For a distributed node, the Items tab shows its work items per worker pod: how many items each pod ran and how many failed, its average item time and its items per second (items finished over the time the pod was busy, first item start to last item end). The node's overall items per second, over the same window, is shown in the header. Each item row shows the worker pod that ran it and links to the item's span in the run's trace, which opens with that span's details shown.

Workflow Statistics​

A workflow's Statistics page is opened from the Statistics action on the workflow's row in Workflow Monitor, or the Statistics button on its Execution History page. For the chosen time range (last hour, 24 hours, 7 days or 30 days) it shows how many runs ended each way (Completed, Partial success, Completed with gaps, Failed, Running, Stopped or cancelled), the success rate (completed runs only), the duration of completed runs (average, fastest, slowest, 95th percentile), trigger sources, node totals, and Recent Problems: the latest runs that did not fully succeed, each with its reason. Execution History also counts the runs that partly succeeded in their own card.

Data flows between nodes through input mappings: each of a node's inputs names a path into the data its directly connected upstream node produced, starting with data. For example, a node might map its text input to data.response from an upstream LLM node.

Execution Logs​

Each execution also writes log entries:

  • info level: execution start, node completions
  • error level: failures with error messages and details
  • Each log entry records the workflow, execution, user, and timestamp

Testing Best Practices​

Test Data Preparation​

  1. Use realistic data that matches production input formats
  2. Test edge cases: empty arrays, null values, missing fields
  3. Test error paths: invalid data that should trigger error handling
  4. Test conditional branches: provide input that exercises each branch of switch-case and conditional nodes

Iterative Testing​

  1. Start simple: Test with basic valid input first
  2. Review spans: Check each node's input/output in the span data
  3. Fix configuration: Update node settings based on span inspection
  4. Test edge cases: Gradually add boundary and error conditions
  5. Verify full flow: Ensure data passes correctly through all connections

Testing Agent Workflows​

Agent nodes (react-agent, supervisor-agent) produce multiple sub-spans:

  1. Review the agent's reasoning loop in the span tree
  2. Check which tools the agent selected and why
  3. Verify LLM call spans for correct prompt and response
  4. Test with different inputs to ensure the agent handles varied scenarios

Testing Data Source Connections​

When testing nodes that connect to external data sources (databases, APIs):

  1. Verify data source credentials are configured correctly
  2. Check that the connection type (addon or datasource) and IDs are set
  3. Test with a small query first before running full data operations
  4. Review the node's span output to confirm data was retrieved correctly

Common Testing Issues​

Issue: Execution Stays in Pending Status​

Possible causes:

  • The workflow already has an execution in flight (workflows run one execution at a time by default)
  • The workflow's compute is still starting
  • STRONGLY_SERVICES generation failed

Solutions:

  • Wait for or cancel the in-flight execution
  • Check the workflow's deployment status
  • Review execution logs for service configuration errors

Issue: Node Fails with Connection Error​

Possible causes:

  • Data source credentials are invalid or expired
  • Add-on service is not running
  • The external service is unreachable

Solutions:

  • Verify data source configuration in the Data Sources section
  • Check add-on health status
  • Review the node's span error message for specific details

Issue: Agent Not Using Tools​

Possible causes:

  • MCP Tools Provider not connected to the agent's tools connector
  • MCP server not enabled or not deployed
  • Agent system prompt does not mention available tools

Solutions:

  • Verify the tools connector (bottom of agent node) has a connection
  • Check the MCP server's status on the Workflow Tools page
  • Update the agent's system prompt to reference the available tools

Issue: Incorrect Data Passed Between Nodes​

Possible causes:

  • Input mapping references wrong field path
  • Upstream node output format changed
  • JSONPath expression is incorrect

Solutions:

  • Inspect span outputs for the upstream node to see actual data structure
  • Update input mappings to use correct JSONPath references (e.g., $.output.field)
  • Use set-fields node to reshape data between nodes

Execution Resume and Checkpoints​

Executions that stop in a resumable state (failed, error, paused) can be resumed from their last checkpoint:

  • The execution retains its ID and accumulated trace data and continues from the checkpoint rather than restarting
  • Resume is only available when the execution recorded a resumable checkpoint

A run waiting at Wait for Input, Human Feedback or Human Checkpoint is not resumed this way: it is still running, and you answer it in the builder's Waiting for you panel (or through the API). See Waiting for a Person.

Testing Checklist​

Before deploying to production, verify:

Functionality​

  • All nodes execute successfully with valid input
  • Data flows correctly through all connections
  • Conditional branches route correctly
  • Agent nodes select appropriate tools

Error Handling​

  • Invalid input is handled gracefully
  • External service failures produce clear error spans
  • Retry logic works for transient failures

Data Validation​

  • Node outputs match expected format
  • Input mappings reference correct JSONPath fields
  • Data transformations produce correct results

Performance​

  • Execution completes within acceptable time
  • No unnecessary sequential dependencies
  • Large data sets are handled without timeout

Next Steps​

Once your workflow is thoroughly tested: