Workflows

Automate your incident response

Build workflows that respond to alerts, schedules, and webhooks. Gather context, make decisions, and take action automatically, with human approval whenever the action could be destructive.

Alert → check → approve → fix — KloudMate correlation Workflow · disk cleanup Alert → check → approve → fix 01 Disk alert fires
webhook from a routing rule
02 Check the host
df -h through the agent
03 On-call approves
one click from their email
04 Fix, then verify
waits until space recovers
Any trigger Alert, schedule, webhook, form, or app event one workflow, one trigger Safe by default Approval before any destructive step nothing changes a resource unasked

An alert tells you what broke. It doesn't fix it.

Most incidents end the same way: someone reads the alert, runs the same three commands, and clicks the same buttons. A workflow does that part for you. It starts from the alert, gathers the same evidence, and makes the same fix, and it waits for a person before it changes anything that matters.

What teams can do with Workflows

Start from any event, act on your cloud and your hosts, route on conditions or a model, and pause for a person before anything destructive.

Start from any event

A workflow runs when an alert fires, a schedule comes due, a webhook or form arrives, or something happens in Slack or Jira. Alerts start one through a routing rule, the same way you'd add a Slack channel.

Act on your cloud and your hosts

Call an AWS or Azure API, or run a shell command on any Linux host through the KloudMate agent. Restart an instance, clear a disk, or rotate a key, not just read about the problem.

Branch, loop, and wait

Take one path or another on a condition, loop over the resources in an alert group, run branches in parallel, or wait until a change takes effect before you report back.

Pause for a human approval

Put an approval in front of anything destructive. The on-call engineer approves or rejects from a link in their email, no KloudMate login needed. Nothing changes a resource unless they say yes.

Work in the tools you already use

Post and edit Slack messages, open and transition Jira issues, call a tool on an MCP server, or send an HTTP request to anything else. Reuse a saved connection instead of pasting a token.

Reuse variables and remember state

Refer to a workspace constant or secret by name instead of pasting it into every workflow, and store a cursor or a marker so a later run knows what the last one already handled.

From an alert to a verified fix, on its own

A routing rule decides who hears about an alert. A workflow is the response you build for it, and runs the same steps every time.

01

Something starts the run

An alert, a schedule, a webhook, a form, or an app event kicks it off. Alerts arrive through a routing rule and carry the whole alert group with them.

02

It gathers the evidence

The workflow queries the host, the cloud API, or your own service, and branches on what it finds, the same lookups you'd run by hand.

03

A person approves the change

Before any step that changes a resource, the run pauses for an approval. The on-call engineer decides from their email, and the run waits as long as it needs to.

04

It makes the fix and verifies

The workflow runs the change, waits until the system reports healthy again, and posts the outcome to Slack. Every run is recorded step by step.

A routing rule starts the run — KloudMate label-based routing Trigger · alert to workflow A routing rule starts the run Alerts fire with labels Routing rules · priority Channel disk > 90% host=prod-db-2 5xx spike service=checkout consumer lag kafka 10 disk > 90% Disk cleanup 20 service=checkout Failover 100 * default Slack #alerts Disk cleanup Failover Slack #alerts

Start from the alert that already fired

There's no new alerting to set up. Point a webhook notification channel at the workflow and attach it to a routing rule, the same way you'd add Slack or email. Every notification starts a run with the full alert group, so the workflow knows the host, the service, and the labels.

  • Five triggers: a manual Run or API call, a schedule or cron, an inbound webhook, a public form, or an event in Slack or Jira
  • Alerts start a workflow through a routing rule, carrying the group's title, labels, and affected resources
  • Filter deliveries and dedupe on an id, so one event doesn't start the same workflow twice
What the run actually did — KloudMate Workflow run KloudMate · Workflow run Steps · one run What the run actually did Steps 5 Duration 3m 12s Approvals 1 Step Type Target Result Read disk usage df -h Agent 94% used ok Post to Slack #ops Slack message sent ok Approve cleanup on-call Approval approved by Ana ok Delete old logs /var/log/*.gz Agent 18 GB freed ok Verify free space wait until < 80% Wait until 71% used ok

Do the fix, not just describe it

A workflow acts on real resources. It runs a shell command on a host through the KloudMate agent, calls an AWS or Azure API, and posts to Slack or opens a Jira issue, all in one run. Every step records its input, output, and timing.

  • Run a command on any Linux host where a Developer has allowed runbook scripts
  • Call cloud APIs through a connected AWS or Azure account: stop an instance, clear a queue, rotate a key
  • Transform payloads between steps with JSONPath, CSV and XML parsing, or sandboxed JavaScript
Decide, repeat, and confirm — KloudMate correlation Control flow · one path or another Decide, repeat, and confirm 01 Check each host
loop over the alert group
02 Branch on result
disk full vs. stuck process
03 Run the matching fix
one path per outcome
04 Wait until healthy
re-check before reporting
Loop Once per resource in the group up to 100, 10 at a time Wait until Confirm the fix took effect then post the all-clear

Decide, repeat, and wait until it's fixed

Real remediation isn't a straight line. Branch on what a lookup returns, loop over every resource in an alert group, run independent work in parallel, and wait until the change takes effect before you post the all-clear.

  • Branch or Route to take one path per outcome, or let AI Route pick when a rule is hard to write
  • Loop over up to 100 items, ten at a time, to act on each host in a group
  • Wait until a condition holds, so you confirm the instance is back before reporting, not just that you asked
The on-call engineer decides — KloudMate Workflow run KloudMate · Workflow run Approval · human in the loop The on-call engineer decides Waited 4m Approver on-call Result approved 03:02 Approval requested emailed to on-call, one link each sent 03:06 Ana Ruiz approved from her email, no login approved 03:06 Cleanup ran deleted 18 GB of old logs 18 GB freed 03:09 Verified and posted free space back under 80% all-clear

Nothing destructive without a yes

Put an approval in front of any step that changes a resource. The on-call engineer gets an email with their own link and approves or rejects without signing in. A rejected run stops cleanly and skips the rest, and a run left waiting costs nothing.

  • Name up to 10 approvers; the first to respond decides, and every other link stops working
  • Collect Input gathers a form instead of a yes or no, and passes the values to later steps
  • A timeout you set falls back your way: fail the run, or continue down a safe path
See every run at a glance — KloudMate Workflows KloudMate · Workflows Run history · every run kept See every run at a glance Runs · 24h 142 Succeeded 138 Failed 2 Workflow Started by Duration Status Disk cleanup prod-db-2 Alert 3m 12s success Grant SSH access sandbox Form 2m 40s approved Nightly report 02:00 Schedule 48s success Restart worker checkout Webhook — failed Onboard hire manual Manual 5m 03s success

Test it before it's live, then see every run

Build against real data. Run a single step to see its actual output, then a full test run end to end before you publish. Once it's live, run history keeps every run step by step, so you can see what a workflow did and why a run failed.

  • Run one step at a time and build the next from its real output
  • A trigger only fires the published version while the workflow is enabled; draft edits change nothing until you publish
  • Run history keeps finished runs and their step detail for 30 days
KloudMate AI

Let a model decide when a rule is easier to describe than to write

Some payloads are too messy for a fixed condition. AI Route reads the situation and sends the run down the right path. An AI extract step pulls clean fields out of a raw webhook or log line, so the next step gets structured data. The model reads the payload as data, never as instructions.

  • Route Send each run down the right path from a plain-English description
  • Extract Pull structured fields out of a messy webhook or log payload
  • Contain The model reads payloads as data, so a payload can't redirect the task
Explore platform
Which team should handle this? — KloudMate Auto-RCA AI Route · pick a path Which team should handle this? Q
A new Jira issue arrived. Route it to the right on-call path.
Assistant · likely cause
  • The description points at checkout failing at payment capture, not a UI bug.
  • AI Route sends it down the payments path with high confidence.
  • The route, the reason, and the confidence are recorded on the step for you to check.
Chosen route payments high confidence Why payment capture errors not a frontend issue Recorded route · reason · confidence on the step output

Get started

From telemetry to root cause,
in one platform.

Connect your OpenTelemetry pipeline, AWS integrations, or eBPF agent. Distributed tracing, log management, alerting, and AI-assisted investigation: unified, with predictable pricing.