Automate your incident response
Build workflows that respond to alerts, schedules, and webhooks. Gather context, make decisions, and take action automatically, with human approval whenever the action could be destructive.
An alert tells you what broke. It doesn't fix it.
Most incidents end the same way: someone reads the alert, runs the same three commands, and clicks the same buttons. A workflow does that part for you. It starts from the alert, gathers the same evidence, and makes the same fix, and it waits for a person before it changes anything that matters.
What teams can do with Workflows
Start from any event, act on your cloud and your hosts, route on conditions or a model, and pause for a person before anything destructive.
Start from any event
A workflow runs when an alert fires, a schedule comes due, a webhook or form arrives, or something happens in Slack or Jira. Alerts start one through a routing rule, the same way you'd add a Slack channel.
Act on your cloud and your hosts
Call an AWS or Azure API, or run a shell command on any Linux host through the KloudMate agent. Restart an instance, clear a disk, or rotate a key, not just read about the problem.
Branch, loop, and wait
Take one path or another on a condition, loop over the resources in an alert group, run branches in parallel, or wait until a change takes effect before you report back.
Pause for a human approval
Put an approval in front of anything destructive. The on-call engineer approves or rejects from a link in their email, no KloudMate login needed. Nothing changes a resource unless they say yes.
Work in the tools you already use
Post and edit Slack messages, open and transition Jira issues, call a tool on an MCP server, or send an HTTP request to anything else. Reuse a saved connection instead of pasting a token.
Reuse variables and remember state
Refer to a workspace constant or secret by name instead of pasting it into every workflow, and store a cursor or a marker so a later run knows what the last one already handled.
From an alert to a verified fix, on its own
A routing rule decides who hears about an alert. A workflow is the response you build for it, and runs the same steps every time.
Something starts the run
An alert, a schedule, a webhook, a form, or an app event kicks it off. Alerts arrive through a routing rule and carry the whole alert group with them.
It gathers the evidence
The workflow queries the host, the cloud API, or your own service, and branches on what it finds, the same lookups you'd run by hand.
A person approves the change
Before any step that changes a resource, the run pauses for an approval. The on-call engineer decides from their email, and the run waits as long as it needs to.
It makes the fix and verifies
The workflow runs the change, waits until the system reports healthy again, and posts the outcome to Slack. Every run is recorded step by step.
Start from the alert that already fired
There's no new alerting to set up. Point a webhook notification channel at the workflow and attach it to a routing rule, the same way you'd add Slack or email. Every notification starts a run with the full alert group, so the workflow knows the host, the service, and the labels.
- Five triggers: a manual Run or API call, a schedule or cron, an inbound webhook, a public form, or an event in Slack or Jira
- Alerts start a workflow through a routing rule, carrying the group's title, labels, and affected resources
- Filter deliveries and dedupe on an id, so one event doesn't start the same workflow twice
Do the fix, not just describe it
A workflow acts on real resources. It runs a shell command on a host through the KloudMate agent, calls an AWS or Azure API, and posts to Slack or opens a Jira issue, all in one run. Every step records its input, output, and timing.
- Run a command on any Linux host where a Developer has allowed runbook scripts
- Call cloud APIs through a connected AWS or Azure account: stop an instance, clear a queue, rotate a key
- Transform payloads between steps with JSONPath, CSV and XML parsing, or sandboxed JavaScript
Decide, repeat, and wait until it's fixed
Real remediation isn't a straight line. Branch on what a lookup returns, loop over every resource in an alert group, run independent work in parallel, and wait until the change takes effect before you post the all-clear.
- Branch or Route to take one path per outcome, or let AI Route pick when a rule is hard to write
- Loop over up to 100 items, ten at a time, to act on each host in a group
- Wait until a condition holds, so you confirm the instance is back before reporting, not just that you asked
Nothing destructive without a yes
Put an approval in front of any step that changes a resource. The on-call engineer gets an email with their own link and approves or rejects without signing in. A rejected run stops cleanly and skips the rest, and a run left waiting costs nothing.
- Name up to 10 approvers; the first to respond decides, and every other link stops working
- Collect Input gathers a form instead of a yes or no, and passes the values to later steps
- A timeout you set falls back your way: fail the run, or continue down a safe path
Test it before it's live, then see every run
Build against real data. Run a single step to see its actual output, then a full test run end to end before you publish. Once it's live, run history keeps every run step by step, so you can see what a workflow did and why a run failed.
- Run one step at a time and build the next from its real output
- A trigger only fires the published version while the workflow is enabled; draft edits change nothing until you publish
- Run history keeps finished runs and their step detail for 30 days
Let a model decide when a rule is easier to describe than to write
Some payloads are too messy for a fixed condition. AI Route reads the situation and sends the run down the right path. An AI extract step pulls clean fields out of a raw webhook or log line, so the next step gets structured data. The model reads the payload as data, never as instructions.
- Route Send each run down the right path from a plain-English description
- Extract Pull structured fields out of a messy webhook or log payload
- Contain The model reads payloads as data, so a payload can't redirect the task
Related Features
Keep the rest of the workflow close by so teams can move between detection, investigation, and response without losing context.
Alerting
Correlate related alerts into one incident with AI, notify the right team once, and attach a likely cause automatically.
Learn moreIncident Management
Coordinate response, ownership, escalation, and telemetry context in one incident workflow.
Learn moreKloudMate Assistant
Use natural language to correlate telemetry, summarize incidents, and guide the next investigation step.
Learn moreReliability & SLOs
Set SLOs, track error budgets, and get notified on burn rate before the budget runs out.
Learn moreGet started
From telemetry to root cause,
in one platform.
Connect your OpenTelemetry pipeline, AWS integrations, or eBPF agent. Distributed tracing, log management, alerting, and AI-assisted investigation: unified, with predictable pricing.