Alerting

One ticket per problem, not a dozen alerts

Get notified once for each real problem. KloudMate correlates related alerts from your metrics, logs, and traces into one on its own, with no group-by keys to define. Turn on a detector and it writes the anomaly and forecast rules for you.

checkout-api · elevated 5xx — KloudMate alert console KloudMate · Alerts Alerts checkout-api · elevated 5xx Alert when the 5xx rate stays above threshold for 5m OverviewInstancesHistoryRule Firing 7 Pending 2 Normal 31 No data 1 threshold env=prod service=checkout team=payments region=us-east-1 7 related alerts → 1 ticket AI-correlated group · on-call notified once

One failure shouldn't open a dozen tickets.

KloudMate's correlation engine links related alerts into one on its own, so a single failure notifies you once and opens one ticket, not one per signal. There are no group-by keys to maintain, and a likely cause is attached before you start digging.

What teams can do with Alerting

Get covered without picking thresholds, let AI correlate what fires into one incident, and arrive with a likely cause already attached.

Zero-config AI correlation

The correlation engine links related alerts into one incident on its own. No group-by keys, no rules to maintain; it learns which alerts fire together and improves from your feedback.

One ticket per real problem

Related alerts append to one group instead of paging again, so on-call gets a single notification and a single ticket, not one per symptom.

Auto-RCA attaches a likely cause on open

Every group it opens gets an AI investigation. The likely cause attaches to the alert and its notifications, so responders start with a lead, not a blank screen.

Detectors you turn on, not thresholds you guess

Smart Alerts covers hosts, Kubernetes, service latency, and Kafka with curated detectors. KloudMate writes each rule, keeps it current as pods and hosts come and go, and removes it when you turn the detector off.

Anomaly and forecast conditions

Anomaly learns each entity's normal range from its own history and fires when a reading lands outside it. Forecast follows the trend and fires days before a disk reaches its limit.

Precise rules when you want them

Start from a template for a common scenario, or build your own from queries and expressions across logs, metrics, traces, and CloudWatch. Fire only after a condition holds.

From raw signals to one explained ticket

The engine decides what counts as one problem, who hears about it, and why it happened, before the first notification goes out.

01

Alerts fire across your signals

Rules you wrote, templates you started from, and detectors KloudMate maintains all fire the same way, one alert per affected host, function, or service.

02

AI correlates them into one

The correlation engine links related alerts into a single group on its own, no group-by keys to define, and records why it grouped them.

03

The group routes to the right team

Each group goes to the team that owns it and notifies once. You set how often it re-notifies until someone acts.

04

It arrives with the cause attached

Auto-RCA investigates the moment the group opens, so the first notification already carries a likely cause.

48 related alerts → 6 tickets — KloudMate alert grouping AI alert correlation 48 related alerts → 6 tickets 48 related alerts fire AI correlates on its own 6 tickets opened Zero-config correlation the engine links related alerts across signals on its own Learns over time co-occurrence history + your Correct / Wrong feedback

Cut ticket volume, with zero config

Most alerts are the same failure seen from ten angles. KloudMate's correlation engine links related alerts across metrics, logs, and traces into one group on its own, with no group-by keys to define, so a single incident opens one ticket instead of a dozen.

  • Nothing to configure and no keys to pick, the engine correlates related alerts on its own
  • New alerts append to the group instead of paging again, so on-call is notified once
  • It learns which alerts fire together, and your Correct or Wrong feedback makes it sharper
How the engine correlates alerts — KloudMate correlation Why these were grouped How the engine correlates alerts 01 Alerts fire
across metrics, logs, and traces
02 AI correlates
links related alerts on its own
03 One group
reason: topology link
04 You confirm
Correct or Wrong trains it
Why grouped Topology + co-occurring history shown in plain English on the group Gets sharper learns what fires together from history and your feedback

See why the AI grouped them, and correct it

When the engine links alerts, it records the reason in plain English, a shared attribute, a topology link, or a history of firing together, so you can trust the group and set it straight when it's wrong.

  • Every correlated group shows its reason: shared attribute, topology, known cascade, or co-occurring history
  • Mark a grouping Correct or Wrong to train the engine on your own environment
  • A quiet workspace starts conservative and tightens as the engine learns what fires together
[Forecast] Filesystem fill · prod-db-2 — KloudMate alert console KloudMate · Alerts Alerts · Smart Alerts [Forecast] Filesystem fill · prod-db-2 Fires when the trend reaches the limit within 4 days OverviewInstancesHistoryRule Firing 1 Pending 2 Normal 46 No data 1 limit env=prod host=prod-db-2 mount=/var/lib team=platform On track to reach the limit in about 3d 4h 4-day horizon · you hear about it days before it fills

Know a disk will fill on Thursday, not at 3am

A forecast condition follows each entity's recent trend and fires while there's still time to act. An anomaly condition does the other half: it learns what normal looks like for that entity and fires when a reading lands outside it. Neither asks you for a number.

  • Forecast horizons of 1, 4, or 7 days, and the alert says how long you have: on track to reach the limit in about 3d 4h
  • One sensitivity setting for an anomaly, High, Medium, or Low, plus a direction when only one side matters
  • A new host collects a baseline before it can fire, so a warm-up ramp isn't an incident
  • Both run on KloudMate metric queries; CloudWatch, log, and trace rules use thresholds
Detectors KloudMate keeps current — KloudMate Alerts KloudMate · Alerts Alerts · Smart Alerts Detectors KloudMate keeps current Detectors on 14 Entities watched 1,240 Firing 1 Detector Method State Entities CPU utilization Hosts & VMs · sensitivity medium anomaly watching 48 hosts Filesystem fill Hosts & VMs · 4-day horizon forecast firing 2 volumes Pod memory vs limit Kubernetes & Containers anomaly watching 312 pods Consumer lag Kafka & Queues anomaly no data yet — Managed for you The rules follow your fleet · pods restart and hosts scale; the rules keep up

Turn on a detector, get the rule

Smart Alerts covers the infrastructure nobody gets around to writing rules for. Pick the detectors you want and KloudMate writes each rule. It follows your fleet as pods restart and hosts scale, and removes the rule when you turn the detector off.

  • Packs for hosts and VMs, Kubernetes and containers, service latency and throughput, and Kafka
  • Managed rules fire into the same grouping, routing, and channels as the ones you write
  • Tune sensitivity or horizon in place; edit the query and the rule becomes yours to own
Query → reduce → condition — KloudMate expression builder Alert rule Query → reduce → condition A Query · checkout latency p95(http.server.duration) by service metrics B Reduce mean(A) over last 5m Math Reduce Condition C Condition B IS ABOVE 300ms Math Reduce Condition Evaluation Every 60s · pending 5m Firing

Alert on real conditions, not a single static threshold

Real problems rarely trip one threshold. KloudMate builds rules from queries and expressions across logs, metrics, traces, and CloudWatch, so you can require math, ratios, and several conditions before anything fires.

  • Start from a template for a common scenario, describe the alert in plain English, or build it from scratch
  • Combine multiple queries with math, reduce (mean, max, min, sum, last, count), and condition expressions
  • Pull from KloudMate telemetry or AWS CloudWatch, and use PromQL for OpenTelemetry data
  • Catch per-resource problems with multi-dimensional alerts: one rule, one alert per affected host or function
KloudMate AI

Every alert group opens with a likely cause attached

Turn on Auto-RCA and KloudMate investigates the moment a group opens. The likely cause attaches to the alert and its notifications, so responders arrive with a lead instead of a blank screen.

  • Correlate Link related alerts into one incident automatically
  • Investigate Run Auto-RCA on group open and attach the likely cause
  • Explain Summarize the result and condition that fired each alert
Explore platform
What caused this alert group? — KloudMate Auto-RCA Auto-RCA on group open What caused this alert group? Q
Summarize this alert group and what caused it to open.
Assistant · likely cause
  • Seven related alerts across prod and staging were correlated into one group.
  • Auto-RCA points at an inventory dependency timing out, with matching error logs in the same window.
  • The group is still open. The notification is in the on-call Slack channel with the investigation attached.
Scope 7 instances · 1 group prod and staging firing Likely cause inventory dependency timeout from Auto-RCA on group open Delivery slack-oncall notified once per-channel outcome logged

Get started

From telemetry to root cause,
in one platform.

Connect your OpenTelemetry pipeline, AWS integrations, or eBPF agent. Distributed tracing, log management, alerting, and AI-assisted investigation: unified, with predictable pricing.