yeke.io · docs · enterprise

Alerts

Stop waiting on the monitoring page: when a threshold is crossed or data stops arriving, word goes to the channel you chose. This guide covers rules, the channel and the limits.

One item, one flag — Enterprise metrics-alerts. Scheduled scans and reports are part of this item now; anomaly detection is measured in shadow mode and doesn't reach a channel yet.

What it's for

Built on top of Community's monitoring and scanning; what was missing was a layer that watches and notifies in your place.

The monitoring page already has a "needs attention" list today — but you have to open the page to see it. Alerts ties that view to a rule: when a threshold is crossed or data stops arriving, word goes to the channel you chose (email, webhook, chat, ITSM).

  • The alert list lives in a new tab on the monitoring page. The count of open alerts also shows as a chip in the top bar.
  • Rules, channels, routing, silences and maintenance windows are written only by an admin. The alert list itself is filtered by each user's own Kubernetes permissions: an alert in a namespace you can't see doesn't show up in the list, but its count stays visible in a chip — it isn't silently swallowed.

Rules: threshold and "data not arriving"

A decision is never made on a stale sample; the data itself going silent is a separate rule.

A threshold rule is made of an entity kind (node, pod, workload, PVC), a metric, a comparison and a threshold value. For a crossing to count as an event it must hold for at least a minimum duration, and for it to count as resolved it must stay below the threshold for a separate duration — a single sample never opens or closes anything.

  • If the latest sample is older than 3 minutes, the threshold rule is not evaluated at all. No verdict — good or bad — is given on a frozen value; the entity is marked stale and its state stays wherever it was.
  • "Data not arriving" is its own rule kind. The collector itself, the node, and the flow of samples (ingest) are each watched separately — the three point at different failures.
  • An unanswered question doesn't count as a failure. A cluster whose state couldn't even be asked for doesn't raise a "no data" event; the two show up as distinct states in the list.
  • Severity (info/warning/critical) only affects routing, not the measurement itself.

Built-in rule set

Derived from the monitoring page's own "needs attention" thresholds; the screen and the alert always agree on the same number.

Every installation ships with 14 built-in rules, enabled by default (11 threshold rules + 3 "data not arriving" rules); when they fire they show up on the monitoring page. None of them has a default channel — to send them to a channel you define a route.

  • The fields that define the measurement are locked. Kind, entity kind, metric, comparison, threshold and scope cannot be changed, and a built-in rule cannot be deleted — only disabled. The point is that what the screen paints red and what the alert fires on never drift apart.
  • Enabled/disabled, minimum duration, resolve duration and severity are free to change. These are the alert's own concepts with no counterpart on the screen, so they can be tuned to your environment.
  • If you want to measure your own threshold, write your own rule instead of editing a built-in one — the same measurement living under two different thresholds leaves you unsure which one to trust.

Routing

Every matching route sends; there's no such thing as a silenced route.

A route carries a minimum severity, an optional cluster/rule filter, and the channels it sends to. If an alert matches more than one route, all of them send — a duplicate notification is visible, a missing one is not.

  • An unresolved alert is re-sent as a reminder every four hours, and a resolution is announced too — when each one actually reaches the channel is covered below.
  • Alerts that fire in a cluster are collected for a 5-minute window and sent as one message for that cluster (critical doesn't wait). The body sent to an HTTP hook (ITSM, etc.) isn't affected by this collection — it's still one line per rule × cluster, carrying at most 50 entities and stating the remainder as a count.
  • A target that fires and resolves in quick succession (flapping) is now suppressed: the close waits 30 minutes, and if the target reopens during that time the close never goes out; the Alerts tab shows a "reopened N times" badge.

Collection window, close hold, flap suppression and HTTP/ITSM timing: Notification timing.

Notification timing

A notification decision is deferred against the event's own due time. The one deliberate suppression: a firing shorter than 5 minutes that doesn't recur within 30 minutes is not sent to the channel; it still shows in the Alerts tab.

An alert bound for a human channel (email, chat) no longer goes out instantly — it's collected per cluster instead. The body sent to an HTTP hook (ITSM, etc.) is unaffected by this; only the timing changes.

  • Collection window: 5 minutes per cluster. Alerts that fire in a cluster are collected for 5 minutes, then sent as one message for that cluster; the window is fixed — a new firing doesn't extend it. A critical firing doesn't wait — it goes out in the same pass, taking any alerts still waiting with it.
  • A close always waits 30 minutes (regardless of severity), then goes out in its own window of at most 5 more minutes — roughly 30–35 minutes in total. If the same target reopens during that wait, the close never goes out and the new firing isn't announced either: as far as the channel is concerned, the alert never closed. A close that is sent says "Not reopened for 30 min," without claiming a cause.
  • A target that flaps open and closed several times is reported once. The close message states how many times it reopened; the Alerts tab shows a "reopened N times" badge too.
  • If monitoring was interrupted during the wait, or the target is no longer read, the message says so: a close's measurement line adds "· monitoring was interrupted for N min"; for an entity that's gone missing it says "closed — target no longer reported" instead.
  • The body sent to an HTTP hook (ITSM, etc.) is UNCHANGED, only the timing is: firing arrives about 5 minutes late, closing about 35 minutes late. On a flapping event the id can change — match your records by fingerprint, not id.
  • Audit logging and SIEM export aren't affected by these windows — both stay real-time.

These figures hold under normal operation. If a cluster's evaluation falls behind (unreachable, license suspended), its close message is deferred until a total of 30 minutes of fresh data has come in from that cluster; the outage itself doesn't count. There is no upper bound: the close is delayed, not lost.

Silences and maintenance windows

Both only silence alerts; neither affects a change-freeze window.

A silence is temporary and requires a reason — you can still pick a scope wide enough to "mute this rule on this cluster until tomorrow." A maintenance window repeats weekly and uses the same calendar format as change-freeze, but its list is separate: a window that forbids changes doesn't also mean alerts go quiet in it — the two are opened independently.

A silence can now be edited in place: the reason, scope and dates can be rewritten, updating the same record instead of opening a new one. Who edited it and when shows up in the list; which fields changed is recorded separately, but the previous reason or dates are not kept.

The "Silence" shortcut on the monitoring page now defaults to that alert's own target only — for a pod, that means the pod's owning workload, which stays valid even if the pod restarts; for a pod whose owner can't be resolved, or for a workload, node or volume alert, it's the target itself. "All targets on this cluster" is still selectable, but it's now the second option, not the default. Because a CronJob's pod is owned by that run's own Job, the narrow option only silences that run on a CronJob; the next run fires again.

Scheduled scans and reports

The scan you used to start by hand can now run on its own, on a regular interval.

Each cluster can have at most one schedule: an interval (every 6 hours, 12 hours, 24 hours, or weekly), an anchor time and a timezone (IANA, local time). The schedule runs with the identity of the admin who created it.

  • The owner check runs again on every occurrence — a recorded permission is never trusted on its own. The schedule pauses, and becomes visible only to admins in the top bar, if: the owner has been disabled, the owner's admin role has been revoked, the cluster's identity no longer resolves, the owner's directory group membership (re-queried from the server) has narrowed, an OIDC owner hasn't logged in for 30 days, or the cluster's AI egress has been turned off.
  • A transient failure doesn't stop it. If the directory is unreachable, the tunnel is down, or the daily AI budget ran out, the same occurrence is retried within a window; only an identity that can't be verified pauses the schedule.
  • When the report finishes, it goes to the notification channel: finding counts by severity, the first 10 findings, and a link to the report itself — the full report is never embedded in the channel body.
  • A single-file HTML report (Enterprise); the scan's JSON output stays free in Community.

Anomaly detection (shadow mode)

Not a fixed threshold — a deviation from each node's and workload's own last 14 days; in this release it's only measured, not delivered.

Four built-in rules ship: node and workload × CPU and memory. A rule looks at the entity's own last 14 days' time-of-day profile (median and median absolute deviation, MAD) and fires once a deviation holds for at least 60 minutes without interruption and exceeds 10% of the node's own capacity, or 1% of the cluster's capacity for a workload. Workloads under 1% of the cluster and deviations shorter than 60 minutes are never seen — by design, not by omission. An entity with less than 7 days of history is never evaluated ("no profile").

  • Shadow mode in this release. A firing anomaly shows up in the Alerts tab and carries its own badge, but it never reaches a channel and is never counted in the top bar's open-alert chip. The goal is to validate the method against live data; delivery to a channel is the subject of a later release.
  • Scope is node and workload. Network and disk are out of scope because neither has a capacity denominator; pods are out of scope too — pod history is kept for at most 48 hours, not enough for a 14-day profile.
  • You can mark an event "expected behavior" (anomaly rules only). It's a reversible feedback record; it doesn't change the profile on the spot.
  • You can also write your own anomaly rule — the same constraints apply (node or workload, two metrics, at least 60 minutes) — and open it to live delivery; the four built-in rules stay locked to shadow.

Email channel and SMTP

The channel is a second transport on the existing notification hook — no new queue was opened.

In the Administration → Notifications → Hooks → New hook screen, choosing Email for Transport replaces the Address field with a Recipients field, where the destination is written. Delivery goes out through the SMTP relay, and the hooks' allow-list is never consulted.

  1. In Administration → Notifications → Hooks, the SMTP server card at the bottom of the screen: enter the server, port, TLS setting, a username and password if required, the From address, and a sender name, then save. Use Send a test email right below it to test the relay. This card is a single record shared by every email hook — saving it alone does not turn on a channel.
  2. On the same screen, New hook: set Type to Notification, Transport to Email, and list one address per line in Recipients. Add a body language if you want one.
  3. Alerts → Routing → New route: pick a severity, clusters and rules, then check this hook in the Channels list. Without a route, an alert shows up on screen but is never sent; every route that matches sends its own copy, so two matching routes mean two emails.
  • The SMTP relay is a single record, living as a card on the same Hooks screen — there's no separate record per cluster or per channel. The password is stored encrypted.
  • A relay on the organization's own internal network can be used. Only loopback and link-local addresses are rejected; internal network ranges are not blocked.
  • The email body's language is chosen per channel (Turkish/English) — the same rule can go out to two channels in two different languages at the same time.
  • The same channel also carries operation notifications (plan applied/rejected and similar): the reason an apply failed, which step it failed on, and the apiserver's actual response each get their own line in the body — nothing silently drops.
  • ITSM integration is not a separate channel — it's the existing webhook transport; connecting to an ITSM applies unchanged.

Without a license

Configuration locks, the engine stops, history stays readable — the three are not the same "403".

ItemWithout a license
Writing a rule / channel / route / silence / window 403 — configuration cannot be changed.
The evaluator (alert engine) Stops, and says so: a chip in the top bar and a record in the audit trail — it does not stop silently.
Existing rules, past events, delivery records Stay readable; a lapsed license does not cut off access to history.

Limits

What this release does not do — the answer if a later round asks to add it in.

  • Anomaly events don't reach a channel. In shadow mode they only show up in the Alerts tab; delivery to email, webhook, chat or ITSM is a later release's subject.
  • Anomaly scope is node and workload. There's no built-in rule for pods, network or disk — pods for a data reason (no history past 48 hours), network/disk for a method reason (neither has a capacity denominator).
  • No automatic action from an alert. An alert does not and will not trigger an operation plan — that would open a second write surface bypassing the approval card and the dry-run.
  • No Prometheus / Alertmanager compatibility. Neither PromQL rule import nor export to Alertmanager.
  • No outbound calling, SMS or voice. PagerDuty and similar services connect through the webhook; YEKE does not become a calling provider.
  • No on-call rotation. Routing goes to a channel, not to a person.
  • No log-based alerting. The product offers a log stream but does not store logs; a rule can't run over something that isn't retained.
  • No event correlation / root-cause analysis. That belongs to the scan-and-diagnose axis.
  • No GitOps management of rules. Rules, channels, routes, silences and maintenance windows live only in the database.

The channel is built on the existing hook mechanism

No new queue or delivery logic — the signed body, retries and dead-letter queue are the same mechanism described on the hook page.