yeke.io · docs · enterprise
Alerts
Stop waiting on the monitoring page: when a threshold is crossed or data stops arriving, word goes to the channel you chose. This guide covers rules, the channel and the limits.
One item, one flag — Enterprise metrics-alerts. Scheduled scans and
reports are part of this item now; anomaly detection is measured in shadow mode and doesn't
reach a channel yet.
What it's for
Built on top of Community's monitoring and scanning; what was missing was a layer that watches and notifies in your place.
The monitoring page already has a "needs attention" list today — but you have to open the page to see it. Alerts ties that view to a rule: when a threshold is crossed or data stops arriving, word goes to the channel you chose (email, webhook, chat, ITSM).
- The alert list lives in a new tab on the monitoring page. The count of open alerts also shows as a chip in the top bar.
- Rules, channels, routing, silences and maintenance windows are written only by an admin. The alert list itself is filtered by each user's own Kubernetes permissions: an alert in a namespace you can't see doesn't show up in the list, but its count stays visible in a chip — it isn't silently swallowed.
Rules: threshold and "data not arriving"
A decision is never made on a stale sample; the data itself going silent is a separate rule.
A threshold rule is made of an entity kind (node, pod, workload, PVC), a metric, a comparison and a threshold value. For a crossing to count as an event it must hold for at least a minimum duration, and for it to count as resolved it must stay below the threshold for a separate duration — a single sample never opens or closes anything.
- If the latest sample is older than 3 minutes, the threshold rule is not evaluated at all. No verdict — good or bad — is given on a frozen value; the entity is marked stale and its state stays wherever it was.
- "Data not arriving" is its own rule kind. The collector itself, the node, and the flow of samples (ingest) are each watched separately — the three point at different failures.
- An unanswered question doesn't count as a failure. A cluster whose state couldn't even be asked for doesn't raise a "no data" event; the two show up as distinct states in the list.
- Severity (info/warning/critical) only affects routing, not the measurement itself.
Built-in rule set
Derived from the monitoring page's own "needs attention" thresholds; the screen and the alert always agree on the same number.
Every installation ships with 14 built-in rules, enabled by default (11 threshold rules + 3 "data not arriving" rules); when they fire they show up on the monitoring page. None of them has a default channel — to send them to a channel you define a route.
- The fields that define the measurement are locked. Kind, entity kind, metric, comparison, threshold and scope cannot be changed, and a built-in rule cannot be deleted — only disabled. The point is that what the screen paints red and what the alert fires on never drift apart.
- Enabled/disabled, minimum duration, resolve duration and severity are free to change. These are the alert's own concepts with no counterpart on the screen, so they can be tuned to your environment.
- If you want to measure your own threshold, write your own rule instead of editing a built-in one — the same measurement living under two different thresholds leaves you unsure which one to trust.
Routing
Every matching route sends; there's no such thing as a silenced route.
A route carries a minimum severity, an optional cluster/rule filter, and the channels it sends to. If an alert matches more than one route, all of them send — a duplicate notification is visible, a missing one is not.
- An unresolved alert is re-sent as a reminder every four hours, and a resolution is announced too — when each one actually reaches the channel is covered below.
- Alerts that fire in a cluster are collected for a 5-minute window and sent as one message for that cluster (critical doesn't wait). The body sent to an HTTP hook (ITSM, etc.) isn't affected by this collection — it's still one line per rule × cluster, carrying at most 50 entities and stating the remainder as a count.
- A target that fires and resolves in quick succession (flapping) is now suppressed: the close waits 30 minutes, and if the target reopens during that time the close never goes out; the Alerts tab shows a "reopened N times" badge.
Collection window, close hold, flap suppression and HTTP/ITSM timing: Notification timing.
Notification timing
A notification decision is deferred against the event's own due time. The one deliberate suppression: a firing shorter than 5 minutes that doesn't recur within 30 minutes is not sent to the channel; it still shows in the Alerts tab.
An alert bound for a human channel (email, chat) no longer goes out instantly — it's collected per cluster instead. The body sent to an HTTP hook (ITSM, etc.) is unaffected by this; only the timing changes.
- Collection window: 5 minutes per cluster. Alerts that fire in a cluster are collected for 5 minutes, then sent as one message for that cluster; the window is fixed — a new firing doesn't extend it. A critical firing doesn't wait — it goes out in the same pass, taking any alerts still waiting with it.
- A close always waits 30 minutes (regardless of severity), then goes out in its own window of at most 5 more minutes — roughly 30–35 minutes in total. If the same target reopens during that wait, the close never goes out and the new firing isn't announced either: as far as the channel is concerned, the alert never closed. A close that is sent says "Not reopened for 30 min," without claiming a cause.
- A target that flaps open and closed several times is reported once. The close message states how many times it reopened; the Alerts tab shows a "reopened N times" badge too.
- If monitoring was interrupted during the wait, or the target is no longer read, the message says so: a close's measurement line adds "· monitoring was interrupted for N min"; for an entity that's gone missing it says "closed — target no longer reported" instead.
- The body sent to an HTTP hook (ITSM, etc.) is UNCHANGED, only the timing is:
firing arrives about 5 minutes late, closing about 35 minutes late. On a flapping event the
idcan change — match your records byfingerprint, notid. - Audit logging and SIEM export aren't affected by these windows — both stay real-time.
These figures hold under normal operation. If a cluster's evaluation falls behind (unreachable, license suspended), its close message is deferred until a total of 30 minutes of fresh data has come in from that cluster; the outage itself doesn't count. There is no upper bound: the close is delayed, not lost.
Silences and maintenance windows
Both only silence alerts; neither affects a change-freeze window.
A silence is temporary and requires a reason — you can still pick a scope wide enough to "mute this rule on this cluster until tomorrow." A maintenance window repeats weekly and uses the same calendar format as change-freeze, but its list is separate: a window that forbids changes doesn't also mean alerts go quiet in it — the two are opened independently.
A silence can now be edited in place: the reason, scope and dates can be rewritten, updating the same record instead of opening a new one. Who edited it and when shows up in the list; which fields changed is recorded separately, but the previous reason or dates are not kept.
The "Silence" shortcut on the monitoring page now defaults to that alert's own target only — for a pod, that means the pod's owning workload, which stays valid even if the pod restarts; for a pod whose owner can't be resolved, or for a workload, node or volume alert, it's the target itself. "All targets on this cluster" is still selectable, but it's now the second option, not the default. Because a CronJob's pod is owned by that run's own Job, the narrow option only silences that run on a CronJob; the next run fires again.
Scheduled scans and reports
The scan you used to start by hand can now run on its own, on a regular interval.
Each cluster can have at most one schedule: an interval (every 6 hours, 12 hours, 24 hours, or weekly), an anchor time and a timezone (IANA, local time). The schedule runs with the identity of the admin who created it.
- The owner check runs again on every occurrence — a recorded permission is never trusted on its own. The schedule pauses, and becomes visible only to admins in the top bar, if: the owner has been disabled, the owner's admin role has been revoked, the cluster's identity no longer resolves, the owner's directory group membership (re-queried from the server) has narrowed, an OIDC owner hasn't logged in for 30 days, or the cluster's AI egress has been turned off.
- A transient failure doesn't stop it. If the directory is unreachable, the tunnel is down, or the daily AI budget ran out, the same occurrence is retried within a window; only an identity that can't be verified pauses the schedule.
- When the report finishes, it goes to the notification channel: finding counts by severity, the first 10 findings, and a link to the report itself — the full report is never embedded in the channel body.
- A single-file HTML report (Enterprise); the scan's JSON output stays free in Community.
Anomaly detection (shadow mode)
Not a fixed threshold — a deviation from each node's and workload's own last 14 days; in this release it's only measured, not delivered.
Four built-in rules ship: node and workload × CPU and memory. A rule looks at the entity's own last 14 days' time-of-day profile (median and median absolute deviation, MAD) and fires once a deviation holds for at least 60 minutes without interruption and exceeds 10% of the node's own capacity, or 1% of the cluster's capacity for a workload. Workloads under 1% of the cluster and deviations shorter than 60 minutes are never seen — by design, not by omission. An entity with less than 7 days of history is never evaluated ("no profile").
- Shadow mode in this release. A firing anomaly shows up in the Alerts tab and carries its own badge, but it never reaches a channel and is never counted in the top bar's open-alert chip. The goal is to validate the method against live data; delivery to a channel is the subject of a later release.
- Scope is node and workload. Network and disk are out of scope because neither has a capacity denominator; pods are out of scope too — pod history is kept for at most 48 hours, not enough for a 14-day profile.
- You can mark an event "expected behavior" (anomaly rules only). It's a reversible feedback record; it doesn't change the profile on the spot.
- You can also write your own anomaly rule — the same constraints apply (node or workload, two metrics, at least 60 minutes) — and open it to live delivery; the four built-in rules stay locked to shadow.
Email channel and SMTP
The channel is a second transport on the existing notification hook — no new queue was opened.
In the Administration → Notifications → Hooks → New hook screen, choosing Email for Transport replaces the Address field with a Recipients field, where the destination is written. Delivery goes out through the SMTP relay, and the hooks' allow-list is never consulted.
- In Administration → Notifications → Hooks, the SMTP server card at the bottom of the screen: enter the server, port, TLS setting, a username and password if required, the From address, and a sender name, then save. Use Send a test email right below it to test the relay. This card is a single record shared by every email hook — saving it alone does not turn on a channel.
- On the same screen, New hook: set Type to Notification, Transport to Email, and list one address per line in Recipients. Add a body language if you want one.
- Alerts → Routing → New route: pick a severity, clusters and rules, then check this hook in the Channels list. Without a route, an alert shows up on screen but is never sent; every route that matches sends its own copy, so two matching routes mean two emails.
- The SMTP relay is a single record, living as a card on the same Hooks screen — there's no separate record per cluster or per channel. The password is stored encrypted.
- A relay on the organization's own internal network can be used. Only loopback and link-local addresses are rejected; internal network ranges are not blocked.
- The email body's language is chosen per channel (Turkish/English) — the same rule can go out to two channels in two different languages at the same time.
- The same channel also carries operation notifications (plan applied/rejected and similar): the reason an apply failed, which step it failed on, and the apiserver's actual response each get their own line in the body — nothing silently drops.
- ITSM integration is not a separate channel — it's the existing webhook transport; connecting to an ITSM applies unchanged.
Without a license
Configuration locks, the engine stops, history stays readable — the three are not the same "403".
| Item | Without a license |
|---|---|
| Writing a rule / channel / route / silence / window | 403 — configuration cannot be changed. |
| The evaluator (alert engine) | Stops, and says so: a chip in the top bar and a record in the audit trail — it does not stop silently. |
| Existing rules, past events, delivery records | Stay readable; a lapsed license does not cut off access to history. |
Limits
What this release does not do — the answer if a later round asks to add it in.
- Anomaly events don't reach a channel. In shadow mode they only show up in the Alerts tab; delivery to email, webhook, chat or ITSM is a later release's subject.
- Anomaly scope is node and workload. There's no built-in rule for pods, network or disk — pods for a data reason (no history past 48 hours), network/disk for a method reason (neither has a capacity denominator).
- No automatic action from an alert. An alert does not and will not trigger an operation plan — that would open a second write surface bypassing the approval card and the dry-run.
- No Prometheus / Alertmanager compatibility. Neither PromQL rule import nor export to Alertmanager.
- No outbound calling, SMS or voice. PagerDuty and similar services connect through the webhook; YEKE does not become a calling provider.
- No on-call rotation. Routing goes to a channel, not to a person.
- No log-based alerting. The product offers a log stream but does not store logs; a rule can't run over something that isn't retained.
- No event correlation / root-cause analysis. That belongs to the scan-and-diagnose axis.
- No GitOps management of rules. Rules, channels, routes, silences and maintenance windows live only in the database.