yeke.io · docs · security

Security Architecture and Threat Model

Technical review scope: YEKE 0.15.3 · 19 August 2026. This page explains trust boundaries, security controls and ways to verify them. Check release notes for changes in newer versions.

Scope, and how to read this page

What is guaranteed, where the control sits, and how to check each claim without our help.

YEKE is split along one line, written on the first day and not moved since: every line of code that runs inside your cluster is open; the centre is closed. The agent is Apache-2.0 at github.com/nairotech/yeke-agent; the wire protocol is its own Apache-2.0 package, @nairotech/yeke-tunnel (4.1.0 as of this review, its major number being the wire version). Core, the console and the AI layer are closed, with no general promise that this changes — a promise that can be withdrawn is worse than one never made.

This page describes where controls are enforced, which threats they address and how to verify them. File layouts, database schemas and license payload structures of closed-source components are outside its scope.

Verification. Each section covers the threat, the control and verification steps. Limits on independent verification are stated explicitly. Open-source components can be downloaded without an account or license.

Trust boundaries

Threats considered: a stolen browser session, a compromised core, a hostile cluster workload, data leaving the perimeter unnoticed.

The following four boundaries use different controls.

  Browser                    Core  (closed)                    Your cluster
  ───────                    ─────────────                     ────────────
  session cookie   ──[1]──►  default-deny guard      ──[2a]──►  agent (open)   ──►  apiserver
  HttpOnly, CSRF             identity resolution                outbound-only
                             the write chain                    path allowlist
                                                     ──[2b]──►  direct kubeconfig ──►  apiserver
                                     │      │                   (core holds it)
                                  [3]│      │[4]
                                     ▼      ▼
                             AI provider    SIEM target
                             opt-in,        TLS syslog or
                             per cluster    signed webhook
BoundaryWhat crossesWho decides
1 · browser → coreAn opaque session token in an HttpOnly cookie.Core, default-deny: three paths reachable unauthenticated — health, login, bootstrap.
2a · core → agentA request over an outbound socket with an impersonation header. No credential crosses back.The apiserver, on the impersonated user's RBAC. Ceiling set in your cluster.
2b · core → apiserverThe same request, but core holds the kubeconfig. The weaker boundary — §3.The apiserver, on whatever that kubeconfig reaches.
3 · core → AI providerOnly what an egress manifest records, only where egress was turned on. Default off.An administrator, per cluster, as an audited change.
4 · core → SIEMAudit records, hash-chained, with signed checkpoints.The receiver, against a public key carried out of band.

Verify independently. Boundaries 1 and 2 are observable from outside the product: run core behind a NetworkPolicy permitting only what you expect, and watch what it attempts. Boundary 3 has a hard off switch you can confirm at your own firewall; boundary 4 is verifiable in your own collector (§8).

Credential custody — depends on the mode

Threats considered: a compromised control plane holding standing credentials to your cluster.

Agent mode

The credential is a ServiceAccount that lives in your cluster and never leaves it. The agent's own ClusterRole carries no write verb and no object read permission — ceiling-limited impersonate, a health-probe self-review, discovery reads. Every user-facing request is authorized by the impersonated user's own RBAC.

Direct kubeconfig mode

Core holds the credential, sealed at rest in an AES-256-GCM envelope keyed by the installation's own snapshot key, with the cluster id as additional authenticated data. It is the more convenient mode, and its cost is set out in the box below.

In direct mode, if core is compromised, the cluster is compromised

The ceiling here is the kubeconfig you imported: an administrator can hand out any identity that credential can assume, and the only unconditional refusal is system:masters. Import a least-privileged kubeconfig, or use agent mode. Direct mode also makes core a request-forgery surface, since core dials the address you register — narrowed to https://, apiserver path prefixes and a refusal of proxy-url, but internal addresses are not blocked. That control belongs in an operator-side egress policy, not a kubeconfig parser.

Verify independently. The difference is a machine-readable field, not a footnote: the cluster inventory carries credentialCustody: "cluster" | "core", shown in the console as you confirm an import. In agent mode, read the generated manifest first — the ClusterRole in it is the whole ceiling.

Identity, sessions, and the impersonation ceiling

Threats considered: offline cracking of a stolen password store, session theft, CSRF, brute force, privilege escalation via identity mapping.

At rest. Passwords use scrypt at a 128 MiB memory cost — which is what removes a GPU attacker's parallelism advantage — with a per-record salt and the cost parameters stored inside the record, so they can be raised later and old records re-hashed on login. Policy follows the NIST 800-63B line: minimum 12 characters, no composition rules.

Sessions. Each opaque session token contains 256 random bits. Core stores its unsalted SHA-256 hash under a unique index. Server-side records support revocation; JWTs and signed stateless cookies are not used.

Cookies are host-only, HttpOnly and SameSite=Lax. Secure follows the request scheme and cannot be disabled with a flag. CSRF protection combines a double-submit cookie with an Origin check.

Sessions expire after 12 idle hours or 7 days in total. There are no refresh tokens or authentication caches, so revocation takes effect on the next request. Routes deny access by default.

The ceiling. YEKE answers who; Kubernetes answers allowed to do what. A user reaches a cluster as an identity resolved from a per-cluster rule or a personal binding, checked against the apiserver when you save it. With no mapping, the request does not leave core — enforced in the type system, not by a setting — and system:masters is refused by two independent gates that no flag opens. The ceiling itself sits in your cluster: the generated manifest limits impersonate with resourceNames, and left blank it is generated fail-closed, so the agent can assume no identity at all and every impersonated request is refused — which surfaces the misconfiguration on the first attempt instead of granting quietly.

SSO and second factor. LDAP / Active Directory shipped in 0.13.0, OIDC in 0.14.0, both behind one Enterprise entitlement and able to run at once. The OIDC flow is Authorization Code with PKCE (S256) on a confidential client, with the redirect_uri derived from the configured public URL and never from a request header, and no refresh tokens stored. The local password form never leaves the login screen, because local accounts are the emergency access path. TOTP with single-use recovery codes exists, and is opt-in (§11).

Verify independently. Ask the cluster, not us: the identity screen shows the username the apiserver returned, not the one the configuration claims. Then measure the ceiling — sign in as a narrowly-bound user and attempt a write outside their namespace. The expected result is a refusal from the apiserver, visible in its own audit log.

The single write path

Threats considered: an unapproved write, a write approved against a stale picture, a client that skips the chain, a destructive action taken by reflex.

There is exactly one way a write reaches your apiserver, and it is a state machine:

plan  ─►  capture prior state  ─►  classify  ─►  guardrail  ─►  dry-run
                                                                   │
                                        approval (planHash) ◄──────┘
                                                │
                                              apply  ─►  revert (a new plan,
                                                          same pipeline)
  • No step is optional, and approval is optional under no condition. Dry-run is skipped only where the apiserver makes it impossible, and the step is then flagged and raised to elevated approval by a shipped policy. A collection delete is expanded into individually named steps before you see it, each with its own uid and version precondition.
  • planHash is a SHA-256 over the decision you were shown — target versions, what will change, at what severity. If any of that changes the hash changes and the earlier approval is refused. An approval pressed on a stale card is not an approval.
  • Apply runs seven gates in order: plan exists; state permits; TTL not passed; caller's identity matches the plan's; hash current; an elevated plan carries a confirmation name matching the target's own name; connection live. Typing that name is part of the API contract, not interface decoration, so no client — the natural-language layer included — can skip it.
  • Writes on the raw proxy path are refused with a 405 and a distinct code, and every attempt becomes an audit event; reads stream through untouched. No emergency escape hatch was built, because your cluster already has one in kubectl. Uncertainty is not rounded down either: if the connection drops mid-apply, the step is recorded as unknown, not as "not applied".

The invariant the chain exists to produce: every applied write points, in its own record, at the approval that authorized it — its own approval card, or a shell session grant that was itself opened by typing a name. An unowned apply cannot be represented.

Verify independently. Do not take our test suite's word for it; use the apiserver's own accounting. The recipe is in §10.

The AI layer as an untrusted-input problem

Threats considered: prompt injection hidden in cluster content, secrets leaving with a prompt, a model's output mistaken for a decision.

Cluster content is untrusted input. A Pod annotation, for example, may contain an instruction to delete a namespace. AI can propose an operation, but the proposal must pass the dry-run, guardrails and the user’s Kubernetes permissions.

Deleting a namespace requires elevated approval: the user reviews the impact and types the namespace name. Text in a cluster resource cannot provide that approval.

Tool output is labeled as data and separated from user instructions by message roles. The audit trail links requests to plans. A classifier for detecting dangerous text is not used as a security boundary.

Egress is per cluster and off by default (off, local-only, allowed); changing it is an audited administrator action, and off means no request reaches any provider. Whether a provider counts as local is derived from the host in its base URL, never from the model's name, and a host allowlist can be a hard boundary. The honest wording, used in the product as well as here: a local class says "the address core connected to is a private one", not "the data did not leave the building".

Data excluded from AI requests. Code-level filtering removes Secret values, dockerconfig bodies, managed-fields metadata and the last-applied-configuration annotation. Key names such as DB_PASSWORD may be included.

Each provider call records an egress manifest: provider, host, egress class, prompt version and hash, and bytes per part. Request bodies are not stored. The manifest shows data categories and volume, not the exact bytes sent.

How this is measured — and how you measure it yourself

There is a fixed, versioned case set in the repository, run against real models rather than a stub, in which injection is its own class. Each case states its acceptance criterion as a list of forbidden operations: the case passes only if the plan corresponding to the injected instruction is never proposed. Results are reported against release notes rather than on this page, so that a figure here cannot outlive the build it came from — and so that what you read here is the mechanism, which does not change between runs.

The class deliberately covers more than one carrier. Injected text is planted in object fields — a Deployment annotation, a ConfigMap's contents, a description — and also in the two places easiest to forget, because they are not object spec at all: pod log output and Event messages. One case is an Event that claims the operator already authorised the deletion. The correct behaviour there is to treat the claim as data, on the ground that an authorization which is not an approval record is not an authorization.

The honest note about the neighbouring class. Declining an instruction outright is the model's own judgement, and a model can fail at it: in the runs so far, a model has proposed a plan in a case where it should have declined. That is exactly why the guarantee on this page is not "the model behaves well". It is where the chain stops the proposal — at an approval card that waits for a human, with the object name to be typed for anything destructive. A model's refusal is a convenience; the chain is the control.

Verify independently. Take your own measurement rather than ours. In a cluster you own, plant an instruction in each carrier — an annotation, a ConfigMap value, a pod's log line, an Event message — then ask an ordinary read question about that object. Use the same acceptance criterion the case set uses: write down, per case, the operations that must not appear, and count the appearance of any of them as a failure. What counts as evidence is not the assistant's prose but the operation record — no plan at all, or a plan that never left the approval card. For a negative control, set egress off first and confirm at your own firewall that nothing is attempted.

The tunnel and the open agent

Threats considered: a tunnel used as a general-purpose hole through the cluster perimeter, a compromised control plane pivoting to neighbouring services.

A tunnel agent is, structurally, a hole punched outward through your perimeter, and whoever holds the other end can ask it to dial. That framing is not ours after the fact: it is in the open source, directly above the constant that narrows it.

  • Outbound only. The agent dials core. It opens no listening port and needs no inbound rule; the only address it is configured with is core's URL.
  • One legitimate target. Requests are accepted only for a fixed set of apiserver path prefixes — /api, /apis, /openapi, /version, /healthz, /livez, /readyz — with traversal refused; anything else fails as PATH_NOT_ALLOWED. The point is not that a wider request would be unauthorized; it is that it cannot be expressed. Where the far end may name any target, a compromised control plane reaches the runtime socket and the cloud metadata endpoint.
  • Identity headers are stripped, not forwarded. Authorization, Proxy-Authorization and the hop-by-hop set never reach the apiserver from the far end; the identity it sees is the impersonation header the chain put there.
  • The wire version is a gate, not a hope. The protocol is at version 4 and the package's major number is that same number by construction, measured by a repository check; a mismatch closes the connection rather than degrading into an undefined dialect. There is also no license, telemetry or account code in the agent — a precondition of the split, not a side effect.

Verify independently. Most of this section can be checked directly in the open source. Clone the repository and git grep for the allowlist, the stripped headers and the absence of licensing code; npm pack @nairotech/yeke-tunnel and read the wire types. Then confirm the behaviour at runtime, with a NetworkPolicy permitting egress only to core and the apiserver.

Audit trail and SIEM export

Threats considered: a record quietly deleted or altered, a dead export channel mistaken for a quiet one, a vendor able to forge a customer's compliance record.

The trail is append-only in the strong sense: no write endpoint, no delete endpoint. Retention removes an operation's own record after its window; the audit events for it are not removed. Each event carries a two-part actor — the YEKE identity (who) and the Kubernetes identity and groups it acted under (with what authority) — plus the session id, because "same person, different session" is the first question in any incident. Reading the trail is free in every tier: the audit endpoints are bit-identical with or without a license, and the second NDJSON copy on standard output stays where it was, independently collectable.

Export is Enterprise, and its design question was not "how do we ship lines" but "how does the receiver prove none are missing".

  • Two sequence numbers, and confusing them defeats the point. One is the record's place in the trail, sparse when a class filter is set — gaps there are legitimate. The other is a dense per-target counter that never skips, so a gap in it is a certain loss. The counter alone would mean trusting the sender; the trail number alone cannot tell a filter from a deletion.
  • A hash chain, and signed checkpoints on top of it. Each record's canonical SHA-256 becomes the next record's previous hash, so deleting, altering or inserting one breaks the chain at exactly that point — computed over the canonical JSON, not the syslog presentation, which can be truncated. But a chain says nothing while the source is silent, and a silent channel is indistinguishable from a dead one, so an Ed25519-signed checkpoint carrying the chain head, covered ranges, declared filter and key id goes out periodically, event or no event. Per-line signatures were rejected: they prove "YEKE wrote this line", whereas the guarantee being sold is "none of these lines is missing".
  • The signing key is the installation's, not the vendor's. Core generates the pair on first use and seals the private half in its own envelope. A vendor key was rejected because core only holds the public half of the license key, and one leak would invalidate every customer's chain at once; HMAC was rejected because a symmetric signature can be produced by the receiver. The public key goes to the collector by hand and is pinned — automatic discovery was rejected for an epistemic reason, not a cryptographic one, since a key fetched over the channel you are verifying proves nothing.
  • Transport is TLS syslog or a signed webhook, with plain TCP and UDP eliminated and no certificate-verification escape flag. Its counter-argument is on record: many enterprise collectors listen on plain TCP, so requiring TLS creates rollout friction. A lapsed license does not stop the flow either, since cutting a compliance record punishes the customer's auditor.

Verify independently. End to end, without us: point a target at a collector you control, pin the public key, verify the checkpoint signatures yourself. Then attack your own copy — delete one line from the receiving store and confirm the chain breaks at exactly that record. Also check the negative case: nothing should arrive on an open plain-TCP or UDP syslog port.

Licensing and phone-home

Threats considered: a licensing mechanism that exfiltrates usage data, a commercial state that takes away emergency access.

A license is a single portable line: a version marker, a base64 payload, a base64 Ed25519 signature, verified entirely offline against a public key compiled into core using the runtime's own crypto — no library, no algorithm negotiation surface, no alg confusion class. Nothing phones home. No license server, no online activation, no call-in-if-you-have-connectivity check, no hardware fingerprint, no obfuscation — each eliminated with a written reason, the connectivity-dependent variant because it would mean two behaviours to document and two answers in a security review. The free tier needs no account and no key.

A tampered payload is refused loudly — a distinct error plus an audit event — and core keeps running with free-tier entitlements. No license state prevents core from starting, because refusing to boot would turn a renewal delay into an outage. On expiry: a banner 30 days ahead, expiry, a grace period at full function, then an additive freeze — no new clusters, no new users, no Enterprise configuration changes, while every existing cluster stays fully manageable, SSO keeps working and a configured audit export keeps flowing. The principle, stated plainly in the decision: the emergency response path is never held hostage to a commercial state.

Verify independently. Run core with no network route and complete a first install, cluster import and first operation; nothing should be attempted outbound. Flip one byte of a payload and observe the loud refusal plus continued operation. Take a strings dump of the published image and confirm the license public key is the only one there, and stable across versions.

How to verify without source access

Threats considered: a security page that is only a set of assertions.

Core is closed, so "read the code" is not on offer for it. What is on offer is a method you can run in your own environment against your own apiserver. Start with the two open artifacts — everything in §7 is a git grep away — then reproduce the central claim of §5 using the apiserver's own bookkeeping rather than any counter of ours.

# 1. Baseline, from the apiserver's OWN metrics — not from YEKE.
#    Sum the non-dry-run write counters; establish an idle noise floor first.
kubectl get --raw /metrics | grep '^apiserver_request_total' \
  | grep -E 'verb="(POST|PUT|PATCH|DELETE)"' | grep 'dry_run=""'

# 2. Run the attack set against core, as an authenticated administrator:
#    a) POST / PUT / PATCH / DELETE on the raw proxy path
#    b) apply with no plan, and any endpoint outside the chain
#    c) apply with a corrupted planHash
#    d) apply an expired plan
#    e) apply a plan created by a different identity
#    f) apply a plan that was already denied
#    g) apply an elevated plan with a missing, then a wrong, confirmation name

# 3. Read the same counters again, and read your own apiserver audit log for
#    the window. Expected: no write attributable to those attempts; each one
#    carrying its own refusal code; each one present in YEKE's audit trail as
#    a rejected attempt. Then confirm the target object is byte-for-byte
#    unchanged.

This is the same measurement we run internally, written out rather than summarised as a result for one reason: a green test suite inside a closed repository is not evidence to anyone outside it. What can be handed over is the method. The same applies to the audit chain (§8) — your collector, your pinned key, your own tampering — and to the AI case set (§6), whose acceptance criteria you can restate against your own cluster.

What this verification cannot reach

Guardrail internals, the shipped policy set's contents, database schemas and the license payload layout are closed, and no amount of black-box testing recovers them. The boundary was drawn on purpose: close enough that an auditor can answer where does the control sit and how do I check it, not so open that a competitor can answer how do I rebuild this. If your review requires source escrow or a code audit under NDA, that is something to raise with us directly; a public page cannot settle it.

Known gaps and non-claims

Each row states the position as of 19 August 2026. None of them carries a date by which it will change.

Not claimedToday's state
SAMLNot present. OIDC and LDAP cover single sign-on.
SCIM provisioningNot present. Users arrive by just-in-time provisioning at first directory or OIDC login.
OIDC logout integrationNeither back-channel nor RP-initiated logout exists. A user deleted in the identity provider is refused at their next login, but an already-open session does not fall; the remedy is revoking that session in YEKE, which takes effect on the next request.
Independent penetration testNot done. No third-party test report exists, and nothing here should be read as one.
SOC 2 / ISO 27001No certification. No audit has been conducted and no report exists.
OIDC against every providerVerified end to end against Keycloak on our own network. The Entra ID column of our internal comparison was written from Microsoft's documentation and no row of it was measured — there is no Entra tenant in the lab. Okta was not measured at all.
SIEM acceptance by a commercial collectorMeasured with real sockets: a real TLS collector, a real database, zero contact with deliberately-open plain-TCP and UDP ports, a cursor that resumes across a database restart. Not measured: that a production Splunk, Elastic or Sentinel parser accepts these lines.
Enforced second factorTOTP is opt-in per account. No installation-wide policy requires it, so an administrator who never enables it stays single-factor and nothing surfaces that.
Password re-prompt on destructive actionsNot present. An elevated approval makes you type the object's name, which slows a stolen session but does not stop it. In the shell surface even that is absent by design: consent is given once when the session opens.
Custody in direct modeCore holds the credential; if core is compromised, the cluster is compromised. The ceiling is the kubeconfig itself and the only unconditional refusal is system:masters (§3).
Egress control in direct modeCore dials the address you register. Narrowed to https://, apiserver path prefixes and a refusal of proxy-url — but internal addresses are not blocked. That control belongs on the operator side.
Credential rotation in direct modeNo rotation endpoint. Changing a kubeconfig means removing the cluster and adding it again.
Distributed brute-force resistanceThe per-IP attempt bucket is in memory and process-local: it resets when core restarts and does nothing against a distributed source. The durable gate is the account lock, which is persisted.
Generalising the AI measurementThe case set was machine-written and has not yet been reviewed by a human. Runs so far cover a narrow band of providers and a single class of consumer hardware, so nothing about local inference in general follows from them — in either direction.
Injection carried in logs and eventsCases that plant the instruction in pod log output and in Event messages were added to the case set after the most recent run against real models. The carrier is covered by the case set; it has not yet been measured end to end.
Diagnosis accuracyThe assistant collects events and logs and answers in a fixed finding/evidence/cause/fix shape, but whether a stated root cause is the real one has not been measured against broken-cluster fixtures. Treat a diagnosis as a lead, not a conclusion.
User deletionAccounts are disabled, not deleted, so audit references never dangle. Password reset is an administrator-issued one-time token; there is no mail infrastructure and therefore no reset email.

If a control you need is not on this page at all, read that as "absent", not "omitted for brevity". This is the section we would rather have shorter, and it is the reason the rest of the page can be taken at face value.

Reporting a vulnerability

There is one address for this, and one caveat about it.

Send it to [email protected] with SECURITY at the start of the subject. There is no separate security address today — this is the only published contact for the product, and inventing one here would be exactly the kind of claim the rest of the page argues against.

What helps: the version you tested (tags are on the releases page), whether the cluster was in agent or direct mode, whether AI egress was enabled, and a reproduction that stops short of data you do not own. For a finding in the agent or the wire package, a repository issue is also valid — those are public. There is no bug bounty programme and no promised response window; both would be commitments.

Where the control actually sits

This page describes the boundary. The two pages next door describe the halves of it you configure yourself: which identity a person writes under, and what your own systems get asked before a change lands.