Kobi Copilot
Kobi's assisted mode: talk to your cluster in natural language. Press ⌘J to open.
Kobi Copilot is the assisted mode of Kobi, KubeBolt’s AI SRE: you ask, Kobi answers. It ships in every edition — in the open-source core you bring your own LLM key (see Enabling Kobi).
How it Works
The copilot sends your question to the configured LLM provider along with the definitions of 39 tools. The backend runs every tool call itself, through the same cluster connection the rest of KubeBolt uses — the LLM never talks to your cluster directly. Kobi analyzes the results and answers with data-backed findings and, where useful, kubectl commands.
Beyond diagnosis, Kobi can propose the fix: restart, scale, roll back, set image, and more. A proposal changes nothing by itself — it runs only when you click Execute, under your KubeBolt role (Editor or Admin), after a server-side dry-run preview (every action except rollback), and every executed action is recorded in the audit trail. See Actions & Governance.
Kobi also answers when no cluster is connected: it still talks about Kubernetes and KubeBolt, and tells you there is no live cluster to inspect.
30 read tools
| Tool | What it does |
|---|---|
get_cluster_overview | Resource counts, CPU/memory, health score, events |
list_resources | List any of 23 resource types (workloads, networking, storage, config, nodes, namespaces, events, Cilium policies) with filtering |
get_resource_detail | Full detail of a specific resource |
get_resource_yaml | Raw YAML definition (secrets redacted) |
get_resource_describe | kubectl-describe output for deep troubleshooting |
get_pod_logs | Pod logs with container, tail, since and grep options |
get_workload_pods | Pods owned by a workload controller |
get_workload_history | Revision history for Deployments/StatefulSets/DaemonSets |
get_workload_metrics | CPU / memory / network over a time range for a workload, pod, or node, and a node’s disk fill |
get_cronjob_jobs | Job children of a CronJob to investigate execution history |
get_topology | Full cluster topology graph |
get_insights | Active insights with severity |
get_events | Events with filtering |
search_resources | Global search by name across 24 resource types |
get_permissions | Detected RBAC permissions |
list_clusters | All available clusters |
get_kubebolt_docs | Product knowledge base (features, navigation, admin pages) |
get_insight_episodes | History of insights: episodes that resolved, expired or are still firing, how each ended and how often it came back |
get_insight_episode | One episode in full: its append-only timeline, who silenced what and when, and the typed evidence the rule recorded |
get_operational_episodes | Bursts of insights that fired together, already classified by shared cause (node rotation, rollout, and so on) |
get_recent_deploys | Deployment rollouts of the last hours, newest first: which workload got a new ReplicaSet, when, with which image |
get_runtime_events | Runtime security events from Falco: a shell spawned in a container, a sensitive file read, an unexpected outbound connection |
get_findings | Security posture: CVEs and misconfigurations from Trivy, policy violations from Kyverno, CIS controls |
get_finding_detail | One finding re-read live from the scanner, with every package carrying a CVE and its fixed version |
get_finding_workloads | Findings grouped by workload and ranked to work through: exposed secrets first, then by severity |
get_right_sizing | Right-sizing recommendations from the same deterministic engine the Capacity and Cost screens read |
get_coverage | What KubeBolt can see right now: how the cluster is connected, which agents report, which metric sources answer |
query_metrics | A PromQL query against KubeBolt’s metrics store, confined to your organization and the current cluster |
get_fleet_summary | Active insights per cluster across the whole organization, worst first. The only tool that sees beyond the selected cluster |
offer_cluster_switch | Offers a one-click switch to another cluster, carrying the question so the new conversation opens already asking it |
Node disk. For DiskPressure, ephemeral-storage evictions or “the disk
filled up”, Kobi reads the node’s disk with get_workload_metrics
(kind=Node, metric=filesystem): the fullest disk at each point, with
perMountpoint naming which one when node-exporter reports several mountpoints.
Each point is the peak of its step, so a disk that filled for a few minutes is
not averaged away, and the chat draws the result as a Disk card. Asking for
a pod’s or workload’s disk is refused rather than approximated — a pod’s disk
usage is not the node’s — and a metric that was not measured comes back with an
error instead of an empty series that would read as idle. Pod-level disk IO is
not exposed; PVC fill is read with query_metrics
(kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes).
9 action proposals
propose_restart_workload, propose_debug_pod, propose_scale_workload,
propose_rollback_deployment, propose_delete_resource,
propose_set_resources, propose_set_image, propose_set_env,
propose_patch_hpa. Each renders an action card in the chat; nothing happens
until you approve it. Admins can withhold all of them
(KUBEBOLT_AI_ACTIONS_ENABLED=false) or only the destructive ones
(KUBEBOLT_AI_DESTRUCTIVE_ACTIONS_ENABLED=false). Details in
Actions & Governance.
Supported Providers
KubeBolt ships two provider adapters, and KUBEBOLT_AI_PROVIDER accepts only these two values:
anthropic— Claude Sonnet 5 (default), Claude Opus 5, Claude Fable 5, Claude Haiku 4.5. Prompt caching on system prompt + tool definitions.openai— OpenAI itself (GPT-5.6 Sol/Terra/Luna, GPT-5.5, GPT-5.4, GPT-4o as the default, GPT-4o Mini) and any OpenAI-compatible Chat Completions API: Azure OpenAI, xAI, DeepSeek, Groq, Mistral, OpenRouter, and self-hosted runners such as Ollama or vLLM.
For anything other than OpenAI itself, set the base URL to the provider’s full chat-completions endpoint (for example https://api.x.ai/v1/chat/completions). KubeBolt posts to it as-is and never appends a path. The model must support tool calling.
Bring your own key. In the open-source edition you configure your own provider key and pay that provider directly; KubeBolt has no AI billing. The backend proxies every LLM request, so the key never reaches the browser, and a key saved from the UI is encrypted at rest. On KubeBolt Cloud the AI is managed — there are no keys to configure, and usage draws from your plan credits.
MCP server (read-only)
Twenty-nine of the read tools are also exposed over the Model Context Protocol, so any MCP host — Claude Code, Cursor, a CI step, another agent — can investigate a live cluster using its own LLM. No KubeBolt AI key is needed: the host’s model does the reasoning, KubeBolt only serves the tools.
- Remote, over HTTP — the server publishes
POST /api/v1/mcpusing the standard Streamable HTTP transport, behind normal KubeBolt authentication with an API token. The token identifies the tenant; sendX-KubeBolt-Cluster: <context-name>(a name fromGET /api/v1/clusters) to target a specific cluster instead of the active one. - Local, over stdio — the standalone
kubebolt-mcpbinary speaks MCP over stdin/stdout against your kubeconfig directly, with no server and no auth. Download it from the GitHub release assets, or build it withmake build-mcp.
Read-only is enforced server-side: the 9 propose_* tools are neither listed
nor callable, even by name. The thirtieth read tool, offer_cluster_switch,
is withheld too: it draws a card only Kobi’s own panel knows how to render, so
to any other host it would be a tool that does nothing.
Which token to use
| Your MCP host reaches KubeBolt through… | Token | Scopes |
|---|---|---|
| The UI URL of a Helm or Docker Compose install (the bundled nginx — the usual case) | API key (kbk_) | Must include /api/v1/mcp or * |
The API directly — the single-container image, or the chart’s <release>-api Service from inside your network | API key or service token (kbs_) | A kbs_ token’s default scopes already include /api/v1/mcp |
An API key has no default scopes. Pick MCP only (read-only tools) in
the create dialog: the key reaches /api/v1/mcp and nothing else, its role is
pinned to viewer on the server, and it expires in 90 days unless you change
it. Tick the clusters it may read in the same dialog, or leave the list empty
for every cluster. Through the API it is the same thing with
"scopes":["/api/v1/mcp"].
Service tokens don’t work through the UI URL. The bundled nginx marks
every request it proxies with X-KubeBolt-Edge: public, and the API rejects
kbs_ service tokens on those requests with 401 invalid or expired token.
That is deliberate: a leaked service token is useless from the internet. If
you want a kbs_ token, point the host at the API directly (for example
http://<release>-api.<namespace>.svc:8080/api/v1/mcp, or a
kubectl port-forward svc/<release>-api 8080:8080).
initialize and tools/list answer even while the cluster is momentarily
disconnected; a tools/call in that window returns an isError result that
says live cluster access is needed, instead of failing the session.
get_kubebolt_docs keeps working, since it needs no cluster.
Contextual “Ask Kobi”
One-click Ask Kobi buttons open the panel with a prompt already loaded with the cluster, namespace, resource and symptom. Surfaces include:
- Insight cards — diagnose the insight and recommend a fix
- Resource detail header — investigate the resource; the prompt adapts to the active tab (Monitor, Logs, YAML…)
- Workload health on the Overview — why a workload is not ready
- Warning events — explain the event and its impact
- Traffic flows on the Cluster Map — explain the flow, drops and error rates
- Dashboard panels — top CPU consumers, right-sizing, recent deploys, error hot-spots, top latency, network drops and more, for the whole panel or a single row
- Stalled actions — when an approved action doesn’t converge, Kobi is asked why
Conversations
Conversations are personal and persist per user, so you can refresh, sign out and resume where you left off. The History button in the panel header lists your past conversations; New conversation starts a fresh one. Persistence needs authentication enabled (the default). Retention and the per-user cap are set by KUBEBOLT_COPILOT_CONVERSATION_RETENTION_HORIZON (default 2160h) and KUBEBOLT_COPILOT_CONVERSATION_MAX_PER_USER (default 200).
Conversation memory
Long sessions stop bleeding context. When the estimated conversation size crosses SESSION_BUDGET_TOKENS × AUTO_COMPACT_THRESHOLD (default 80%), the handler folds older turns into a summary generated by the provider’s cheap-tier model (claude-haiku-4-5 / gpt-4o-mini) and stubs bulky tool_results in the preserved tail. The active turn’s tool_results are always protected so mid-flight compacts never truncate a response.
A Scissors icon in the panel header exposes the same primitive on demand — “new session with summary” collapses the whole transcript into a single summary message so you can pivot topics without losing context.
| Env var | Default | Purpose |
|---|---|---|
KUBEBOLT_AI_AUTO_COMPACT | true | Master switch for auto-compaction |
KUBEBOLT_AI_SESSION_BUDGET_TOKENS | context window of the model | Total ceiling. Trigger fires at budget × threshold |
KUBEBOLT_AI_AUTO_COMPACT_THRESHOLD | 0.80 | Fraction of the budget at which compact fires |
KUBEBOLT_AI_COMPACT_MODEL | claude-haiku-4-5 / gpt-4o-mini | Override the summarisation model — required on any OpenAI-compatible endpoint other than OpenAI itself, since gpt-4o-mini only exists there |
KUBEBOLT_AI_COMPACT_PRESERVE_TURNS | 3 | Turns kept intact after compaction |
Scope guardrail
The system prompt defines in-scope (Kubernetes operations, DevOps/SRE topics that support the user’s cluster, KubeBolt itself) and out-of-scope (general coding unrelated to Kubernetes, non-technical topics, cloud products not running on Kubernetes). The LLM refuses out-of-scope questions with a one-sentence polite redirect in the user’s language — never answers partially.
Usage analytics
Administration → AI (Kobi) → Usage (Admin role): sessions, tokens billed, cache hit rate, estimated USD cost (a best-effort price table — your real bill is whatever your provider charges; unknown models show no estimate), top tools with error rates, and a per-session drill-down with tool breakdown and compaction events. Range selector 24h / 7d / 30d. Stored locally in BoltDB with a 30-day / 5,000-record retention cap. Requires authentication to be enabled.