Open-source SRE copilot — observability, FinOps, runbook automation, and incident response across Kubernetes and AWS / Azure / GCP.
Nudgebee is an open-source SRE copilot that watches your Kubernetes clusters and AWS / Azure / GCP accounts, turns raw signals into ranked findings, and walks operators through investigation and remediation. It bundles:
- Observability ingestion — Kubernetes events, metrics, traces, plus cloud-provider scans across AWS, Azure, and GCP.
- FinOps & cost optimization — surfaces unused / underutilized resources (idle workloads, oversized pods, stale snapshots, dangling volumes) and right-sizing recommendations across cloud and Kubernetes.
- LLM-powered triage — agentic planners that reproduce, root-cause, and propose fixes for incidents.
- ChatOps — Slack / Teams chatbot for SRE workflows: query state, run runbooks, ack alerts, and drive investigations from the channel where on-call already lives.
- Runbook automation — codifies recurring fixes as reusable runbooks and triggers them from chat, alert, or schedule.
- Ticketing + notifications — bidirectional sync with Jira, ServiceNow, PagerDuty, Zenduty; alert delivery to Slack, Teams, email.
Dashboard screenshot — to be added. Track discussion thread or contribute via a PR.
This is the fastest way to run Nudgebee from source. Infra in containers, backend and frontend from source on the host. This is the path contributors should use.
- Docker (with
docker compose) or Podman Desktop (withpodman-compose) - Go 1.26+
- Node 25+ and npm
git clone https://github.com/nudgebee/nudgebee.git
cd nudgebeedocker compose up -dThe default compose profile starts Postgres, Redis, RabbitMQ, Qdrant, Temporal, and a one-shot migrations container that applies the Postgres + RabbitMQ schema and then exits. Re-runs are safe — golang-migrate is idempotent against an up-to-date tracker. To also run the backend and frontend in containers (instead of from source), use docker compose --profile full up -d. The full profile mounts the host Docker socket into llm-server so it can launch an isolated code-analysis workspace container per account; access to that socket is equivalent to host-level Docker control. Workspace containers join the internal nudgebee-workspace network and do not publish host ports.
To connect a Kubernetes agent to the Compose relay and K8s collector, use the local agent configuration.
See api-server/migrations/README.md for how migration tracking works and how to add a new migration.
# macOS / Linux / WSL
cp api-server/services/.env.example api-server/services/.env
# Windows PowerShell
Copy-Item api-server\services\.env.example api-server\services\.envThen generate the encryption key and replace the __REPLACE__ placeholder for
NUDGEBEE_ENCRYPTION_KEY in the new .env:
openssl rand -hex 32Keep this value — you'll paste the same key into app/.env in step 5, and
into every other service's .env if you later run more from source (see
Local Stack Bootstrap below). Rotating it after data is written makes
previously-encrypted DB rows unreadable, so treat it like a database master
password.
Other defaults work as-is against the compose stack from step 2. Read the inline comments before any non-local deploy — a few other values (private keys) also need rotation.
# macOS / Linux / WSL (requires make)
cd api-server/services
make run
# Windows (no make required — runs the same command directly)
cd api-server\services
go run ./cmdListens on http://localhost:8000. Leave it running.
In a new terminal:
# macOS / Linux / WSL
cp app/.env.example app/.env
# Windows PowerShell
Copy-Item app\.env.example app\.envReplace __REPLACE__ for NUDGEBEE_ENCRYPTION_KEY with the same value you
generated in step 3. The app can't decrypt what services-server writes unless
these match.
The NEXTAUTH_SECRET in the example is a dev-only sample; rotate it for any
non-local deploy.
cd app
npm install --legacy-peer-deps
npm run devOpen http://localhost:3000.
On the sign-in page, click Admin Login. Then:
- Email: any address (e.g.
dev@example.com) — a tenant and an admin user are created automatically on first sign-in. - Password: literally
Test!24#5— the value ofNEXTAUTH_DUMMY_CREDS_PASSWORDshipped inapp/.env.example. Type it exactly; this is the dummy-credentials provider, not your own password.
The sample values in steps 3 and 5 above are fine for local dev. For any non-local deployment, generate fresh values and review the notes below.
| Var | Used by | How to generate | Notes |
|---|---|---|---|
APP_DATABASE_URL |
services-server | — | Compose default: postgres://postgres:postgrespassword@localhost:5432/nudgebee?sslmode=disable. Use localhost from the host, postgres hostname from inside the compose network. |
NUDGEBEE_ENCRYPTION_KEY |
services-server and app | openssl rand -hex 32 |
Encrypts integration credentials and other sensitive columns. Must match between services-server and app. Rotating it makes previously-encrypted rows unreadable — there is no automatic re-encryption migration. |
ACTION_API_SERVER_TOKEN |
services-server and app | openssl rand -hex 32 |
Optional. Shared secret for internal app↔services-server action calls. Defaults to empty on both sides, which disables the check (fine for local dev). If you set it, the value must match in both files. |
NEXTAUTH_SECRET |
app | openssl rand -base64 32 |
Signs both NextAuth session cookies and the inner HS256 session JWT (used by nbctl / Bearer-flow callers). Rotating it logs everyone out and invalidates outstanding bearer tokens. |
NEXTAUTH_DUMMY_CREDS_ENABLED / _PASSWORD |
app | — | Enables the any-email/password provider. Use for local development only; turn it off in any deployment exposed beyond your laptop. |
RABBIT_MQ_USERNAME / _PASSWORD / _HOST / _PORT |
services-server | — | Compose defaults: guest / guest / localhost / 5672. |
error pinging postgres: lookup postgres: no such hostfrom backend →APP_DATABASE_URLinapi-server/services/.envstill uses the container hostname. Replace@postgres:5432with@localhost:5432.migrate: error: pq: relation "..." already exists→ the tracker schema drifted from actual tables. InspectSELECT version, dirty FROM nudgebee.schema_migrations;and usemigrate force <version>to align. See api-server/migrations/README.md.- Action call from frontend returns 502
RPC gateway could not handle the operation→ the requested action isn't registered inapp/src/lib/actions.yaml, or it's a subscription / fragment / parse error. Check the frontend dev-server console for the unhandled reason.
The umbrella chart is published as a public OCI artifact at oci://ghcr.io/nudgebee/charts/nudgebee and bundles Postgres, RabbitMQ, Redis, Qdrant, and Temporal as subcharts.
# 1. Generate a permanent encryption key — store this securely.
# Losing it makes previously-encrypted DB rows unreadable.
export NUDGEBEE_ENC_KEY=$(openssl rand -hex 32)
echo "Save this key: $NUDGEBEE_ENC_KEY"
# 2. Install
helm install nudgebee oci://ghcr.io/nudgebee/charts/nudgebee \
--namespace nudgebee --create-namespace \
--set nudgebee_secret.NUDGEBEE_ENCRYPTION_KEY="$NUDGEBEE_ENC_KEY" \
--set admin.email="you@example.com" \
--set agent.enabled=true \
--wait --timeout 20madmin.email creates the admin and its organisation during install, so the first
login has nothing to set up. agent.enabled installs the NudgeBee agent alongside
the server and connects the cluster hosting it — without it, that first cluster
needs a separate agent install and an auth key copied out of the UI. Both are
optional; leave them out and the deployment provisions itself when you first sign
in, exactly as before. Additional clusters always use the normal agent install.
Two caveats for agent.enabled. The agent shares this Helm release, so
helm uninstall removes it too. And if you render offline (Argo CD, Flux), set
agent.accessKey and agent.accessSecret explicitly — the chart cannot read the
existing credential back and refuses to re-issue it on upgrade rather than
silently breaking the agent.
To pin a specific version, pass --version <X.Y.Z> (latest is used by default). To install from source instead — useful when iterating on chart changes — clone the repo, run helm dep update deploy/kubernetes/nudgebee, and point helm install at the local path.
The post-install hook applies database migrations automatically. Once the pods are ready:
kubectl -n nudgebee port-forward svc/app 3000:80
# Retrieve the bootstrap admin password
kubectl -n nudgebee get secret nudgebee \
-o jsonpath='{.data.NEXTAUTH_DUMMY_CREDS_PASSWORD}' | base64 -dOpen http://localhost:3000 and sign in as your admin.email with that password — the shared password works for that address only (or the licence address on a licensed install). See deploy/kubernetes/README.md for production-grade configuration (ingress, TLS, external Postgres, ClickHouse, observability sidecars).
The platform is live but empty. A quick tour that takes ~10 minutes:
- Connect a cloud account or a Kubernetes cluster — Settings → Integrations. Onboarding a real cluster lets the K8s collector populate the knowledge graph; an AWS / GCP / Azure account lights up the spend + recommendations surfaces. Without at least one of these, most of the dashboard is intentionally empty.
- Wire up a notification channel — Settings → Integrations → Slack / Teams / Email. Without one, you can still try the product, but notifications + ChatOps flows won't reach you.
- Run a sample runbook — Runbooks → Library. The bundled library has runnable examples (health-check loops, K8s investigation, cost spotlights). Run one manually to see end-to-end orchestration.
- Ask the AI assistant something — bottom-right corner of the dashboard. It can answer questions about your connected clusters, walk you through a recommendation, or kick off an investigation.
- Want to chat? Join us on Discord — async help, design discussions, contributor coordination.
- Hit a bug? Open an issue using the bug template.
- Idea for a feature? Use the feature template.
- Want to contribute code? Read CONTRIBUTING.md — covers CLA, branch model, PR conventions, and local-dev debugging tips. Look for issues tagged
good first issue. - Question that doesn't fit a template? Email
dev@nudgebee.comor use the contact links on the new-issue page.
Nudgebee is a Kubernetes-native monorepo of Go, Python, and TypeScript services. The high-level flow:
┌─────────────────────────────────────────────────┐
│ Browser (Next.js dashboard — `app/`) │
└────────────────────┬────────────────────────────┘
│ HTTP / WebSocket
▼
┌─────────────────────────────────────────────────┐
│ app (Next.js server) — RPC gateway + NextAuth │
└──┬─────────────┬──────────────┬────────────┬────┘
│ │ │ │
▼ ▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ api-server │ │ llm-server │ │ ticket-server│ │ ml-k8s- │
│ services │◀│ + rag-server │ │ notifications│ │ server │
│ (Go / Gin) │ │ + code- │ │ runbook │ │ (right- │
│ │ │ analysis │ │ │ │ sizing ML) │
└─┬──────┬─────┘ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘
│ │ │ │ │
│ └──────┬───────┴────────────────┴────────────────┘
│ ▼ ▼ ▼ ▼
│ Postgres RabbitMQ events Qdrant Temporal
│ (state) (cross-service) (vectors) (workflows)
│ ▲
│ proxy / │ publish + consume
│ control │
▼ ┌─────┴──────────────────────┐
┌──────────────┐ │ cloud-collector │
│ relay-server │ │ k8s-collector │
│ (in-cluster │ └────────────────────────────┘
│ agent gw) │
└──────┬───────┘
│ wss tunnels (also used by api-server / runbook to command agents)
▼
in-cluster agents
app/— Next.js dashboard; in-process RPC gateway at/api/graphqlforwards client calls to backend/rpc/*handlers.api-server/services/— core Go backend (Gin) for tenants, accounts, recommendations, integrations.llm/llm-server+llm/rag-server+llm/code-analysis— LLM session state, retrieval-augmented context, on-demand code analysis.ticket-server/— bidirectional sync with external ticketing (Jira, ServiceNow, PagerDuty, Zenduty).runbook-server/— runbook orchestration via Temporal workflows. See runbook-server/README.md for architecture, env reference, API, and task framework.notifications-server/— Slack/Teams/email alert delivery.collector-server/— cloud-scan + Kubernetes metrics pipelines;relay-serverbridges in-cluster agents back to the central plane.ml-k8s-server/— Python ML pipelines for workload right-sizing.
For per-service detail, see the Project Structure table below or each module's README.md.
The repo ships a docker-compose.yaml that wires every service against the public container registry. You rarely need to bring all 18 containers up — start only the services that match what you're working on. The table below lists each service and the minimum set of upstream services required for it to boot successfully.
| Service | Image | Notes |
|---|---|---|
postgres |
postgres:16 | Primary RDBMS. App schema applied by golang-migrate on deploy. |
rabbitmq |
rabbitmq:3-management | Message bus. UI at :15672. |
redis |
redis:7-alpine | Cache. |
qdrant |
qdrant/qdrant:v1.19.0 | Vector store for RAG / LLM. |
temporal |
temporalio/auto-setup:1.29.1 | Workflow engine. Backed by postgres (creates temporal + temporal_visibility DBs on first boot). Required by workflow-server. |
temporal-ui |
temporalio/ui:2.44.0 | Optional Temporal Web UI at :8233. |
| Service | Min upstream deps | Why |
|---|---|---|
api-server-services |
postgres, rabbitmq, redis |
Core backend. Won't bootstrap RabbitMQ consumers without RabbitMQ; queries fail without Postgres; cache pulls hit Redis. |
ticket-server |
postgres, rabbitmq |
DB writes + async ticketing sync. |
workflow-server (runbook-server) |
postgres, rabbitmq, temporal |
Workflow state in PG; events on RMQ; Temporal SDK calls Dial(7233) at startup and crashes if unreachable. |
notifications-server |
postgres, rabbitmq |
Persists messages, consumes RMQ events. |
cloud-collector |
rabbitmq |
Publishes scrape events to RMQ. |
relay-server |
postgres, rabbitmq |
K8s gateway; tunnel state in PG, events on RMQ. |
k8s-collector-app |
rabbitmq |
Publishes K8s metrics to RMQ. |
ml-k8s-server |
postgres |
Reads/writes scaling features. |
llm-server |
postgres, qdrant |
LLM session state + vector lookups. Spawns code-analysis per account on demand (not a long-running compose service). |
rag-server |
postgres, qdrant |
RAG retrieval against Qdrant; metadata in PG. |
llm-gateway |
postgres, api-server-services |
Not in docker-compose.yaml — run from source (llm/gateway, default port 8000). Backs the AI Gateway tab and Admin → AI & Tools → Gateway. Redis is optional (cache_provider defaults to in-memory). |
vulnerability-server |
none | Not in docker-compose.yaml — run from source (vulnerability-server, default port 8080). Stateless matcher; needs a local vulnerability DB on disk (~1.8 GB unpacked). api-server-services calls it to produce VM vulnerability findings. |
| Service | Min upstream deps | Why |
|---|---|---|
app (Next.js) |
api-server-services |
Client calls are served by the in-process RPC gateway in the Next.js server (/api/graphql), which fans out to six upstreams — api-server-services, llm-server, workflow-server, llm-gateway, ticket-server, notifications-server — plus a direct proxy to relay-server. Only api-server-services is needed to boot and sign in; the rest gate individual tabs (see the next section). RabbitMQ/Redis come in transitively via api-server-services' own deps. |
Baseline — every tab needs these three: postgres, api-server-services, app. They carry sign-in, tenant/account context, the left nav, and the RPC gateway itself. rabbitmq and redis come along transitively (api-server-services won't boot without them). The table lists only what a tab needs on top of that baseline.
| Nav section | Tab / surface | Extra services beyond the baseline | Notes |
|---|---|---|---|
| Home | Home, Account Overview | — | |
| Dashboards | Dashboard List, Application Grouping | — | |
| Troubleshoot | All Events (Triage Inbox, Events, group-by-type / -app, Triage Rules, Alert Tuning) | — | The Create ticket action on an event needs ticket-server; the Investigate button needs llm-server. |
Investigations (Auto / Manual Investigated), /investigate |
llm-server |
Add relay-server when an investigation pulls live cluster evidence (pod logs, live resources). |
|
| Event Resolutions | ticket-server |
Resolution rows create and comment on tickets. | |
| Knowledge Graph, Analytics | — | ||
Agent Health (/agentHealth) |
— | Agent connectivity is read from Postgres, so the page works with the collectors down — it just reports them down. | |
| Automations | Automations, Executions, Task Runner | workflow-server + temporal |
workflow-server dials Temporal on :7233 at startup and exits if it's unreachable. |
| Optimize | Summary, Cost, Configuration, Security, Resolutions | — | Reading is baseline-only. Generating the findings needs ml-k8s-server (right-sizing) and vulnerability-server (VM vulnerabilities) — see below. |
| Auto Optimize (Optimizations, Approvals) | workflow-server + temporal |
||
| LLM Analyser | llm-server |
Also gated by the per-tenant LLM_ANALYSER feature flag (Tenant Settings → Feature Flags). |
|
| AI Gateway | llm-gateway |
Also gated by UI_ENABLE_LLM_GATEWAY=true in app/.env. |
|
| Infra | K8s, Cloud, VM (list + detail) | — | VM → Vulnerabilities / Packages read stored matches; producing them needs vulnerability-server. |
| K8s detail: live resources, pod logs, pod shell, embedded Grafana | relay-server |
These proxy through /api/proxy/relay and /api/proxy/grafana, not the RPC gateway. |
|
| Tickets | All Tickets, Assigned to me | ticket-server |
|
| Admin | Access & Users, Tenant Settings | — | |
| Notification Rules | notifications-server |
Channel pickers and delivery-mode checks call it directly. | |
| Integrations | — | Test connection needs the owning service: notifications-server for Slack / Teams / Google Chat, ticket-server for Jira / ServiceNow / PagerDuty / Zenduty. |
|
| AI & Tools → Agents, Tools & MCP, Functions, Providers, Egress Filter, Memory Policy, RCA Format | llm-server |
||
| AI & Tools → Gateway, Budgets & Limits | llm-gateway (Gateway), llm-server (Budgets & Limits) |
||
| Nubi | Assistant panel (header) and /ask-nudgebee |
llm-server |
Knowledge-base answers additionally need rag-server + qdrant; llm-server spawns code-analysis on demand. |
What a missing service looks like. app never crashes on an unreachable upstream — the page renders and only the calls that need that service fail:
- Service down, env var set → GraphQL response carries
Upstream unreachable for <action>;/api/rpcanswers 502. - Env var unset (e.g. no
LLM_GATEWAY_URL) →Handler URL unresolved for <action>and a 500. The dev-server console logs the missing variable name.
Both come out of app/src/lib/rpcGateway.ts and app/src/pages/api/rpc.ts. /status in the running app probes API, Data Ingest (relay), Nubi AI, Automation (workflow-server + Temporal) and Notifications, which is the fastest way to see what's actually up.
Services that supply data rather than gate a tab. cloud-collector, k8s-collector-app, relay-server, ml-k8s-server and vulnerability-server don't sit on the request path for most tabs — they populate the tables those tabs read. Without them the tabs render empty, not broken. ml-k8s-server owns right-sizing (krr_scan), unused-volume analysis and anomaly detection, so Optimize → Cost stays empty until it has run; vulnerability-server matches VM package inventories against the vulnerability database, and api-server-services calls it via VULN_MATCHER_SERVER_ENDPOINT.
Deriving this yourself. app/src/lib/actions.yaml is the source of truth: each action's handler: names the upstream via a {{SERVICE_URL}} placeholder. The six the gateway can route to are SERVICE_API_SERVER_URL, LLM_SERVER_URL, WORKFLOW_SERVER_URL, LLM_GATEWAY_URL, TICKET_SERVICE_URL and NOTIFICATION_SERVICE_URL. A tab's calls live under app/src/api1/<domain>/ — cross-reference the action names in that directory against actions.yaml to get the tab's service set.
docker compose publishes several services on different host ports than app/.env.example defaults to. Running the frontend from source against the compose stack means overriding these in app/.env:
| Service | Compose host port | app/.env.example default |
|---|---|---|
api-server-services |
8000 | SERVICE_API_SERVER_URL — 8000 ✅ |
llm-server |
8005 | LLM_SERVER_URL — 8005 ✅ |
ticket-server |
8001 | TICKET_SERVICE_URL — 8004 ❌ |
cloud-collector |
8002 | CLOUD_COLLECTOR_SERVER_URL — 8001 ❌ |
relay-server |
8004 | RELAY_SERVER_ENDPOINT — 8006 ❌ |
workflow-server |
8007 | WORKFLOW_SERVER_URL — 8002 ❌ |
notifications-server |
8090 | NOTIFICATION_SERVICE_URL — 8003 ❌ |
llm-gateway |
not in compose | LLM_GATEWAY_URL — 8007 (collides with compose workflow-server) |
Inside compose this doesn't bite — the full profile wires every upstream by service name through the x-app-common anchor at the top of docker-compose.yaml. The one exception is LLM_GATEWAY_URL: llm-gateway has no compose service, so the AI Gateway tab has no upstream in a pure-compose stack.
- DB exploration:
postgres(connect withpsql/ DBeaver). - Login + dashboard render:
postgres+api-server-services+app(pulls RabbitMQ/Redis transitively). This is the baseline every tab needs. - Cloud findings pipeline: add
rabbitmq+cloud-collector. - Recommendation generation (Optimize → Cost): add
ml-k8s-server. - LLM/RAG flows + Nubi: add
qdrant+llm-server+rag-server. - Automations / Auto Optimize: add
temporal+workflow-server. - Tickets: add
ticket-server. Notification Rules: addnotifications-server. - Live K8s (pod logs, pod shell, embedded Grafana): add
relay-server.
Start a subset with docker compose up -d <service> [<service> ...] (or podman-compose up -d ...); transitive deps are pulled in automatically via depends_on.
Each module has its own README with setup and development instructions.
| Module | Description | README |
|---|---|---|
app/ |
Frontend dashboard (Next.js + React, NextAuth) | app/README.md |
| Module | Description | README |
|---|---|---|
api-server/ |
GraphQL API layer overview | api-server/README.md |
api-server/services/ |
Core backend Go services (Gin) | api-server/services/README.md |
api-server/migrations/ |
DB migrations (Postgres via golang-migrate, ClickHouse, RabbitMQ) | api-server/migrations/README.md |
| Module | Description | README |
|---|---|---|
collector-server/cloud-collector/ |
Cloud data collection (AWS / Azure / GCP) | collector-server/cloud-collector/README.md |
collector-server/k8s-collector/app/ |
K8s metrics aggregation (Python) | collector-server/k8s-collector/app/README.md |
collector-server/k8s-collector/relay-server/ |
K8s relay gateway (WebSocket) | collector-server/k8s-collector/relay-server/README.md |
| Module | Description | README |
|---|---|---|
ml-k8s-server/ |
ML models & K8s autoscaling | ml-k8s-server/README.md |
llm/llm-server/ |
LLM inference service | llm/llm-server/README.md |
llm/code-analysis/ |
Code analysis engine | llm/code-analysis/README.md |
llm/rag-server/ |
RAG (Retrieval Augmented Generation) | llm/rag-server/README.md |
llm/benchmark/ |
LLM benchmarking | llm/benchmark/README.md |
| Module | Description | README |
|---|---|---|
runbook-server/ |
Runbook orchestration + automation engine (Temporal) | runbook-server/README.md |
ticket-server/ |
External ticketing integration (Jira, ServiceNow, PagerDuty, Zenduty) | ticket-server/README.md |
notifications-server/ |
Notification delivery (Slack, Teams, email) | notifications-server/README.md |
| Module | Description | README |
|---|---|---|
deploy/kubernetes/ |
Helm charts & Kubernetes config files | deploy/kubernetes/README.md |
| Module | Description |
|---|---|
app-e2e-tests/ |
End-to-end integration tests |
We welcome contributions! Before opening your first PR:
- Read CONTRIBUTING.md for the development workflow, conventional-commit format, and PR guidelines.
- Review the Code of Conduct.
- Browse open issues — look for
good first issueandhelp wantedlabels if you're getting started.
By contributing, you agree that your contributions are licensed under the Business Source License 1.1 (and its Change License, Apache 2.0). On your first PR, the CLA Assistant bot will post a one-click sign link — subsequent PRs need no further action.
If you believe you have found a security vulnerability in Nudgebee, please do not open a public GitHub issue. Instead, follow the responsible disclosure process in SECURITY.md.
Nudgebee ships with no telemetry or product analytics. No data leaves your cluster except what you explicitly configure — notification webhooks, ticket-system sync, LLM provider calls, and any outbound integrations you wire up.
- Chat / questions / contributor coordination — Discord
- Bugs and feature requests — GitHub Issues
- Security reports — SECURITY.md
Nudgebee is source-available under the Business Source License 1.1. You can self-host and run it in production for your own internal purposes for free; offering it to third parties as a hosted/managed service, or using it to deliver services to your customers, requires a commercial license. Each version converts to Apache 2.0 four years after release. See LICENSING.md for a plain-language summary, or contact licensing@nudgebee.com.