🧠 Cerebro β€” Platform & Security 🌐 Loom β€” Infrastructure & Deploy 🧢 Weave β€” Plugins & Knowledge ⚑ Workflows πŸ” Identity πŸ“Š Presentation

🌐 Loom β€” One Command to Deploy, Upgrade, or Tear Down on Any Cloud

Terraform provisions the infrastructure Β· Helm packages the application Β· Kubernetes runs the workloads Β· Same code, any cloud

πŸ—οΈ Terraform Infrastructure as Code πŸ“‹ Plan terraform plan πŸš€ Apply terraform apply ⬆️ Upgrade helm upgrade πŸ’₯ Teardown terraform destroy TERRAFORM MODULES β€” Cloud-Agnostic Building Blocks πŸ†” Identity Managed ID / IAM ☸️ Kubernetes AKS / EKS / GKE πŸ’Ύ Database Cosmos / Dynamo / Fire πŸ” Search AI Search / OpenSearch πŸ€– LLM OpenAI / Bedrock / Vertex πŸ” Vault Key Vault / Secrets Mgr πŸ“¦ Registry ACR / ECR / GAR 🌐 Networking VNet / VPC / Ingress 🌍 DNS Azure DNS / R53 πŸ“§ Messaging ACS / SES / SendGrid πŸ”’ Private EP PE + DNS / VPC EP HELM CHARTS β€” Application Packaging & Deployment ⎈ cerebro-app Deployment Β· Service Β· Ingress Β· HPA ConfigMaps Β· Secrets Β· ServiceAccount Β· RBAC values-azure.yaml Β· values-aws.yaml Β· values-local.yaml ⎈ mcp-engine Deployment Β· Service Β· Namespace Β· HPA Plugin volume mounts Β· Liveness/Readiness probes Per-engine Helm release: mcp1, mcp2, mcp-healthcare... ⎈ monitoring Prometheus Β· Grafana Β· Loki Β· Promtail ServiceMonitors Β· AlertRules Β· Dashboards kube-prometheus-stack + loki + promtail charts KUBERNETES CLUSTER β€” Runtime Workloads ☸ Kubernetes Cluster πŸ“¦ cerebro namespace 🧠 cerebro-app Flask Β· HPA 2–8 Β· Gunicorn βš›οΈ React HPA 2–5 οΏ½οΏ½ Valkey Sessions Services: cerebro-app:5080 Β· Ingress: app.brainzbytes.com (Let's Encrypt TLS) πŸ”‘ Secrets (env vars) πŸ“„ ConfigMaps πŸ” cert-manager Β· contextweaver-tls (Let's Encrypt auto-renew) πŸ“¦ mcp1 namespace πŸ”Œ mcp-engine FastMCP Β· Dashboard Β· Plugins πŸ”Œ mcp2 (future) Additional engines Services: mcp1:5000 (dashboard) Β· mcp1:5001 (MCP transport) πŸ”‘ Engine Secrets πŸ“š Plugin Volumes πŸ”— ServiceMonitor β†’ Prometheus scraping πŸ“¦ monitoring namespace Prometheus Grafana Loki Promtail Alertmanager Β· node-exporter kube-state-metrics Dashboards: Platform Overview Β· MCP Engine Β· RAG Β· Security Β· User Activity Β· Loki Logs 🚨 Alert Rules: High error rate Β· Slow tools Β· Pod crashes Β· Disk full Β· Memory pressure πŸ“¦ workspaces namespace πŸ’» ws-{username} code-server Β· 4C/4Gi πŸ› οΈ Copilot + Claude Azure Disk PVC Auth-gated Β· zsh + Oh-My-Zsh Β· kubectl/helm/git πŸ“¦ jupyterhub namespace 🏠 JupyterHub Multi-user hub πŸ‘€ jupyter-{user} Per-user pods Entra ID auth Β· Python/R/Data Science kernels πŸ“¦ ingress-nginx namespace 🌐 NGINX Controller πŸ” TLS / cert-mgr πŸ”’ Auth-URL Let's Encrypt auto-renew Β· SSL redirect Β· per-path auth gating CLOUD PROVIDER β€” One terraform apply away ☁️ Azure CURRENT β€” Production AKS Β· Cosmos DB Β· AI Search Azure OpenAI Β· Key Vault Β· ACR Β· ACS Managed Identity β€” zero keys in pods ☁️ AWS PLANNED EKS Β· DynamoDB Β· OpenSearch Bedrock Β· Secrets Manager Β· ECR IAM Roles for Service Accounts ☁️ GCP PLANNED GKE Β· Firestore Β· Vertex Search Vertex AI Β· Secret Mgr Β· Artifact Reg Workload Identity Federation 🏠 Local / On-Prem AVAILABLE Docker Β· SQLite Β· Ollama File vault Β· docker-compose Air-gapped Β· No cloud needed ⇄ ⇄ ⇄
Terraform (Provision + Lifecycle)
Helm (Package + Deploy)
Kubernetes (Runtime)
Upgrade Path
Teardown
Cloud Providers
βœ•

πŸ—οΈ Terraform β€” Infrastructure as Code

What it does: Declares the entire cloud infrastructure in code files. Run one command β†’ everything exists. Run another β†’ everything is gone.

Directory: terraform/ β€” 9 modules, 3+ environment configs

Key commands:

terraform plan β€” Preview what will change (safe, read-only)

terraform apply β€” Create or update all infrastructure

terraform destroy β€” Tear everything down cleanly

State: Local backend (terraform.tfstate on deployment machine). Previously Azure Blob β€” migrated to local for simplicity since deployments run from a single PC.

βœ•

πŸ“‹ Plan Phase

terraform plan β€” Shows exactly what will be created, modified, or destroyed before making any changes.

Like a construction blueprint review β€” you see every change before the builders start. No surprises.

Output example: "Plan: 23 to add, 2 to change, 0 to destroy"

βœ•

πŸš€ Apply Phase

terraform apply β€” Creates all infrastructure: Kubernetes cluster, databases, search service, vault, messaging, networking, DNS, private endpoints.

Then helm install deploys the application workloads into the cluster.

Full stack from zero: One terraform apply + helm install = entire platform running.

βœ•

⬆️ Upgrade Phase

helm upgrade cerebro-app ./helm/cerebro-app β€” Rolling update, zero downtime.

What can be upgraded:

β†’ Application code (new Docker image)

β†’ Configuration (env vars, secrets, replicas)

β†’ Infrastructure (Terraform detects drift and reconciles)

β†’ Plugins (hot-reload via MCP engine API)

Rollback: helm rollback cerebro-app 1 β€” instant revert to previous version.

βœ•

πŸ’₯ Teardown Phase

terraform destroy β€” Removes ALL infrastructure cleanly. Every resource, every service, every DNS record.

Useful for: dev/test environments, cost optimization, migration between clouds.

Safety: Requires explicit confirmation. State file records what was created, so destroy is precise β€” no orphaned resources.

βœ•

🧱 Terraform Modules

Each module abstracts one cloud service into a cloud-agnostic interface:

πŸ†” Identity: Azure Managed Identity / AWS IAM Roles / GCP Workload Identity

☸️ Kubernetes: AKS / EKS / GKE β€” same K8s API, different provisioning

πŸ’Ύ Database: Cosmos DB / DynamoDB / Firestore

πŸ” Search: Azure AI Search / OpenSearch / Vertex AI Search

πŸ€– LLM: Azure OpenAI / AWS Bedrock / Vertex AI

πŸ” Vault: Azure Key Vault / AWS Secrets Manager / GCP Secret Manager

πŸ“¦ Registry: ACR / ECR / Artifact Registry

πŸ“§ Messaging: Azure Communication Services / Amazon SES / SendGrid

🌐 Networking: VNet / VPC / Ingress controllers

🌍 DNS: Azure DNS / Route 53 / Cloud DNS

πŸ”’ Private Endpoints: Azure PE + Private DNS / AWS VPC Endpoints / GCP Private Service Connect β€” zero public access to data services

The app code never references cloud-specific APIs β€” only the Terraform modules change per cloud.

βœ•

⎈ Helm Charts

Helm packages Kubernetes manifests into versioned, configurable releases:

cerebro-app: The main platform β€” Flask app, ingress, TLS, HPA autoscaling, service account with workload identity.

mcp-engine: Each MCP engine gets its own Helm release in its own namespace. Deploy new engines by adding a release.

monitoring: Full observability stack β€” Prometheus, Grafana, Loki, Promtail, AlertManager, dashboards.

Values files: values-azure.yaml, values-aws.yaml, values-local.yaml β€” same chart, different cloud configs.

βœ•

πŸ“¦ cerebro namespace

Runs the main Cerebro app β€” the Flask web server, OIDC + Magic Link auth, agent loop, Visual Designer, all API endpoints.

Pods: cerebro-app (HPA 2-8, Gunicorn), contextweaver-frontend (React SPA, HPA 2-5), Valkey (sessions + rate limiting), code-server (standalone workspace)

TLS: cert-manager auto-provisions Let's Encrypt certificates

Secrets: Cosmos endpoint, search keys, LLM keys, SSO secret β€” all injected via Helm values

βœ•

πŸ“¦ mcp1 namespace

Each MCP engine runs in its own Kubernetes namespace for isolation. Deploy additional engines (mcp2, mcp-healthcare, mcp-finance) as separate Helm releases.

Ports: 5000 (dashboard/API), 5001 (MCP transport β€” SSE/HTTP)

Plugins: Mounted as volumes or uploaded via API. Each plugin is a ZIP with manifest + Python code.

Scaling: Each engine scales independently based on load.

βœ•

πŸ“¦ monitoring namespace

Full observability stack running alongside the application:

Prometheus: Scrapes metrics from all pods via ServiceMonitors

Grafana: Dashboards for platform, MCP, RAG, security, users, Loki logs

Loki + Promtail: Aggregates all pod logs, searchable with LogQL

Alertmanager: Fires alerts on high error rates, slow tools, pod crashes, disk/memory pressure

βœ•

☁️ Azure β€” Current Production

Fully wired and running. cd terraform/environments/azure && terraform apply

AKS (D4as_v5 nodes), Cosmos DB, Azure AI Search, Azure OpenAI (GPT-5), Key Vault, ACR, ACS Email, Private Endpoints

Security: Managed Identity for all service-to-service auth β€” zero API keys stored in pods.

βœ•

☁️ AWS β€” Planned

Same Terraform modules, AWS provider. cd terraform/environments/aws && terraform apply

EKS, DynamoDB, OpenSearch, Bedrock (Claude/Titan), Secrets Manager, ECR, SES Email, VPC Endpoints

Security: IAM Roles for Service Accounts (IRSA) β€” same zero-key pattern as Azure.

βœ•

☁️ GCP β€” Planned

Same Terraform modules, GCP provider. cd terraform/environments/gcp && terraform apply

GKE, Firestore, Vertex AI Search, Vertex AI (Gemini), Secret Manager, Artifact Registry, SendGrid, Private Service Connect

Security: Workload Identity Federation β€” same zero-key pattern.

βœ•

🏠 Local / On-Premises

For development, demos, and air-gapped deployments. docker-compose up

Docker Compose, SQLite (no Cosmos needed), Ollama (free local LLM), file-based vault

No cloud account required. Everything runs on a single machine.

☸️ Production AKS Topology β€” Current State

contextweaver-aks Β· 2–6 nodes (Standard_D4as_v5) Β· Cluster Autoscaler Β· HPA Β· Valkey

🧠 cerebro namespace

⚑ cerebro-app β€” Flask backend 2–8 replicas (HPA)
πŸ–₯️ contextweaver-frontend β€” React SPA 2–5 replicas (HPA)
πŸ”΄ Valkey β€” Redis-compatible (sessions + rate limiting)
πŸ’» code-server β€” Standalone workspace (data source)
πŸ“‹ ConfigMaps Β· Secrets (tokens, provider config)

☁️ workspaces namespace

πŸ‘€ ws-{username} β€” code-server 4.115.0 per user
πŸ’Ύ Azure Disk PVC per user (ext4, persistent)
🐚 zsh + Oh-My-Zsh + kubectl/helm/git autocomplete
πŸ› οΈ Copilot CLI Β· Node.js 20 Β· Claude Code
πŸ“Š 4 cores / 4Gi RAM per workspace

πŸ““ jupyterhub namespace

🏠 JupyterHub β€” Multi-user notebook hub
πŸ‘€ jupyter-{username} β€” Per-user notebook pods
πŸ” Entra ID authentication
🐚 Custom zsh prompts with username

πŸ“Š monitoring

πŸ“ˆ Prometheus
πŸ“Š Grafana
πŸ“‹ Loki + Promtail
πŸ”” Alertmanager

πŸ”’ ingress-nginx

🌐 NGINX Controller
πŸ” TLS termination
↔️ SSL redirect
⬆️ Cluster Autoscaler: 2–6 nodes πŸ“¦ ACR: aiopmacr.azurecr.io πŸ”΄ Valkey: sessions + rate limiting πŸ›‘οΈ 20 req/min per-user rate limit

πŸ†• Recent Infrastructure Changes

● Storage: Azure File (CIFS) β†’ Azure Disk (ext4)
● Workspace memory: 256Mi β†’ 4Gi
● Build context: 336MB β†’ 3.5MB via .dockerignore
● GitHub Copilot OAuth device flow integration
● Multi-LLM: GitHub Models / Azure OpenAI / Ollama
● Flask-Session with Redis backend for HA
● πŸ”’ Private endpoints on Storage, Cosmos DB, Key Vault (public access disabled)
● Terraform state: Azure Blob β†’ Local backend
● Per-user sliding window rate limiter (20 req/min)
βœ•

πŸ“¦ workspaces namespace

Per-user cloud IDE environments. Each user gets their own ws-{username} pod running code-server 4.115.0.

Resources: 4 CPU cores, 4Gi RAM, Azure Disk PVC (ext4, persistent across restarts).

Tools: GitHub Copilot CLI, Claude Code, Node.js 20, kubectl, helm, git, zsh + Oh-My-Zsh with autocomplete.

Auth: nginx auth-url annotation gates access through /api/auth/verify.

Lifecycle: Spawn/stop/delete via admin API. Each workspace is a separate Helm release.

βœ•

πŸ“¦ jupyterhub namespace

Multi-user Jupyter notebook server. JupyterHub manages per-user jupyter-{username} pods with persistent storage.

Auth: Entra ID (Azure AD) β€” same identity as the main platform.

Kernels: Python 3, R, Julia, custom kernels. Pre-loaded data science libraries.

Integration: Notebooks can call ContextWeaver APIs and interact with MCP engines programmatically.

βœ•

πŸ“¦ ingress-nginx namespace

NGINX Ingress Controller handles all external traffic routing with TLS termination.

TLS: cert-manager auto-provisions and renews Let's Encrypt certificates for app.brainzbytes.com.

Auth gating: auth-url annotations on workspace and JupyterHub ingresses route through /api/auth/verify β€” only authenticated users can access their resources.

Public IP: 4.152.21.202 via Azure Load Balancer.