[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-93302":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":9,"language":10,"languages":9,"totalLinesOfCode":9,"stars":11,"forks":12,"watchers":13,"openIssues":14,"contributorsCount":15,"subscribersCount":15,"size":15,"stars1d":15,"stars7d":16,"stars30d":16,"stars90d":15,"forks30d":15,"starsTrendScore":15,"compositeScore":17,"rankGlobal":9,"rankLanguage":9,"license":18,"archived":19,"fork":19,"defaultBranch":20,"hasWiki":21,"hasPages":19,"topics":22,"createdAt":9,"pushedAt":9,"updatedAt":35,"readmeContent":36,"aiSummary":37,"trendingCount":15,"starSnapshotCount":15,"syncStatus":14,"lastSyncTime":38,"discoverSource":39},93302,"Flawless","William-Lu-stack\u002FFlawless","William-Lu-stack","AI SRE AgenticOps for Kubernetes and cloud infrastructure.",null,"Python",663,119,59,2,0,42,74.44,"Other",false,"main",true,[23,24,25,26,27,28,29,30,31,32,33,34],"agenticops","ai","aiops","aisre","cloud","cloud-native","devops","kubernetes","llm","mcp","observability","sre","2026-07-22 04:02:08","# Flawless\n\n[![Python](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPython-3.11+-3776AB?logo=python&logoColor=white)](#quick-start)\n[![TypeScript](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FTypeScript-5.x-3178C6?logo=typescript&logoColor=white)](#development)\n[![Kubernetes](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FKubernetes-ready-326CE5?logo=kubernetes&logoColor=white)](#kubernetes-deployment)\n[![Docker](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDocker-flawless-2496ED?logo=docker&logoColor=white)](#docker)\n[![Langfuse](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLangfuse-optional-111827)](#langfuse)\n[![License](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Flicense-PolyForm%20Noncommercial-red)](#license)\n\n**Your infrastructure can explain itself, heal safely, and prove it recovered.**\u003Cbr>\n**让基础设施自己解释故障、安全完成修复，并证明它真的恢复了。**\n\n**Flawless** is an AI-native SRE control plane for Kubernetes and cloud infrastructure. It connects alerts, evidence, topology, human approval, controlled remediation, and recovery verification in one auditable AgenticOps loop.\n\n它不是另一个只会给建议的运维聊天框。Flawless 将“发现问题、收集证据、生成预演、人工授权、执行变更、恢复验证、经验沉淀”连接成一个可审计闭环。\n\nCreated in Shanghai by **陆宣宇 (Xuanyu Lu)**.\n\n![Flawless connects alerts, evidence, approval, remediation, and verified recovery](blog\u002Fassets\u002Fimages\u002Fluxyai-agenticops-loop.png)\n\n## Product Preview \u002F 产品实景\n\nThese are real console captures from the running platform, not conceptual mockups.\n\n| AI inspection and executable preview | SRE conversation and live evidence |\n|---|---|\n| ![Flawless AI inspection produces an evidence-backed remediation preview](docs\u002Fassets\u002Fflawless-ai-inspection.png) | ![Flawless SRE chat keeps the selected workload and live risk context](docs\u002Fassets\u002Fflawless-sre-chat.png) |\n\n![Flawless 3D topology impact analysis](docs\u002Fassets\u002Fflawless-topology-impact.png)\n\nCurrent release: **3.2.0**.\n\nRelease 3.2 adds persistent remediation lineage: every failed strategy, action,\nverification result, and replacement plan stays linked across operator-approved\nfollow-up jobs. The effectiveness ledger is persisted on the runtime volume so\nmodel comparisons and recovery records survive Pod restarts.\n\n> **Compatibility note:** the public product name is **Flawless**. Existing\n> `LUXYAI_*` environment variables, the `charts\u002Fluxyai` directory, storage\n> paths, and Kubernetes resource identifiers remain supported so current\n> installations can upgrade without a destructive migration.\n\n\n## The AgenticOps Loop\n\n`discover → diagnose → preview → approve → execute → verify → learn`\n\n- **Evidence first**: connect alerts, events, logs, metrics, topology, runbooks, and recent changes.\n- **Guarded action**: keep RBAC, policy, dry-run, human approval, and audit outside the model boundary.\n- **Verified recovery**: test the original symptom after execution instead of treating a successful command as success.\n\n## Field Notes \u002F 实战手记\n\nOnly published notes are listed below; links point to versioned files in this repository:\n\n- [From Alert to Verified Recovery \u002F 从告警到可验证恢复](blog\u002Fposts\u002F2026-07-13-from-alert-to-verified-recovery.md)\n- [Should AI Be Allowed to Fix Kubernetes? \u002F AI 可以修 Kubernetes 吗？](blog\u002Fposts\u002F2026-07-13-ai-should-earn-the-right-to-act.md)\n- [The Next SRE Control Plane Is More Than a Chat Box \u002F 下一代 SRE 控制平面](blog\u002Fposts\u002F2026-07-13-not-another-chatbox.md)\n- [Building AgenticOps from Shanghai \u002F 在上海构建 AgenticOps](blog\u002Fposts\u002F2026-07-13-building-agenticops-in-shanghai.md)\n\n## Why This Exists\n\nModern cloud systems fail in ways that are hard to reason about from a single log line:\n\n- a Pod restart can hide a PVC, image, scheduling, network, quota, or rollout issue;\n- a small workload change can affect services, data pipelines, middleware, and downstream users;\n- repeated human firefighting leaves valuable operational knowledge outside the platform;\n- model output is useful only when it is constrained by evidence, policy, permissions, and rollback.\n\nFlawless is built as an SRE control plane. It uses a model as a planner and explainer, but the platform keeps the execution boundary: RBAC, action catalog, dry-run, approval, audit, and recovery verification.\n\n## Core Features\n\n- **SRE Chat**: ChatGPT-style operations console with cluster, namespace, workload, and risk context.\n- **Inspection Queue**: scheduled or manual scans across Rancher\u002FKubernetes scopes with severity ranking.\n- **Controlled Remediation**: evidence collection, change preview, human approval, execution, post-change verification, and evidence-driven replanning that remembers failed strategies across follow-up jobs.\n- **Topology Impact**: 2D\u002F3D topology, CMDB-style dependencies, eBPF\u002Fdata-flow adapters, blast-radius analysis.\n- **Release Governance**: SLO, error budget, canary\u002Frisk gate, emergency fix path, and release audit chain.\n- **Skills Library**: portable operation skills that encode expert knowledge and can be reused by other agents.\n- **Knowledge Base**: upload text, Markdown, PDF, Word, Excel, logs, YAML, and runbooks for operations RAG.\n- **Model Lab**: configure multiple OpenAI-compatible or OAuth-protected model gateways and compare outcomes.\n- **Measurable Outcomes**: persistent remediation lineage, changed-resource history, recovery evidence, and model effectiveness comparisons.\n- **Observability**: Prometheus metrics, Loki logs, Tempo traces, Grafana links, and optional Langfuse traces.\n- **Extensible Infrastructure**: adapters for Kubernetes, Rancher, databases, virtual machines, storage, and middleware.\n\n## Architecture\n\n```text\nFrontend Console\n  ├─ SRE Chat \u002F Inspection \u002F Topology \u002F Release \u002F Skills \u002F Models\n  │\n  ▼\nControl Plane API\n  ├─ Evidence pipeline\n  ├─ Remediation job state machine\n  ├─ Release gate and SLO budget\n  ├─ Knowledge and model registries\n  ├─ Observability store\n  └─ Integration health checks\n  │\n  ├─ MCP Kubernetes tools\n  ├─ A2A healing \u002F incident \u002F postmortem agents\n  ├─ Rancher \u002F Prometheus \u002F CMDB \u002F eBPF flow adapters\n  ├─ Langfuse \u002F Loki \u002F Tempo \u002F Grafana adapters\n  └─ Optional custom algorithm extension\n```\n\n## Quick Start\n\nThe quick-start command performs prerequisite checks, creates `.env` when\nneeded, builds the console, starts the API plus all local agents\u002FMCP services,\nand waits for the complete core health check to pass.\n\nYou need Git, plus either:\n\n- Docker Engine or Docker Desktop with Compose v2; or\n- Python 3.11+ and Node.js 20+ for the automatic native fallback.\n\n### Always-Latest One-Click Deployment\n\nEach operating-system installer pulls only the canonical `main` branch, safely\nfast-forwards an existing clean checkout, verifies that `HEAD` exactly matches\n`origin\u002Fmain`, and starts Flawless with Docker.\n\n#### macOS (Docker Desktop)\n\n```bash\ncurl -fsSL --retry 3 \\\n  https:\u002F\u002Fraw.githubusercontent.com\u002FWilliam-Lu-stack\u002FFlawless\u002Fmain\u002Fscripts\u002Finstall-macos.sh \\\n  | bash -s -- --china\n```\n\n#### Linux or WSL (Docker Engine)\n\n```bash\ncurl -fsSL --retry 3 \\\n  https:\u002F\u002Fraw.githubusercontent.com\u002FWilliam-Lu-stack\u002FFlawless\u002Fmain\u002Fscripts\u002Finstall-linux.sh \\\n  | bash -s -- --china\n```\n\n#### Windows (PowerShell + Docker Desktop)\n\nRun in PowerShell, not Command Prompt:\n\n```powershell\n& ([scriptblock]::Create((irm `\n  \"https:\u002F\u002Fraw.githubusercontent.com\u002FWilliam-Lu-stack\u002FFlawless\u002Fmain\u002Fscripts\u002Finstall-windows.ps1\"))) -China\n```\n\nRemove `--china` or `-China` when mainland China mirrors are not needed. The\ndefault install directory is `~\u002FFlawless` on macOS\u002FLinux and\n`$HOME\\Flawless` on Windows.\n\nThe portable macOS\u002FLinux installer remains available when automatic Docker\nfallback is preferred:\n\n```bash\ncurl -fsSL --retry 3 \\\n  https:\u002F\u002Fraw.githubusercontent.com\u002FWilliam-Lu-stack\u002FFlawless\u002Fmain\u002Fscripts\u002Finstall.sh | bash\n```\n\nThe installer never resets or overwrites local changes. It stops and restarts\nan existing stack only when the verified revision actually changed.\n\n### Manual Clone\n\n```bash\ngit clone https:\u002F\u002Fgithub.com\u002FWilliam-Lu-stack\u002FFlawless.git\ncd Flawless\n.\u002Fscripts\u002Fquickstart.sh\n```\n\nOpen `http:\u002F\u002F127.0.0.1:8080` after the command reports `ready`.\n\nTo prove that an existing checkout is current, both `Latest status: latest`\nand `Worktree: clean` must be shown:\n\n```bash\n.\u002Fscripts\u002Fquickstart.sh version\n.\u002Fscripts\u002Fquickstart.sh update\n.\u002Fscripts\u002Fquickstart.sh doctor\n```\n\n`doctor` reports the exact commit, remote comparison, platform, Docker Compose,\nDocker daemon, Python, Node.js, npm, and active runtime without printing `.env`\nor credential values. Attach its output when reporting an installation failure.\n\nUseful controls:\n\n```bash\n.\u002Fscripts\u002Fquickstart.sh status\n.\u002Fscripts\u002Fquickstart.sh logs\n.\u002Fscripts\u002Fquickstart.sh stop\n.\u002Fscripts\u002Fquickstart.sh --china\n.\u002Fscripts\u002Fquickstart.sh --port 18080 --open\n.\u002Fscripts\u002Fquickstart.sh --mode native\n.\u002Fscripts\u002Fquickstart.sh --mode docker\n```\n\n`auto` mode prefers Docker when its daemon is running and otherwise uses the\nnative toolchain. Runtime data and logs stay under `.flawless\u002F` in native mode\nor in the `flawless-data` Docker volume. Re-running the command is safe.\n\n### Model Configuration\n\nThe console and baseline workflows start without a live model endpoint. To use\nAI chat, configure `.env` for an OpenAI-compatible local endpoint such as\nOllama, then restart the stack:\n\n```env\nLLM_API_BASE=http:\u002F\u002Flocalhost:11434\u002Fv1\nLLM_API_KEY=ollama\nLLM_MODEL=qwen2.5:7b\nLLM_AUTH_TYPE=api_key\n```\n\nIn Docker mode, the quick-start script automatically maps `localhost` model\nURLs to `host.docker.internal`. For an OAuth client-credentials gateway:\n\n```env\nLLM_AUTH_TYPE=oauth_client_credentials\nOAUTH_TOKEN_URL=https:\u002F\u002Fyour-iam\u002Frealms\u002Fmain\u002Fprotocol\u002Fopenid-connect\u002Ftoken\nOAUTH_CLIENT_ID=your-client\nOAUTH_CLIENT_SECRET=your-secret\nLLM_API_BASE=https:\u002F\u002Fyour-llm-gateway\u002Fengines\u002Fdefault\nLLM_MODEL=your-model\nLLM_VERIFY_SSL=true\n```\n\n### Kubernetes Access\n\nThe local console can start without Kubernetes credentials. In that state the\nUI, API, agents, MCP gateway, and dry-run planning are available, while cluster\ntools return a clear `Kubernetes access is not configured` response.\n\nNative mode automatically reads `KUBECONFIG` or `~\u002F.kube\u002Fconfig`:\n\n```bash\n.\u002Fscripts\u002Fquickstart.sh --mode native\n```\n\nDocker mode intentionally does not mount host cluster credentials by default.\nFor a real cluster, use the Kubernetes\u002FHelm deployment below so the workloads\nreceive a scoped ServiceAccount, or explicitly provide a read-only kubeconfig\nmount appropriate for your local cluster.\n\n## Docker\n\nThe one-click command uses [`compose.yaml`](compose.yaml) automatically when\nDocker is available. The Compose service builds one image, runs the complete\nlocal service group, persists runtime state, and publishes only the console on\nthe loopback interface.\n\nDirect Compose usage is also supported:\n\n```bash\ncp .env.example .env\ndocker compose up -d --build\ndocker compose ps\n```\n\nOr build and run the all-in-one image without Compose:\n\n```bash\ndocker build --target backend-runtime -t flawless:latest .\ndocker run --rm \\\n  --env-file .env \\\n  --add-host host.docker.internal:host-gateway \\\n  -p 127.0.0.1:8080:8080 \\\n  -v flawless-data:\u002Fvar\u002Flib\u002Fflawless \\\n  flawless:latest\n```\n\nPush to GHCR or an internal registry:\n\n```bash\nIMAGE=ghcr.io\u002Fwilliam-lu-stack\u002Fflawless:latest .\u002Fscripts\u002Fbuild-push.sh\n```\n\nFor air-gapped or China mainland networks, the `Dockerfile` and `scripts\u002Fbuild-push.sh` already expose mirror build args:\n\n```bash\nNODE_IMAGE=docker.m.daocloud.io\u002Flibrary\u002Fnode:24-slim \\\nPYTHON_IMAGE=docker.m.daocloud.io\u002Flibrary\u002Fpython:3.13-slim \\\nNPM_REGISTRY=https:\u002F\u002Fregistry.npmmirror.com \\\nPIP_INDEX_URL=https:\u002F\u002Fmirrors.aliyun.com\u002Fpypi\u002Fsimple \\\nIMAGE=your-registry\u002Fflawless:latest \\\n.\u002Fscripts\u002Fbuild-push.sh\n```\n\n## Kubernetes Deployment\n\n### Recommended: Helm\n\nThe chart installs the control plane, six agent services, controlled cluster-wide\nRBAC, Generic NAS persistence, NodePort access, optional Ingress, and Secret hooks:\n\n```bash\nkubectl create namespace k8s-agent\n\nhelm upgrade --install flawless .\u002Fcharts\u002Fluxyai \\\n  --namespace k8s-agent \\\n  --set image.repository='\u003Cinternal-registry>\u002Fflawless' \\\n  --set image.tag='\u003Crelease-tag>' \\\n  --set persistence.storageClass=standard\n```\n\nUse `rbac.mode=controlled` in production. `rbac.mode=cluster-admin` exists only\nfor isolated validation environments and should require a security exception.\n\nRender and review before deployment:\n\n```bash\nhelm lint .\u002Fcharts\u002Fluxyai --strict\nhelm template flawless .\u002Fcharts\u002Fluxyai -n k8s-agent > rendered.yaml\n```\n\n### Raw Manifests\n\n#### 1. Create Namespace\n\n```bash\nkubectl create namespace k8s-agent\n```\n\n#### 2. Create Secrets\n\nDo not commit real secrets. Create them from your terminal or secret manager:\n\n```bash\nkubectl -n k8s-agent create secret generic k8s-agent-oauth \\\n  --from-literal=OAUTH_CLIENT_ID='\u003Cclient-id>' \\\n  --from-literal=OAUTH_CLIENT_SECRET='\u003Cclient-secret>' \\\n  --from-literal=RANCHER_TOKEN='\u003Coptional-rancher-token>'\n```\n\nFor Langfuse:\n\n```bash\nkubectl -n k8s-agent create secret generic k8s-agent-langfuse \\\n  --from-literal=public-key='\u003Cpk-lf-...>' \\\n  --from-literal=secret-key='\u003Csk-lf-...>'\n```\n\nAdministrator configuration is disabled by default. The password must never be\nstored in `values.yaml`, a ConfigMap, or the repository:\n\n```bash\nkubectl -n k8s-agent create secret generic luxyai-console-auth \\\n  --from-literal=CONSOLE_BASIC_AUTH_USERNAME='admin' \\\n  --from-literal=CONSOLE_BASIC_AUTH_PASSWORD='\u003Cpassword-from-vault>'\n\nhelm upgrade --install flawless .\u002Fcharts\u002Fluxyai \\\n  -n k8s-agent \\\n  --reuse-values \\\n  --set admin.enabled=true\n```\n\nWhen admin mode is off, the console remains readable and operational approvals\ncontinue to work, but model, knowledge, and Skill writes are blocked. When it is\non, those three write surfaces require the Secret-backed administrator identity.\n\n#### 3. Configure Runtime\n\nEdit `manifests\u002Fdeployment.yaml`:\n\n- `LLM_AUTH_TYPE`\n- `OAUTH_TOKEN_URL`\n- `LLM_API_BASE` \u002F `LLM_GATEWAY_BASE`\n- `LLM_MODEL`\n- `PROMETHEUS_URL`\n- `CMDB_URL`\n- `RANCHER_URL`\n- `RANCHER_CLUSTER_IDS`\n- `ALLOWED_NAMESPACES`\n- `OPS_MUTATION_ENABLED`\n- `AUTO_HEALING_ENABLED`\n\n#### 4. Apply Manifests\n\n```bash\nkubectl apply -f manifests\u002Frbac.yaml\nkubectl apply -f manifests\u002Fdeployment.yaml\nkubectl apply -f manifests\u002Ffrontend.yaml\n```\n\nThe default service exposes the console through NodePort `30080`:\n\n```bash\nkubectl get svc -n k8s-agent\n```\n\nFor production, use your company Ingress\u002FGateway with TLS and identity middleware instead of exposing an unauthenticated public NodePort.\n\n## Model Configuration\n\nFlawless supports two common model access patterns.\n\n### OAuth Token URL + Base URL\n\nUse this when your gateway requires a dynamic bearer token:\n\n```json\n[\n  {\n    \"id\": \"primary\",\n    \"provider\": \"oauth-gateway\",\n    \"model\": \"your-model\",\n    \"base_url\": \"https:\u002F\u002Fyour-gateway\u002Fengines\u002Fdefault\",\n    \"auth_type\": \"oauth_client_credentials\",\n    \"token_url\": \"https:\u002F\u002Fyour-iam\u002Ftoken\",\n    \"client_id\": \"your-client\",\n    \"client_secret\": \"from-secret\",\n    \"role\": \"primary\",\n    \"max_tokens\": 4096,\n    \"verify_ssl\": true\n  }\n]\n```\n\n### Base URL + API Key\n\nUse this for OpenAI-compatible providers:\n\n```json\n[\n  {\n    \"id\": \"openai-compatible\",\n    \"provider\": \"openai-compatible\",\n    \"model\": \"your-model\",\n    \"base_url\": \"https:\u002F\u002Fapi.example.com\u002Fv1\",\n    \"auth_type\": \"api_key\",\n    \"api_key\": \"from-secret\",\n    \"role\": \"candidate\",\n    \"max_tokens\": 4096\n  }\n]\n```\n\nSet the JSON in `MODEL_PROFILES_JSON` or add models from the **Model Lab** page. Secrets should come from Kubernetes Secret or your enterprise secret platform.\n\n## Langfuse\n\nLangfuse is optional. When configured, the platform records model calls, latency, token usage, cost estimates, tool spans, quality scores, and trace IDs.\n\nEnvironment variables:\n\n```env\nLANGFUSE_ENABLED=true\nLANGFUSE_HOST=http:\u002F\u002Flangfuse-web.langfuse.svc.cluster.local:3000\nLANGFUSE_PUBLIC_KEY=pk-lf-...\nLANGFUSE_SECRET_KEY=sk-lf-...\n```\n\nDeployment references:\n\n- Docker Compose: https:\u002F\u002Flangfuse.com\u002Fself-hosting\u002Fdeployment\u002Fdocker-compose\n- Kubernetes Helm: https:\u002F\u002Flangfuse.com\u002Fself-hosting\u002Fdeployment\u002Fkubernetes-helm\n\nThis repository also contains `manifests\u002Flangfuse-local.yaml` as a local reference manifest. It is ignored by Git by default because production Langfuse credentials and storage settings should be managed separately.\n\n## Custom Algorithm Extension\n\nThe public repository includes a runnable baseline algorithm module at `agents\u002Faiops_algorithms.py`.\n\nIf you have a custom scoring implementation, keep it outside the repository and load it at runtime:\n\n```bash\nexport LUXYAI_CUSTOM_ALGORITHM_PATH=.local\u002Fcustom_algorithms\u002Faiops_algorithms_custom.py\nuvicorn backend.app.main:app --host 0.0.0.0 --port 8080\n```\n\nDocker:\n\n```bash\ndocker run --rm \\\n  --env-file .env \\\n  -e LUXYAI_CUSTOM_ALGORITHM_PATH=\u002Fvar\u002Flib\u002Fluxyai-custom\u002Faiops_algorithms_custom.py \\\n  -v \"$PWD\u002F.local\u002Fcustom_algorithms:\u002Fvar\u002Flib\u002Fluxyai-custom:ro\" \\\n  -p 8080:8080 \\\n  flawless:latest \\\n  python -m uvicorn backend.app.main:app --host 0.0.0.0 --port 8080\n```\n\nKubernetes:\n\n```bash\nkubectl -n k8s-agent create secret generic luxyai-custom-algorithms \\\n  --from-file=aiops_algorithms_custom.py=.local\u002Fcustom_algorithms\u002Faiops_algorithms_custom.py\nkubectl rollout restart deploy\u002Fluxyai -n k8s-agent\nkubectl rollout restart deploy\u002Fk8s-agent-api -n k8s-agent\n```\n\nIf the secret is absent, the platform still runs with the open baseline.\n\nTo explicitly test the open baseline path:\n\n```bash\nLUXYAI_DISABLE_CUSTOM_ALGORITHMS=1 uvicorn backend.app.main:app --host 0.0.0.0 --port 8080\n```\n\n## Skills\n\nSkills are portable operational knowledge packages. They describe:\n\n- symptoms and trigger conditions;\n- required evidence;\n- allowed objects;\n- allowed actions;\n- recovery criteria;\n- rollback guidance;\n- optional references and runbooks.\n\nSkills are stored under `OPS_SKILL_ROOT` and can be created from the console. They are intentionally portable so they can be reused by other agents or moved between environments.\n\nEach Skill is persisted as an independent package directory, rather than being\ncompiled into the application. This keeps the Skill repository separately\nversionable, reviewable, exportable, and reusable by other compatible agents.\n\n## Unified Resource API\n\n`GET \u002Fapi\u002Fresources` exposes Kubernetes, databases, virtual machines,\nmiddleware, storage, and cloud resources through one stable contract:\n\n```text\nGET \u002Fapi\u002Fresources?resource_type=pod&cluster=prod&namespace=orders&limit=200\n```\n\nThe response uses contract `luxyai.resource.v1`, includes source and health\nsummaries, and supports cursor pagination. New infrastructure teams should add\nan adapter and normalize into this contract instead of introducing a parallel\nresource API.\n\n## Safety Model\n\nThe platform is designed around least privilege:\n\n- no browser-side shell or arbitrary command execution;\n- no mutation unless server switches allow it;\n- high-risk actions require explicit operator confirmation;\n- action catalog limits what the model can request;\n- namespace\u002Fworkload scope is controlled by RBAC and allowlists;\n- every change records preview, actor, diff, result, and verification status;\n- secrets are never committed and should be supplied through Kubernetes Secret or a secret manager.\n\n## Repository Layout\n\n```text\nFlawless\u002F\n├── agents\u002F                  # SRE workflow agents and execution engines\n├── backend\u002Fapp\u002F             # Control plane API, schemas, services, domain logic\n├── mcp_servers\u002F             # Kubernetes MCP tools\n├── cmdb\u002F                    # Local CMDB\u002Ftopology service\n├── cloud\u002F                   # Cloud and infrastructure adapter contracts\n├── frontend\u002Fmodern\u002F         # React + TypeScript + Vite console\n├── charts\u002Fluxyai\u002F      # Production Helm chart\n├── manifests\u002F               # Kubernetes manifests\n├── scripts\u002F                 # Build and image helper scripts\n├── docs\u002F                    # Architecture and maintainer documentation\n├── examples\u002F                # Sample alerts\n└── tests\u002F                   # Backend and workflow tests\n```\n\n## Development\n\n```bash\npython -m compileall -q backend agents mcp_servers cmdb cloud a2a openwebui\npython -m unittest discover -s tests -v\ncd frontend\u002Fmodern && npm run build\n```\n\nRun the security gate locally:\n\n```bash\npython -m pip install pip-audit\npip-audit -r requirements.lock --no-deps --disable-pip\n```\n\n## GitHub\n\nThe public engineering baseline lives at\n[`William-Lu-stack\u002FFlawless`](https:\u002F\u002Fgithub.com\u002FWilliam-Lu-stack\u002FFlawless). Custom algorithm\nextensions, credentials, production topology, and company data are loaded at\nruntime and are intentionally excluded from the public repository.\n\n## Roadmap\n\nFlawless is designed to grow from a Kubernetes SRE console into an AgenticOps operating system for modern infrastructure.\n\n- **Kubernetes Autopilot**: cover the full lifecycle from alert, evidence, root-cause analysis, remediation preview, approval, execution, rollback, and recovery verification.\n- **Rancher Multi-Cluster Fleet**: make every cluster, namespace, workload, event, metric, and operation record searchable and governable from one control plane.\n- **Full-Stack Infrastructure Operations**: extend the same evidence-to-action loop to databases, virtual machines, storage, middleware, ingress, service mesh, and hybrid-cloud resources.\n- **Runtime Data-Flow Intelligence**: fuse Kubernetes inventory, CMDB, Prometheus, Loki, Tempo, and eBPF flow data into a living dependency graph that explains impact, blast radius, and traffic direction.\n- **Release Governance Control Plane**: turn SLO, error budget, canary scope, image risk, YAML policy, topology risk, and emergency repair paths into a programmable release gate.\n- **Operations Skills Network**: let engineers package hard-won troubleshooting experience as portable Skills, so the platform becomes stronger every time an incident is solved.\n- **Model Benchmark Arena**: evaluate different models by remediation success rate, MTTR reduction, safety score, evidence quality, token cost, and rollback correctness.\n- **Digital Twin for Change Risk**: simulate changes against topology, historical incidents, dependency paths, and SLO budgets before production is touched.\n- **100k-Node Scale Architecture**: move heavy discovery to event-driven collectors, sharded caches, async job queues, streaming evidence pipelines, and pluggable stores.\n- **Enterprise Trust Layer**: ship production Helm charts, air-gapped packages, OIDC\u002FSSO, RBAC presets, audit retention, policy-as-code, secret-manager integration, and compliance reports.\n- **Cloud & Edge Expansion**: add first-class adapters for Alibaba Cloud, Generic Cloud, Tencent Cloud, AWS, Azure, private cloud, edge clusters, and cross-region disaster recovery.\n- **Self-Healing Platform Runtime**: let the platform inspect and repair its own agents, collectors, queues, stores, and integrations under strict approval and audit boundaries.\n\n## License\n\nThis project is released under the standardized\n[PolyForm Noncommercial License 1.0.0](LICENSE).\n\nYou may use, study, modify, and redistribute the source code for non-commercial purposes. Commercial use, hosted commercial services, resale, enterprise product bundling, paid support services, and removal of author attribution require prior written authorization from **陆宣宇 (Xuanyu Lu)**.\n\nThis is a standardized source-available non-commercial license, not an\nOSI-approved open-source license. Commercial authorization requests can be sent\nto the repository owner through GitHub.\n","Flawless 是一个面向 Kubernetes 与云基础设施的 AI 原生 SRE 控制平面，实现可审计的自动化运维闭环。其核心功能包括：基于多源证据（告警、日志、指标、拓扑等）的故障诊断与修复预演；严格遵循 RBAC 与人工审批的受控执行；以及以原始症状是否消失为依据的恢复验证机制。技术上采用 AgenticOps 架构，将发现、诊断、预演、审批、执行、验证与经验沉淀集成于单一闭环，并支持持久化修复血缘与 Pod 重启后状态恢复。适用于中大型云原生环境下的生产级 SRE 团队，尤其适合对变更安全性和运维可追溯性要求严格的金融、电信及企业级云平台场景。","2026-07-16 02:30:02","CREATED_QUERY"]