c2bb0465d0
Adds diagrams and expands the ATC demo path so the cockpit is fully described in-repo.
164 lines
6.6 KiB
Markdown
164 lines
6.6 KiB
Markdown
# Architecture
|
||
|
||
## 1. System context
|
||
|
||
OME Cockpit is a **BFF + SPA** running on `atc-portal01` (Docker). It does **not** embed OME; it continuously aggregates OME REST data, enriches with GPU telemetry and LLM chat, and pushes live snapshots to browsers over WebSocket.
|
||
|
||
### Actors
|
||
|
||
| Actor | Role |
|
||
|-------|------|
|
||
| ATC admin (browser) | Operates fleet via UI |
|
||
| OpenManage Enterprise | Source of truth for devices/alerts/inventory |
|
||
| vLLM (`atc-gpu-prod:8000`) | LLM inference for Cockpit chat |
|
||
| Open WebUI (`:3080`) | Model catalog + virtual models (`ome-copilot`, `arena-model`) |
|
||
| GPU metrics sidecar (`:9110`) | NVIDIA V100 utilization/power |
|
||
| Target servers / iDRACs | SSH endpoints opened from the browser via cockpit bridge |
|
||
|
||
---
|
||
|
||
## 2. Runtime components
|
||
|
||
```
|
||
┌─────────────────────────────────────────────────────────────┐
|
||
│ Docker: ome-cockpit │
|
||
│ uvicorn → FastAPI (main.py) :8090 → host :3090 │
|
||
│ │
|
||
│ Background tasks: │
|
||
│ • poll_loop() ~POLL_INTERVAL (default 12–15s) │
|
||
│ • gpu_loop() ~2s │
|
||
│ │
|
||
│ In-memory STATE + WebSocket fan-out │
|
||
│ SQLite tickets → /data/ops.db │
|
||
│ Static UI → /ui (bind-mount) │
|
||
└─────────────────────────────────────────────────────────────┘
|
||
```
|
||
|
||
### Process model
|
||
|
||
- **Single uvicorn worker** (important): fleet `STATE`, WebSocket client sets, and SSH session registry are in-process memory. Do not scale to multiple workers without an external shared store.
|
||
- UI assets and `main.py` are bind-mounted for lab hot-reload; image rebuild needed when Python dependencies change (`asyncssh`, etc.).
|
||
|
||
---
|
||
|
||
## 3. Data flows
|
||
|
||
### 3.1 Fleet snapshot (OME → UI)
|
||
|
||
1. `poll_loop` authenticates to OME `SessionService/Sessions`
|
||
2. Fetches devices, groups, alerts (paged `$top`)
|
||
3. For a capped set of connected servers, fetches `/Power`
|
||
4. Normalizes into cockpit device nodes (`is_server`, `is_idrac`, subnet, watts, …)
|
||
5. Diffs against `PREV_DEVICES` → `EVENT_FEED` (connection/power/status + **device_new** / **device_removed**)
|
||
6. Writes `STATE`, broadcasts `{type:"snapshot", data: STATE}` to all `/ws/fleet` clients
|
||
|
||
Browser `app.js` applies snapshot → layout → canvas draw loop; `ops.js` updates GPU, tickets, notifications.
|
||
|
||
### 3.2 Device inventory (on demand)
|
||
|
||
`GET /api/devices/{id}`:
|
||
|
||
1. Resolves node from current `STATE`
|
||
2. Opens OME session
|
||
3. Reads `InventoryTypes` for that device and pulls each `InventoryDetails('…')`
|
||
4. Builds normalized `landscape` (OS, firmware, drivers, applications, management, licenses)
|
||
5. Caches in `DETAIL_CACHE` (~300s)
|
||
|
||
### 3.3 Chat
|
||
|
||
`POST /api/chat`:
|
||
|
||
1. Builds compact fleet system prompt (`build_fleet_context`, char-budgeted for 4096-token models)
|
||
2. Routes:
|
||
- `llama3-70b-gptq` (and similar) → direct vLLM `/v1/chat/completions`
|
||
- `ome-copilot` / `arena-model` → Open WebUI `/api/chat/completions` (service login)
|
||
3. Retries with tighter context on vLLM context-length 400 errors
|
||
|
||
### 3.4 SSH bridge
|
||
|
||
Browser `ssh.js` (xterm) → `WS /ws/ssh`:
|
||
|
||
1. First message: JSON `{type:"auth", host, port, username, password, cols, rows}`
|
||
2. Server allowlists `host` against current fleet IPs only
|
||
3. `asyncssh.connect` + interactive shell process
|
||
4. Bidirectional pump; resize via JSON control messages
|
||
5. Concurrency: global semaphore (25) + max 3 sessions per client IP; idle timeout 30 min
|
||
|
||
### 3.5 GPU telemetry
|
||
|
||
`gpu_loop` polls `GPU_METRICS_URL/api/gpu` every ~2s and broadcasts `{type:"gpu"}` (and embeds GPU in fleet snapshots).
|
||
|
||
---
|
||
|
||
## 4. Frontend architecture
|
||
|
||
| File | Responsibility |
|
||
|------|----------------|
|
||
| `index.html` | App chrome, rails, stage, drawers, modals (KPI, triage, connect, SSH, tickets) |
|
||
| `app.js` | Canvas topology, camera, filters, KPIs, inspector, inventory render, KPI popup |
|
||
| `ops.js` | Ops desk CRUD, chat drawer, triage modal, live-feed collapse/drag, GPU matrix |
|
||
| `ssh.js` | SSH modal + xterm session lifecycle |
|
||
| `styles.css` | Design tokens + component styles |
|
||
| `vendor/xterm/*` | Local terminal assets (CDN-independent) |
|
||
|
||
### State
|
||
|
||
- Primary client state lives in `app.js` `state` (data, filters, camera, layoutMode, inventoryCache, …)
|
||
- Shared bridge: `window.cockpit` and `window.cockpitSsh`
|
||
|
||
### Rendering
|
||
|
||
- Topology is an **immediate-mode canvas** (`requestAnimationFrame`), not DOM nodes for devices
|
||
- Camera = translate + scale; layouts write target coordinates (`tx`/`ty`) with lerp for motion
|
||
- Pulse packets on links are decorative; toggle via **Animate link pulse**
|
||
|
||
---
|
||
|
||
## 5. Persistence
|
||
|
||
| Store | Path | Contents |
|
||
|-------|------|----------|
|
||
| SQLite | `/data/ops.db` (`COCKPIT_DATA`) | Tickets + ticket_messages |
|
||
| Browser localStorage | client | `feed_collapsed`, `feed_bottom`, `cockpit_ssh_user`, `cockpit_chat_model`, actor |
|
||
| In-memory | process | Fleet STATE, DETAIL_CACHE, SSH_SESSIONS, EVENT_FEED |
|
||
|
||
---
|
||
|
||
## 6. External contracts
|
||
|
||
### OME (OpenManage Enterprise 4.7)
|
||
|
||
- Session auth header `X-Auth-Token`
|
||
- Device types: server `Type == 1000`; iDRAC via `SubDeviceType == "iDRAC"`
|
||
- Inventory via per-device `InventoryTypes` + `InventoryDetails`
|
||
|
||
### vLLM OpenAI-compatible API
|
||
|
||
- `GET /v1/models`, `POST /v1/chat/completions`
|
||
- Lab model `llama3-70b-gptq` with **max context 4096** → prompt budgeting is mandatory
|
||
|
||
### Open WebUI
|
||
|
||
- Sign-in → Bearer token (cached ~1h)
|
||
- `GET /api/models`, `POST /api/chat/completions`
|
||
|
||
---
|
||
|
||
## 7. Failure domains
|
||
|
||
| Failure | UX impact | Mitigation |
|
||
|---------|-----------|------------|
|
||
| OME down | Stale/empty fleet | Health + last `updated_at`; poll retries |
|
||
| vLLM 400 context | Chat errors | Slim context + retry |
|
||
| Open WebUI auth missing | Virtual models unavailable | Falls back to vLLM list / error message |
|
||
| SSH to non-fleet IP | Rejected | Allowlist |
|
||
| Multi-user SSH overload | Queue/reject | Semaphore + per-IP cap |
|
||
| Soft inspector refresh | Inventory preserved | `softRefreshInspector` + inventory cache |
|
||
|
||
---
|
||
|
||
## 8. Related systems (out of repo)
|
||
|
||
- **OpenManage AI** (`mo/openmanage-ai`): Open WebUI branding + OME REST tool for demos
|
||
- Native OME MCP at `https://ome/mcp` exists but session stickiness is unreliable in this lab — Cockpit does not depend on it
|