164 lines
6.6 KiB
Markdown
164 lines
6.6 KiB
Markdown
|
|
# Architecture
|
|||
|
|
|
|||
|
|
## 1. System context
|
|||
|
|
|
|||
|
|
OME Cockpit is a **BFF + SPA** running on `atc-portal01` (Docker). It does **not** embed OME; it continuously aggregates OME REST data, enriches with GPU telemetry and LLM chat, and pushes live snapshots to browsers over WebSocket.
|
|||
|
|
|
|||
|
|
### Actors
|
|||
|
|
|
|||
|
|
| Actor | Role |
|
|||
|
|
|-------|------|
|
|||
|
|
| ATC admin (browser) | Operates fleet via UI |
|
|||
|
|
| OpenManage Enterprise | Source of truth for devices/alerts/inventory |
|
|||
|
|
| vLLM (`atc-gpu-prod:8000`) | LLM inference for Cockpit chat |
|
|||
|
|
| Open WebUI (`:3080`) | Model catalog + virtual models (`ome-copilot`, `arena-model`) |
|
|||
|
|
| GPU metrics sidecar (`:9110`) | NVIDIA V100 utilization/power |
|
|||
|
|
| Target servers / iDRACs | SSH endpoints opened from the browser via cockpit bridge |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 2. Runtime components
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
┌─────────────────────────────────────────────────────────────┐
|
|||
|
|
│ Docker: ome-cockpit │
|
|||
|
|
│ uvicorn → FastAPI (main.py) :8090 → host :3090 │
|
|||
|
|
│ │
|
|||
|
|
│ Background tasks: │
|
|||
|
|
│ • poll_loop() ~POLL_INTERVAL (default 12–15s) │
|
|||
|
|
│ • gpu_loop() ~2s │
|
|||
|
|
│ │
|
|||
|
|
│ In-memory STATE + WebSocket fan-out │
|
|||
|
|
│ SQLite tickets → /data/ops.db │
|
|||
|
|
│ Static UI → /ui (bind-mount) │
|
|||
|
|
└─────────────────────────────────────────────────────────────┘
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Process model
|
|||
|
|
|
|||
|
|
- **Single uvicorn worker** (important): fleet `STATE`, WebSocket client sets, and SSH session registry are in-process memory. Do not scale to multiple workers without an external shared store.
|
|||
|
|
- UI assets and `main.py` are bind-mounted for lab hot-reload; image rebuild needed when Python dependencies change (`asyncssh`, etc.).
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 3. Data flows
|
|||
|
|
|
|||
|
|
### 3.1 Fleet snapshot (OME → UI)
|
|||
|
|
|
|||
|
|
1. `poll_loop` authenticates to OME `SessionService/Sessions`
|
|||
|
|
2. Fetches devices, groups, alerts (paged `$top`)
|
|||
|
|
3. For a capped set of connected servers, fetches `/Power`
|
|||
|
|
4. Normalizes into cockpit device nodes (`is_server`, `is_idrac`, subnet, watts, …)
|
|||
|
|
5. Diffs against `PREV_DEVICES` → `EVENT_FEED` (connection/power/status + **device_new** / **device_removed**)
|
|||
|
|
6. Writes `STATE`, broadcasts `{type:"snapshot", data: STATE}` to all `/ws/fleet` clients
|
|||
|
|
|
|||
|
|
Browser `app.js` applies snapshot → layout → canvas draw loop; `ops.js` updates GPU, tickets, notifications.
|
|||
|
|
|
|||
|
|
### 3.2 Device inventory (on demand)
|
|||
|
|
|
|||
|
|
`GET /api/devices/{id}`:
|
|||
|
|
|
|||
|
|
1. Resolves node from current `STATE`
|
|||
|
|
2. Opens OME session
|
|||
|
|
3. Reads `InventoryTypes` for that device and pulls each `InventoryDetails('…')`
|
|||
|
|
4. Builds normalized `landscape` (OS, firmware, drivers, applications, management, licenses)
|
|||
|
|
5. Caches in `DETAIL_CACHE` (~300s)
|
|||
|
|
|
|||
|
|
### 3.3 Chat
|
|||
|
|
|
|||
|
|
`POST /api/chat`:
|
|||
|
|
|
|||
|
|
1. Builds compact fleet system prompt (`build_fleet_context`, char-budgeted for 4096-token models)
|
|||
|
|
2. Routes:
|
|||
|
|
- `llama3-70b-gptq` (and similar) → direct vLLM `/v1/chat/completions`
|
|||
|
|
- `ome-copilot` / `arena-model` → Open WebUI `/api/chat/completions` (service login)
|
|||
|
|
3. Retries with tighter context on vLLM context-length 400 errors
|
|||
|
|
|
|||
|
|
### 3.4 SSH bridge
|
|||
|
|
|
|||
|
|
Browser `ssh.js` (xterm) → `WS /ws/ssh`:
|
|||
|
|
|
|||
|
|
1. First message: JSON `{type:"auth", host, port, username, password, cols, rows}`
|
|||
|
|
2. Server allowlists `host` against current fleet IPs only
|
|||
|
|
3. `asyncssh.connect` + interactive shell process
|
|||
|
|
4. Bidirectional pump; resize via JSON control messages
|
|||
|
|
5. Concurrency: global semaphore (25) + max 3 sessions per client IP; idle timeout 30 min
|
|||
|
|
|
|||
|
|
### 3.5 GPU telemetry
|
|||
|
|
|
|||
|
|
`gpu_loop` polls `GPU_METRICS_URL/api/gpu` every ~2s and broadcasts `{type:"gpu"}` (and embeds GPU in fleet snapshots).
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 4. Frontend architecture
|
|||
|
|
|
|||
|
|
| File | Responsibility |
|
|||
|
|
|------|----------------|
|
|||
|
|
| `index.html` | App chrome, rails, stage, drawers, modals (KPI, triage, connect, SSH, tickets) |
|
|||
|
|
| `app.js` | Canvas topology, camera, filters, KPIs, inspector, inventory render, KPI popup |
|
|||
|
|
| `ops.js` | Ops desk CRUD, chat drawer, triage modal, live-feed collapse/drag, GPU matrix |
|
|||
|
|
| `ssh.js` | SSH modal + xterm session lifecycle |
|
|||
|
|
| `styles.css` | Design tokens + component styles |
|
|||
|
|
| `vendor/xterm/*` | Local terminal assets (CDN-independent) |
|
|||
|
|
|
|||
|
|
### State
|
|||
|
|
|
|||
|
|
- Primary client state lives in `app.js` `state` (data, filters, camera, layoutMode, inventoryCache, …)
|
|||
|
|
- Shared bridge: `window.cockpit` and `window.cockpitSsh`
|
|||
|
|
|
|||
|
|
### Rendering
|
|||
|
|
|
|||
|
|
- Topology is an **immediate-mode canvas** (`requestAnimationFrame`), not DOM nodes for devices
|
|||
|
|
- Camera = translate + scale; layouts write target coordinates (`tx`/`ty`) with lerp for motion
|
|||
|
|
- Pulse packets on links are decorative; toggle via **Animate link pulse**
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 5. Persistence
|
|||
|
|
|
|||
|
|
| Store | Path | Contents |
|
|||
|
|
|-------|------|----------|
|
|||
|
|
| SQLite | `/data/ops.db` (`COCKPIT_DATA`) | Tickets + ticket_messages |
|
|||
|
|
| Browser localStorage | client | `feed_collapsed`, `feed_bottom`, `cockpit_ssh_user`, `cockpit_chat_model`, actor |
|
|||
|
|
| In-memory | process | Fleet STATE, DETAIL_CACHE, SSH_SESSIONS, EVENT_FEED |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 6. External contracts
|
|||
|
|
|
|||
|
|
### OME (OpenManage Enterprise 4.7)
|
|||
|
|
|
|||
|
|
- Session auth header `X-Auth-Token`
|
|||
|
|
- Device types: server `Type == 1000`; iDRAC via `SubDeviceType == "iDRAC"`
|
|||
|
|
- Inventory via per-device `InventoryTypes` + `InventoryDetails`
|
|||
|
|
|
|||
|
|
### vLLM OpenAI-compatible API
|
|||
|
|
|
|||
|
|
- `GET /v1/models`, `POST /v1/chat/completions`
|
|||
|
|
- Lab model `llama3-70b-gptq` with **max context 4096** → prompt budgeting is mandatory
|
|||
|
|
|
|||
|
|
### Open WebUI
|
|||
|
|
|
|||
|
|
- Sign-in → Bearer token (cached ~1h)
|
|||
|
|
- `GET /api/models`, `POST /api/chat/completions`
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 7. Failure domains
|
|||
|
|
|
|||
|
|
| Failure | UX impact | Mitigation |
|
|||
|
|
|---------|-----------|------------|
|
|||
|
|
| OME down | Stale/empty fleet | Health + last `updated_at`; poll retries |
|
|||
|
|
| vLLM 400 context | Chat errors | Slim context + retry |
|
|||
|
|
| Open WebUI auth missing | Virtual models unavailable | Falls back to vLLM list / error message |
|
|||
|
|
| SSH to non-fleet IP | Rejected | Allowlist |
|
|||
|
|
| Multi-user SSH overload | Queue/reject | Semaphore + per-IP cap |
|
|||
|
|
| Soft inspector refresh | Inventory preserved | `softRefreshInspector` + inventory cache |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 8. Related systems (out of repo)
|
|||
|
|
|
|||
|
|
- **OpenManage AI** (`mo/openmanage-ai`): Open WebUI branding + OME REST tool for demos
|
|||
|
|
- Native OME MCP at `https://ome/mcp` exists but session stickiness is unreliable in this lab — Cockpit does not depend on it
|