Add Project Nessie on lake01 with Command Center Iceberg tab (snapshots,
manifests, data files, time-travel SQL). Data Flow pulses only when endpoints
are reachable; offline nodes/edges render red.
Pass DOCKHAND_API_TOKEN to container inventory calls so pipeline_active
and topology animation work after Authentik. Replace ring-buffer-only CDC
stats with minute rollups (no 1000 cap), add 15m/1h/6h/24h window selector
on the Live Changes tab, and poll recent events on an interval.
Co-authored-by: Cursor <cursoragent@cursor.com>
Adds native Command Center features (no new containers) integrated as sub-tabs
in the existing Data Explorer and Data Quality views:
- Continuous Data Quality (dq_monitor.py): live completeness/uniqueness/validity/
freshness scorecards via Trino with rolling trends → DataQuality "Live Monitoring".
- Ownership & stewardship (catalog_governance.py): owner/steward/tier matrix,
orphan detection, business glossary; local store best-effort synced to
OpenMetadata (owner PATCH) → Data Explorer "Ownership".
- Access & policy posture: per-dataset compliance combining PII masking, ownership,
live DQ and observability alerts vs data contracts → Data Explorer "Access & Policies".
- Lineage (lineage.py): staged source→CDC→Spark→S3→Iceberg→Trino→serving graph with
live row counts and column-level PII/masking tracing → Data Explorer "Lineage".
- Observability (observability.py): volume/freshness/schema-drift monitoring with
alerts → Data Explorer "Observability".
- Shared lake_meta.py dataset registry + bounded Trino client; fast native row-count
and PK-indexed freshness so monitors stay cheap on 25-54M-row tables.
- LLM context (lab_context.py) enriched with DQ scores, ownership and active alerts.
- streaming_ops: add connector status discovery, manual resync endpoint
(POST /api/pipeline/streaming/resync), connectors status endpoint, and a
background connector_autoheal_loop that restarts FAILED tasks automatically
- main: wire connector_autoheal_loop into app lifespan
- api.ts: add resyncSources() helper
- ChangesView: add "Re-sync sources" header button with live status
Agent terminals were idle (one-shot probe) while agents were busy in the
background. Now every agent streams what it is actually doing:
- agent_terminal: emit_threadsafe() so background threads can stream lines.
- agent_ops: Data Custodian DML loop logs the real INSERT/UPDATE/DELETE SQL
(+ Mongo ops) and Hadoop-offload Trino CTAS/INSERT to its terminal; ETL
Guardian announces each orchestrated movement.
- movements: trigger_and_watch streams the Airflow DAG / API call, conf,
before/after Trino counts and result to the owning agent terminal.
- etl_offload: per-dataset read + pyarrow->S3 parquet writes and cycle
summaries stream to the ETL Guardian terminal.
- agent_activity (new): round-robin live probes for Lakehouse Ops, Hadoop
Ranger (NameNode JMX + YARN), Infra Sentinel (Dockhand inventory + host
load) and Network Watcher (VLAN 20/21 path checks).
- fix: YARN ResourceManager runs on 10.0.21.62:8088 (was .61).
- ui: terminal dock merges the selected agent ops stream with the node probe.
- etl_offload.py: autonomous agent backfills/tails source DBs (PG/MySQL/
Mongo/Cassandra) to S3 as Parquet in small chunks, accumulates a live
federated business matrix (/api/etl/status, /api/etl/business, /run, /config).
- storage_s3.py: buffer generated CDC + masked curated rows to S3, overlay
live last-write into analytics; put_object_bytes for Parquet parts.
- trino_federated.py: capture generated rows + archive to S3; generator_active.
- dataflow.py: pulse generate + kafka/spark->S3 archive edges when active.
- StorageView: realtime ETL ingest panel; TrinoFederationView: realtime
business KPIs/charts from /api/etl/business.
- ChangesView: top KPIs/charts now overlay the live WS stream on server stats
so they update in lock-step with the bottom feed; faster 2.5s refresh.
- useCommandCenter: retain 800 live CDC changes.
Generated data is now visible flowing into the source systems on the Data
Flow graph: a manual "Generate data" burst (and the continuous live loop)
stamps a last_tick, and a new generator_active() helper lights up the
generator -> postgres/mysql/mongodb/cassandra edges for ~12s. Neo4j stays
dark since the live generator does not write to it.
Redesigned the Changes tab into a "New & Changed Data" dashboard driven by
the live CDC stream: animated KPI cards (new/updated/deleted + throughput),
a change-volume area chart, an operation-mix donut, per-system and
top-table breakdown bars, plus the existing filterable change feed.
Previously Run/Pause/Stop only drove the pulse animation, so on the Data Flow
tab nothing was generated (the streamer is heartbeat-gated and only the Live
dashboard was sending heartbeats). Now:
- New POST /api/federated/live/heartbeat keep-alive endpoint.
- Run seeds an immediate burst into the sources and keeps the generator alive
(5s heartbeat) while the tab is open; Pause/Stop halts generation.
Data Flow tab:
- Prominent "Generate data" button (500 / 2K / 10K) that inserts a fresh
burst of business rows into all source DBs on demand via a new
POST /api/federated/generate (fresh connections, safe alongside the
background streamer); result toast shows what was inserted, CDC streams it.
- "Scripts" button + a "View generation scripts" action on the Data Generator
node open a modal listing every generator script with full source, served by
GET /api/dataflow/scripts. Sources are the real files: the live streaming
generator (sliced live out of trino_federated.py) and the Airflow per-source
DAGs + Faker scripts (mounted read-only from infra/airflow into the API).
Knowledge Chat:
- New "Vector DB" explorer modal: shows the ChromaDB chunking config
(RecursiveCharacterTextSplitter 800/120, all-MiniLM-L6-v2, 384-dim, HNSW),
collections & documents, and the actual stored chunks with text, metadata and
an embedding preview (bars + values) so you can see exactly how files are
split and written as vectors.
Refactor: generator row-builders shared by the streamer and the on-demand burst.
Live dashboard now feels truly real-time:
- Background generator streams randomly-sized bursts of real rows into
PostgreSQL, MySQL, MongoDB & Cassandra every ~4s (CDC picks them up).
Throughput rises and falls; counters move in lock-step (base snapshot +
generated). Runs only while the Live tab is polling (heartbeat-gated) so
source tables do not grow unbounded; on/off toggle exposed in the UI.
- New /api/federated/live/generator toggle; /live returns per-tick activity
(last burst sizes, orders by region/status, event feed).
- LiveDashboard: live-activity panel, orders-per-tick sparkline, event
stream feed, burst-by-region/status charts, generator status + control.
Data Flow graph now explains how data reaches the assistant:
- Added ChromaDB -> RAG (LangChain) -> vLLM Gateway -> Knowledge Chat lane,
with Trino / OpenMetadata / curated-masked feeding LLM context. Live model
& embed metrics pulled from the RAG /config. New node/edge kinds + legend.
- New "Live" tab: auto-polls /api/federated/live every 2.5s with
animated counters, ingestion throughput sparkline, per-source
write-rate bars, a region scorecard matrix (heat-shaded) and live
business breakdown charts (region/channel/status/customers/
telemetry/supply)
- Backend /api/federated/live: instant source estimates (Postgres),
monotonic max(event_id) for MySQL and Mongo estimated count for
immediate movement, Cassandra from cached matrix; business aggs
cached over the small Hadoop lake tables (short TTL)
- Fix embedded Trino panels being squeezed with internal scrollbars
by making panels/grids shrink-0 so the page scrolls instead
The marquee panel only joined 3 region-keyed sources. Add a
"one SQL across every database" reach matrix that fans a single
Trino query out to PostgreSQL, MySQL, MongoDB, Cassandra and the
Hadoop/Iceberg lake in one UNION ALL (telemetry has no region, so
a per-source summary is used instead of a misleading join).
- New MATRIX_SQL + concurrent execution alongside the region
scorecard so total latency stays ~ the slower query
- Federated tab shows the 5-source matrix (records + headline
metric per engine) above the relabelled 3-source region scorecard
- Data Explorer now hosts Business Overview, Federated (Trino),
Hadoop Lake and Data Dictionary as tabs instead of a separate
top-level view (richer, less stacked layout)
- TrinoFederationView supports an embedded mode driven by an
external active tab
- Remove the standalone Trino Federation sidebar entry/route
- Remove the Kibana shortcut from the sidebar (it already lives in
the Elasticsearch view)
Trino Federation tab (3 sub-views):
- Federated: catalog landscape + a single cross-source SQL that joins
PostgreSQL + MySQL + MongoDB (region scorecard) — the federation proof,
computed in the background and cached (large full scans take ~2 min).
- Hadoop Lake: all federated business data materialized as external Iceberg
tables on HDFS (iceberg.hadoop.*_ext, ~120k rows each) with live, fast
business analytics (revenue by region/channel, top customers, HR by
department, supply by type, telemetry averages). Includes a one-click
"rebuild external tables" job.
- Data Dictionary: every business table + column with masked / visible PII
badges and categories.
Backend trino_federated.py: /catalogs, /marquee(+refresh), /lake,
/materialize(+status), /dictionary. Name-based PII detection flags raw PII
in derived/lake tables as visible vs physically-masked curated layer.
LLM context: platform_context now emits a full BUSINESS DATA CATALOG section
(tables, columns, types, source row counts, federated scorecard) with exact
per-column masked/visible status, so the assistant knows the data in detail
and what is masked vs not.
Add a dedicated business-data view (separate from infra/search):
- New "Data Explorer" tab: KPIs (customers, employees, products, source
row totals, sample revenue) + charts driven by Elasticsearch aggregations
(revenue over time, by region/channel/status, top customers & products,
HR by department/role/event, supply chain by type, telemetry averages),
with region + free-text filters.
- Backend /api/search/business endpoint: battery of ES aggregations with
numeric/date index template (so amount sums and date histograms work),
source totals via cheap planner estimates; biased the orders sample to
rows carrying customer names so people are searchable/visible.
Manual data entry on source systems:
- "Insert row" action in the Data Hub browser opens a column-aware form;
POST /api/sql/row/insert writes to Postgres/MySQL/Mongo/Cassandra/Neo4j.
- Inserts into CDC sources (PG/MySQL/Mongo) are captured by Debezium and
streamed to Kafka in real time; UI flags this and pulses the data flow.
Build dedicated de-duplicated entity indices (atc-customers ~39k,
atc-employees ~39k, atc-products) from Trino so users can search by
customer/employee name, id and email directly instead of scanning raw
transaction rows. Raise per-table sample caps for the big source tables
(sales_orders, employee_events, supplychain events, device_metrics) and
all iceberg/Hadoop tables for broader row-level coverage.
UI: add Datasets quick-filter chips (Customers, Employees, Products,
Sales orders, HR events, Supply chain, Telemetry, Lakehouse) and default
the Search tab to business data instead of an infra-looking empty view.
- Each pipeline node is now a button; clicking shows a detail card (what the step does + the LangChain/tech component)
- Stronger active-stage pulse (expanding ring + brighter glow)
- Chat auto-opens the How-it-works panel so live per-stage pulsing is always visible while answering
- rag-api: new POST /chat/stream SSE endpoint emits live pipeline stages (embed/retrieve/context/llm/answer) and streams LLM tokens
- Knowledge Chat: document chat now streams answers token-by-token and pulses each pipeline stage in real time as it executes
- How-it-works panel: active stage glows/scales, completed stages settle, agent mode also drives the pulse
- LangChain made visible: orchestrated-by-LangChain badge + LC markers on LangChain-native nodes (TextSplitter, Embeddings, as_retriever, ChatOpenAI, Chroma)
Add a collapsible, pulsing architecture diagram to the Knowledge Chat that
visualizes the real RAG pipeline in two rows: Index (Upload -> Docling ->
Split -> Embed/MiniLM -> ChromaDB) and Query (Question -> Embed -> Retrieve
-> Context -> LLM -> Answer), badged with the actual stack (LangChain,
ChromaDB, Docling, sentence-transformers).
Connectors have a flowing dashed track plus a traveling glow dot and nodes
pulse; the whole flow speeds up while the chat/ingest is busy. Honors
prefers-reduced-motion.
When switching source engine, the sample effect fired with the previous
engine selected object (e.g. public.sales_orders against MySQL), surfacing
"Table hr.sales_orders doesn't exist". Track which engine the selected
object belongs to and only sample when it matches the active engine, plus
guard against stale responses overwriting newer ones.
Row-count for the browser used an exact count(*) which full-scanned huge
tables (~30s on 24M rows). Use planner/statistics estimates and only run a
time-bounded exact count for small (<=50k) tables.
Add a /preview endpoint that range-reads the first chunk of an object and
returns it as text. Browser gets an eye action (list + gallery) opening a
modal that pretty-prints JSON, renders CSV/TSV as a table, and shows logs/
text/yaml/xml verbatim, with a truncation note and download link.
Add a List/Gallery toggle to the bucket browser that auto-switches to a
thumbnail grid when a listing is mostly images, and a keyboard-navigable
lightbox (prev/next/esc) for full-size previews. Download endpoint serves
inline with the correct image/* content-type (ECS stores octet-stream) so
previews render instead of forcing a download.
Analytics now walks every top-level prefix with its own budget instead of a
flat alphabetical scan, so a single huge prefix (kafka/ CDC json) no longer
hides the rest — composition now correctly reflects images, logs, csv,
parquet and avro, and reports true total size.
Browser gains recursive (whole-subtree) listing, free-text name search,
file-type filter chips, type column and Load-more pagination so every object
across all folders is discoverable.
Add /api/storage/s3/analytics endpoint that scans buckets (bounded +
cached) to compute total size/objects, per-bucket distribution, file-type
and size-class breakdowns, cumulative data-growth timeline and largest /
recent objects. Track every S3 op routed through the API for a live
storage-activity timeline.
Rebuild StorageView into a tabbed view: Overview (KPI cards + SVG charts:
growth area, bucket donut, type/size bars, activity sparkline, top folders,
largest & recent objects) and the original bucket Browser.
Add Graph sub-tab in Data Sources UI for Neo4j: force-directed SVG view of
Product-Supplier nodes and SUPPLIES/RELATED_TO/PART_OF/COMPATIBLE_WITH edges.
Backend GET /api/sql/graph/neo4j with rel_type filter and edge limit.
New "Data Sources UI" tab with per-database Browser (catalog + sample data),
Query Console (SqlWorkbench) and embedded interactive Shell (DbShell via SSH).
Backend:
- Extend sql_console.py with Cassandra (CQL) + Neo4j (Cypher) engines
- Add GET /api/sql/catalog/{engine} and GET /api/sql/sample/{engine}
- ssh_terminal: optional initial_command for auto-launching DB CLIs
Frontend:
- DataSourcesView with 5-DB rail, health dots, Browser/Console/Shell sub-tabs
- DbShell embedded xterm terminal with docker exec CLI per engine
- Deep-link topology DB nodes to Data Sources UI (no SQL dock on platform)
- WorkbenchPanel restricted to agent mode only — frees dashboard space
Assign owners to all 177 tables and descriptions + tier tags to the 18
business tables, then trigger SearchIndexing + DataInsights apps so the
platform Data Insights page shows real totals, description coverage and
tier distribution instead of empty plots.
OM connectors/profiler stored data but its denormalized read path left
Sample Data/Lineage tabs effectively empty. This script populates OM directly:
- real 50-row sample data for source + iceberg curated tables
- table/column profiles (column profiles read back correctly in UI)
- full traceable lineage: generator -> source -> Debezium/Kafka CDC topic
-> S3 archive + Iceberg curated -> Trino query layer
Profiler:
- pg/mysql: table-level metrics via DB statistics (exact-ish row counts) only;
source tables hold 24-53M rows and OM's column profiler full-scans per column
(no TABLESAMPLE pushdown), which would hammer the live CDC source.
- trino/iceberg: full column metrics + sample data on the small (20-25k row)
curated_masked + hadoop tables.
Data quality: 16 test cases across 5 tables (row counts on big sources; row
count + uniqueness/not-null/range on curated/masked + historical), all passing.
This populates OM's Data Quality dashboard (coverage, healthy assets,
dimensions, test results).
The OM MySQL datadir was bind-mounted on the 448MB /opt partition and kept
filling up. Moved db-data to /var/lib/docker/openmetadata-db-data (32GB, ~31GB
free) and kept binlog disabled + redo log capped. /opt back to 7% used.
OM 1.13 cannot deserialize Airflow 3.0.6 serialized DAGs, so pipelines had no
tasks. af_tasks.py parses serialized_dag and PATCHes task lists; af_lineage.py
adds generator->table and mask_to_curated/hadoop_to_trino lineage edges.
Note: OM MySQL datadir lives on the 448MB /opt partition which had filled up
(103MB binlog). Disabled binary logging + shrank redo log capacity in the OM
docker-compose to restore headroom.
- ES: connect over https with verify_certs=false (self-signed CA); 16 indexes incl source_data
- Airflow: ingest 8 DAGs as pipelines from SQLite metadata DB snapshot
- es_run.py: standalone runner that pins the working ES TLS config
- Command Center API now wired with ELASTIC_PASSWORD
Widen Trino ingestion to every catalog (postgres_sales/mysql_hr/mongodb/iceberg)
and add messaging (Kafka), storage (S3 object store), dashboards (Superset) and
search (Elasticsearch) service ingestion configs so OpenMetadata reflects the
whole running platform. Secrets are injected at run time, not committed.
Align nodes into pipeline columns (producers -> sources -> CDC -> lakehouse ->
Trino) with governance centred at the bottom, so lineage reads in order instead
of Iceberg/Trino floating mid-canvas.
Query postgres/mysql/mongo directly (fast, early LIMIT) instead of full Trino
table scans; Trino remains the path for the curated lakehouse + as a fallback.
- pii_catalog: persistent per-column masking policy (default masked); GET/POST
/api/pii/policy and POST /api/pii/lookup which redacts masked values server-side.
- get_pii masked flag now reflects the policy; dataflow exposes the dataset key.
- Data Flow PII inspector: per-column lock/unlock toggles + mask-all/unmask-all,
so operators control exactly which data the assistant may reveal.
- decide_approval now executes the underlying data movement when an approved
request carries an executor=movement payload (real human-gated executor).
- rag-api service gets OPENMETADATA_* (via atc.env) + COMMAND_CENTER_URL so it can
sync the catalog and call platform tools.
- Knowledge Chat gains an Agent-mode toggle (SSE tool-loop with step chips) and a
'Sync catalog' button.
lab_context now includes an in-process governance section so the LLM knows the
live CDC volume, data movements + last-run state, the PII catalog (OM+heuristic)
with masked/unmasked status, and source->masked lineage.
- deploy/airflow/mask_to_curated_dag.py: Trino-SQL masking (hash/redact/generalize)
of PII from postgres/mysql sources into iceberg.curated_masked.* (verified:
14k+11k masked rows, all PII masked).
- OM source->masked lineage edges created; curated tables cataloged.
- Version OM compose + ingestion configs + om_api helper under deploy/openmetadata.