Adds native Command Center features (no new containers) integrated as sub-tabs
in the existing Data Explorer and Data Quality views:
- Continuous Data Quality (dq_monitor.py): live completeness/uniqueness/validity/
freshness scorecards via Trino with rolling trends → DataQuality "Live Monitoring".
- Ownership & stewardship (catalog_governance.py): owner/steward/tier matrix,
orphan detection, business glossary; local store best-effort synced to
OpenMetadata (owner PATCH) → Data Explorer "Ownership".
- Access & policy posture: per-dataset compliance combining PII masking, ownership,
live DQ and observability alerts vs data contracts → Data Explorer "Access & Policies".
- Lineage (lineage.py): staged source→CDC→Spark→S3→Iceberg→Trino→serving graph with
live row counts and column-level PII/masking tracing → Data Explorer "Lineage".
- Observability (observability.py): volume/freshness/schema-drift monitoring with
alerts → Data Explorer "Observability".
- Shared lake_meta.py dataset registry + bounded Trino client; fast native row-count
and PK-indexed freshness so monitors stay cheap on 25-54M-row tables.
- LLM context (lab_context.py) enriched with DQ scores, ownership and active alerts.
- New "Live" tab: auto-polls /api/federated/live every 2.5s with
animated counters, ingestion throughput sparkline, per-source
write-rate bars, a region scorecard matrix (heat-shaded) and live
business breakdown charts (region/channel/status/customers/
telemetry/supply)
- Backend /api/federated/live: instant source estimates (Postgres),
monotonic max(event_id) for MySQL and Mongo estimated count for
immediate movement, Cassandra from cached matrix; business aggs
cached over the small Hadoop lake tables (short TTL)
- Fix embedded Trino panels being squeezed with internal scrollbars
by making panels/grids shrink-0 so the page scrolls instead
- Data Explorer now hosts Business Overview, Federated (Trino),
Hadoop Lake and Data Dictionary as tabs instead of a separate
top-level view (richer, less stacked layout)
- TrinoFederationView supports an embedded mode driven by an
external active tab
- Remove the standalone Trino Federation sidebar entry/route
- Remove the Kibana shortcut from the sidebar (it already lives in
the Elasticsearch view)
Add a dedicated business-data view (separate from infra/search):
- New "Data Explorer" tab: KPIs (customers, employees, products, source
row totals, sample revenue) + charts driven by Elasticsearch aggregations
(revenue over time, by region/channel/status, top customers & products,
HR by department/role/event, supply chain by type, telemetry averages),
with region + free-text filters.
- Backend /api/search/business endpoint: battery of ES aggregations with
numeric/date index template (so amount sums and date histograms work),
source totals via cheap planner estimates; biased the orders sample to
rows carrying customer names so people are searchable/visible.
Manual data entry on source systems:
- "Insert row" action in the Data Hub browser opens a column-aware form;
POST /api/sql/row/insert writes to Postgres/MySQL/Mongo/Cassandra/Neo4j.
- Inserts into CDC sources (PG/MySQL/Mongo) are captured by Debezium and
streamed to Kafka in real time; UI flags this and pulses the data flow.