21af36a591
Profiler: - pg/mysql: table-level metrics via DB statistics (exact-ish row counts) only; source tables hold 24-53M rows and OM's column profiler full-scans per column (no TABLESAMPLE pushdown), which would hammer the live CDC source. - trino/iceberg: full column metrics + sample data on the small (20-25k row) curated_masked + hadoop tables. Data quality: 16 test cases across 5 tables (row counts on big sources; row count + uniqueness/not-null/range on curated/masked + historical), all passing. This populates OM's Data Quality dashboard (coverage, healthy assets, dimensions, test results).
29 lines
759 B
YAML
29 lines
759 B
YAML
source:
|
|
type: TestSuite
|
|
serviceName: atc_data_quality
|
|
sourceConfig:
|
|
config:
|
|
type: TestSuite
|
|
entityFullyQualifiedName: "atc_postgres.postgres.public.sales_orders"
|
|
processor:
|
|
type: orm-test-runner
|
|
config:
|
|
testCases:
|
|
# Cheap row-count check only: this is the 53M-row live CDC source table.
|
|
- name: sales_orders_has_rows
|
|
testDefinitionName: tableRowCountToBeBetween
|
|
parameterValues:
|
|
- name: minValue
|
|
value: "1"
|
|
- name: maxValue
|
|
value: "10000000000"
|
|
sink:
|
|
type: metadata-rest
|
|
config: {}
|
|
workflowConfig:
|
|
openMetadataServerConfig:
|
|
hostPort: http://openmetadata-server:8585/api
|
|
authProvider: openmetadata
|
|
securityConfig:
|
|
jwtToken: "__JWT__"
|