7.9 KiB
name, description
| name | description |
|---|---|
| metadata-validator | Use when validating that a data requirement's candidate tables, fields, joins, and time columns actually exist in the project's metadata. Triggers include "validate metadata", "check if table exists", "field mapping", "join key validation", "确认表/字段", "校验 schema", "这个字段在哪张表", "口径对得上吗". Reads metadata files recursively across the workspace (JSON / Markdown / CSV / Excel / YAML / DDL / data dictionary) and produces a standardized Validation Result and Field Mapping. |
Metadata Validator
Overview
Bridge between the requirement's expected tables/fields and the actual schema living in workspace metadata. Output is consumed verbatim by logic-planner and sql-context-builder — they must NOT re-validate.
Core principle: never guess a table name, field name, or join relationship. If a candidate is not provably present in metadata, surface it as a pending_question and stop at status NEED_USER_CONFIRMATION.
When to Use
Use when:
requirements-analysishas produced candidate tables/fields and you need to confirm they exist in real metadata- The user asks "does this table/field exist", "which file describes table X", "map business field to physical column"
- You need to detect join keys, time columns, or aggregation-suitable numeric columns
- The user is about to write SQL and you must guard against hallucinated column names
Do NOT use when:
- The user has already supplied exact table+column names and confirmed them — go straight to
logic-planner - The task is code generation without any data warehouse context
Inputs
- From
requirements-analysis:candidate_tables: list of table names mentioned or impliedcandidate_fields: list of business field names with their proposed tablebusiness_logic: short summary of what is being computed
- Workspace metadata: any file under the working directory matching metadata conventions (see below)
Workflow
digraph metadata_validator {
"Receive candidate tables/fields" [shape=box];
"Recursively scan workspace for metadata files" [shape=box];
"Parse each file into canonical (table, field, type) records" [shape=box];
"Check candidate tables exist" [shape=box];
"Check candidate fields exist in their table" [shape=box];
"Fuzzy match missing fields against parsed metadata" [shape=box];
"Detect join keys & time columns" [shape=box];
"All checks pass?" [shape=diamond];
"Build Field Mapping" [shape=box];
"Emit Validation Result (VALIDATED)" [shape=box];
"Emit Validation Result (NEED_USER_CONFIRMATION)" [shape=box];
"Receive candidate tables/fields" -> "Recursively scan workspace for metadata files";
"Recursively scan workspace for metadata files" -> "Parse each file into canonical (table, field, type) records";
"Parse each file into canonical (table, field, type) records" -> "Check candidate tables exist";
"Check candidate tables exist" -> "Check candidate fields exist in their table";
"Check candidate fields exist in their table" -> "Fuzzy match missing fields against parsed metadata";
"Fuzzy match missing fields against parsed metadata" -> "Detect join keys & time columns";
"Detect join keys & time columns" -> "All checks pass?";
"All checks pass?" -> "Build Field Mapping" [label="yes"];
"All checks pass?" -> "Emit Validation Result (NEED_USER_CONFIRMATION)" [label="no"];
"Build Field Mapping" -> "Emit Validation Result (VALIDATED)";
}
Metadata Discovery
-
Recursively scan the current working directory (and any subdirectories) — do not assume a
metadata/folder exists. -
Filenames often resemble table names (e.g.
loan_order.json,customer_info.md,dim_product.csv) but never trust the filename alone — open the file and confirm the table identifier inside. -
Accepted formats (auto-detect by content, not extension):
- JSON / JSONL
- YAML
- Markdown tables / headings
- CSV (with header)
- Excel (
.xlsx,.xls) — useopenpyxlorpandas.read_excel - DDL /
CREATE TABLEstatements - Free-form data dictionaries (parse key-value lines or
field | type | commenttables)
-
Convert every file into the canonical record shape before validating:
tables: - table: <physical_table_name> source_file: <relative_path> fields: - name: <field_name> type: <string|int|long|double|decimal|date|timestamp|boolean|...> comment: <optional>
Validation Rules
Run all of the following. Any failure ⇒ status NEED_USER_CONFIRMATION.
- Table existence — every candidate table must appear in the parsed metadata.
- Field existence — every candidate field must appear in its candidate table.
- Field ownership — a field may exist but belong to a different table; never silently swap.
- Type sanity — aggregation targets (sum / avg / count) must be numeric; time filters must be date/timestamp/string-encoded-date.
- Necessary fields — flag if common fields are missing for the business logic (stat metric, time, join key, dimension).
- Fuzzy match — for every missing field, score candidates by name similarity (
customer_id↔cust_id,cust_no,customer_no) and Levenshtein distance; surface top-N. - Join detection — across the validated tables, look for shared
*_id,*_no,*_codecolumns. If multiple plausible join keys exist, ask. - Time column disambiguation — list every time-like field per table (
create_time,apply_time,txn_date,update_time,dt); ask which one defines the time grain. - Coverage — verify that dimensions, metrics, filters, and sort keys are all present.
Forbidden Behaviors
| Rationalization | Reality |
|---|---|
"The filename is loan_order.json, must be that table" |
Filename is a hint. Parse the file to confirm. |
| "Field exists somewhere, so the join is fine" | Wrong table ownership invalidates the join. Always check ownership. |
"The user probably means cust_no, just use it" |
Surface as a candidate and ASK. Never auto-substitute. |
| "I'll guess the time field" | Multiple time fields ⇒ ASK. The grain defines the entire query. |
| "This metadata is enough, no need to scan further" | Always scan recursively. One file can describe many tables. |
Output Contract
The output is a single YAML document. Downstream skills consume it as-is.
validation_result:
validated_tables:
- <table_name>
validated_fields:
- <table_name>.<field_name>
missing_tables: []
missing_fields:
- <field_name_not_found_anywhere>
candidate_fields:
<missing_field_name>:
- <candidate_1>
- <candidate_2>
candidate_tables:
<missing_table_name>:
- <candidate_1>
join_candidates:
- left: <table_a>.<field>
right: <table_b>.<field>
confidence: <high|medium|low>
time_field_candidates:
<table_name>:
- <field>
- <field>
metadata_sources:
- <relative_path_to_metadata_file>
pending_questions:
- "<human-readable question, ideally with options A/B/C>"
field_mapping:
<business_field_name>:
table: <physical_table>
column: <physical_field>
type: <data_type>
status: <VALIDATED | NEED_USER_CONFIRMATION>
field_mapping is the single source of truth for downstream logic-planner and sql-context-builder. If a business field has no mapping, the request cannot proceed — add it to pending_questions.
Completion Criteria
Status VALIDATED is only allowed when all of the following hold:
- Every candidate table is present in
validated_tables - Every candidate field has a
field_mappingentry - All join keys are decided (or trivially obvious with
highconfidence) - The time column for the query grain is decided
pending_questionsis empty
Anything else ⇒ NEED_USER_CONFIRMATION and stop. Do not proceed to logic-planner.