{"openapi":"3.1.0","info":{"title":"CityTap API","version":"1.0.0","description":"The public read surface for the citytap catalog of mirrored NYC Open Data: search, canonical metadata with both freshness axes, conformed dimensions, the ontology bundle, read-only SQL and catalog-confirmed relates through a sandboxed executor, the graph's named-query registry, standards feeds, and the citytap MCP server. Anonymous throughout; routes needing sign-in are listed under x-citytap-authenticated. Rate-limit semantics per class are under x-citytap-rate-classes; a 429 always carries Retry-After in seconds. Named refusals with a fixed status and wait are indexed under x-citytap-refusals."},"servers":[{"url":"/"}],"paths":{"/api/analyses":{"get":{"operationId":"analyses-list","summary":"Published cross-dataset analyses, newest first, superseded ones withheld.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/datasets/suggest":{"get":{"operationId":"datasets-suggest","summary":"Type-ahead for dataset address fields: name matches first, metadata matches second.","x-citytap-auth-class":"anonymous-suggest-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/dataset-suggest"}}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/datasets/{fourFour}":{"get":{"operationId":"dataset-detail","summary":"Canonical metadata, both freshness axes, columns, LL11 attribution, lineage.","x-citytap-auth-class":"anonymous-rate-limited","parameters":[{"name":"fourFour","in":"path","required":true,"schema":{"type":"string"},"description":"A Socrata four-by-four dataset identifier."}],"responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/dataset-detail"}}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/datasets/{fourFour}/bindings/{partner}":{"get":{"operationId":"pair-bindings","summary":"Confirmed shared dimension bindings between two datasets, from the catalog's own tags.","x-citytap-auth-class":"anonymous-rate-limited","parameters":[{"name":"fourFour","in":"path","required":true,"schema":{"type":"string"},"description":"A Socrata four-by-four dataset identifier."},{"name":"partner","in":"path","required":true,"schema":{"type":"string"},"description":"The partner dataset's Socrata four-by-four."}],"responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/datasets/{fourFour}/partners":{"get":{"operationId":"partners","summary":"Ranked, dimension-tag-confirmed relatives of one dataset, plus a same-domain bonus.","x-citytap-auth-class":"anonymous-rate-limited","parameters":[{"name":"fourFour","in":"path","required":true,"schema":{"type":"string"},"description":"A Socrata four-by-four dataset identifier."}],"responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/partners"}}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/dimensions":{"get":{"operationId":"dimensions-list","summary":"Every conformed-dimension family, each with its LIVE tagged count.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/dimensions"}}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/dimensions/{key}/datasets":{"get":{"operationId":"dimension-datasets","summary":"One family's tagged members, paginated, with live counts — \"everywhere this field exists\".","x-citytap-auth-class":"anonymous-rate-limited","parameters":[{"name":"key","in":"path","required":true,"schema":{"type":"string"},"description":"A conformed-dimension family key."}],"responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/field-coverage":{"get":{"operationId":"field-coverage","summary":"Field-coverage rollup: the quadruple, the per-dataset bands, and the graph web's reach.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/field-coverage-response"}}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/field-coverage/datasets":{"get":{"operationId":"field-coverage-datasets","summary":"Datasets ranked by remaining analysis work, filterable by band, dimension and agency.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/field-coverage-datasets-response"}}}},"400":{"description":"band names something outside the six, or order names something outside the six orderings. Refused rather than answered with an empty list, which is what a real band holding no datasets returns byte for byte.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/field-coverage/datasets/{fourFour}":{"get":{"operationId":"field-coverage-dataset","summary":"Every column of one dataset with its verdict, its method, and whether anybody read it.","x-citytap-auth-class":"anonymous-rate-limited","parameters":[{"name":"fourFour","in":"path","required":true,"schema":{"type":"string"},"description":"A Socrata four-by-four dataset identifier."}],"responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/field-coverage-dataset-response"}}}},"400":{"$ref":"#/components/responses/refused"},"404":{"description":"No such dataset in this catalog. Distinct from a dataset that HAS no columns, which answers 200 with an empty fields array and band no-inventory — 603 datasets on prod are in that state and it is a harvest gap rather than a typo.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/field-coverage/web":{"get":{"operationId":"field-coverage-web","summary":"The dataset-to-dimension web as nodes and edges, for a rendered graph.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/field-coverage-web-response"}}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/geometry/{fourFour}":{"get":{"operationId":"geometry-layer","summary":"One dataset's geometry as a GeoJSON FeatureCollection, resolved from the mirrored mart and capped. No view ever receives a raw wktLiteral; the conversion happens here.","x-citytap-auth-class":"anonymous-rate-limited","parameters":[{"name":"fourFour","in":"path","required":true,"schema":{"type":"string"},"description":"A Socrata four-by-four dataset identifier."}],"responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/graph/expand":{"post":{"operationId":"graph-expand","summary":"One node's expansion, through the instance's own five-slot explore-graph configuration.","x-citytap-auth-class":"anonymous-rate-limited","requestBody":{"required":true,"content":{"application/json":{"schema":{"$ref":"#/components/schemas/graph-expand-request"}}}},"responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"400":{"$ref":"#/components/responses/refused"},"429":{"$ref":"#/components/responses/rate-limited"},"503":{"description":"The graph store cannot answer this now; the caller's request is not at fault. refusal names the class and Retry-After gives the seconds to wait: graph-credential-refused (the store refused serve's reader credential, which a wrong password and a user granted nothing both cause; 60s), graph-store-unreachable (the store did not answer at all; 15s), graph-store-fault (the store answered with a fault; 15s), graph-store-timeout (the store ran past the query's wall-clock budget; retry once, and if it times out again, narrow the request (a smaller limit or a more specific node), because a query that always exceeds the time cap never succeeds by retrying; 15s), graph-generation-missing (the published graph generation is gone and no successor answered; 15s), graph-unpublished (this environment has published no graph generation yet; 300s), graph-unconfigured (expand only: this environment publishes no explore-graph configuration; 3600s). A 503 with no refusal is serve's own graph query queue, not the store: it is full, or deep with no cached answer for this request, and Retry-After gives the seconds to wait.","headers":{"Retry-After":{"schema":{"type":"integer"},"description":"Seconds to wait before retrying."}},"content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}},"x-citytap-request-body-cap-bytes":16384,"x-citytap-truncation":"Answers carry truncated and true_degree. When truncated is true, the link set is a clipped, deterministic subset and true_degree is the node's real edge count."}},"/api/graph/query":{"post":{"operationId":"graph-query","summary":"A named query from the published registry, with typed arguments. Never caller-supplied command text.","x-citytap-auth-class":"anonymous-rate-limited","requestBody":{"required":true,"content":{"application/json":{"schema":{"$ref":"#/components/schemas/graph-query-request"}}}},"responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/graph-query-response"}}}},"400":{"$ref":"#/components/responses/refused"},"429":{"$ref":"#/components/responses/rate-limited"},"503":{"description":"The graph store cannot answer this now; the caller's request is not at fault. refusal names the class and Retry-After gives the seconds to wait: graph-credential-refused (the store refused serve's reader credential, which a wrong password and a user granted nothing both cause; 60s), graph-store-unreachable (the store did not answer at all; 15s), graph-store-fault (the store answered with a fault; 15s), graph-store-timeout (the store ran past the query's wall-clock budget; retry once, and if it times out again, narrow the request (a smaller limit or a more specific node), because a query that always exceeds the time cap never succeeds by retrying; 15s), graph-generation-missing (the published graph generation is gone and no successor answered; 15s), graph-unpublished (this environment has published no graph generation yet; 300s), graph-unconfigured (expand only: this environment publishes no explore-graph configuration; 3600s). A 503 with no refusal is serve's own graph query queue, not the store: it is full, or deep with no cached answer for this request, and Retry-After gives the seconds to wait.","headers":{"Retry-After":{"schema":{"type":"integer"},"description":"Seconds to wait before retrying."}},"content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}},"x-citytap-request-body-cap-bytes":16384,"x-citytap-truncation":"Answers carry truncated, true_count and row_limit. When truncated is true, every count in that answer is a floor computed over the clipped row set, never a total. true_count is the size of the whole result set — a number on an unclipped answer, and null on a clipped one, because the registry publishes no count form of a named query, so null means nothing measured this rather than zero. row_limit names the cap whether or not it bit, so rows.length landing on a round number is separable from rows.length landing on the ceiling."}},"/api/mcp":{"post":{"operationId":"mcp","summary":"The citytap MCP server, JSON-RPC 2.0. A curated tool set over TYPED arguments — no caller text ever becomes command text — reaching the graph only through the metered answer path. The session passcode is open and is a fairness key, never a credential.","x-citytap-auth-class":"anonymous-session-labelled","responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"}},"x-citytap-protocol":"JSON-RPC 2.0 over one POST endpoint: tools/list enumerates the curated tools, tools/call runs one. Send a stable, RANDOM x-citytap-session value — a UUID or 16+ random characters. It is an open fairness key between concurrent sessions, never a credential, and callers who pick the same string share one bucket, so a guessable value means sharing a bucket with strangers."}},"/api/ontology":{"get":{"operationId":"ontology","summary":"The shared ontology for a scope: the concept scheme, one row per column bound or not, and the coverage quadruple, labelled with the population it counts. Unscoped returns the vocabulary and the catalog-wide coverage rather than every column. A concept key that names nothing in the vocabulary is refused with 404, never an empty bundle.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/ontology-response"}}}},"400":{"$ref":"#/components/responses/refused"},"404":{"description":"The ?concept= key names no concept in the vocabulary. Refused rather than answered with an empty bundle, which a real concept that nothing binds would produce byte for byte. Every valid key ships in this endpoint's own concepts array.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}},"429":{"$ref":"#/components/responses/rate-limited"}},"x-citytap-caps":{"datasets_per_bundle":25,"bindings_per_bundle":5000},"x-citytap-coverage":"coverage and catalog_coverage each carry a describes sentence naming the population they count; they are different populations and must not be compared without reading both. A bucket that cannot be computed at the requested scope is null with an entry in not_computable, never 0 — under ?concept= that is insufficient and unanalysed, neither of which can name a concept."}},"/api/openapi.json":{"get":{"operationId":"openapi","summary":"The OpenAPI 3.1 manifest for this surface: every public route with its request and response schemas, caps, rate classes and truncation semantics; sign-in-only routes are listed as stubs.","x-citytap-auth-class":"anonymous-manifest-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/query":{"post":{"operationId":"query","summary":"Read-only SQL, executed inside the sandboxed query-executor process (R10). Never in brain, never in serve.","x-citytap-auth-class":"anonymous-query-rate-limited","requestBody":{"required":true,"content":{"application/json":{"schema":{"$ref":"#/components/schemas/query-request"}}}},"responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/query-response"}}}},"400":{"$ref":"#/components/responses/refused"},"429":{"$ref":"#/components/responses/rate-limited"},"503":{"description":"The read plane cannot take this statement now, with Retry-After in seconds: refusal busy (saturated, shutting down, or no answer inside serve's backstop) or replica-conflict (the replica cancelled it to apply replication). A 503 with no refusal means the read plane is not answering at all: no read replica, or the query executor is unreachable.","headers":{"Retry-After":{"schema":{"type":"integer"},"description":"Seconds to wait before retrying."}},"content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}},"x-citytap-request-body-cap-bytes":25000,"x-citytap-caps":{"estimated_rows_gate":5000000,"returned_rows_cap":50000,"statement_timeout_ms":30000},"x-citytap-refusals":"A statement that runs past its wall-clock budget is cancelled and refused 400 with refusal timeout; narrow it or ask for fewer rows. When the read plane is saturated, or the read replica cancelled the statement to apply replication, the answer is 503 with refusal busy or replica-conflict and a Retry-After header; retry after that many seconds.","x-citytap-routing-note":"routing.read_host_degraded reports which read-host candidate answered. true means the environment is serving below its preferred replica under reduced caps — an environment posture, not an error, and normal in single-instance environments."}},"/api/refresh/capacity":{"get":{"operationId":"refresh-capacity","summary":"Whether the public refresh lane is open and how much of this environment's hourly portal allowance is left. One scalar read, no portal cost.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"}}}},"/api/refresh/{fourFour}":{"post":{"operationId":"refresh-submit","summary":"R8's path-param refresh submit — same gatekeeper proxy as POST /api/refresh, URL-addressed.","x-citytap-auth-class":"anonymous-refresh-rate-limited","parameters":[{"name":"fourFour","in":"path","required":true,"schema":{"type":"string"},"description":"A Socrata four-by-four dataset identifier."}],"responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"},"500":{"description":"refusal refresh-gatekeeper-uninterpretable: the refresh gatekeeper answered with something other than a verdict. It may have recorded the request, so no wait is given; a repeat is deduplicated as in_progress if it was recorded.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}},"503":{"description":"refusal refresh-gatekeeper-unreachable: the refresh gatekeeper, which owns the portal budget, is not answering, so no verdict exists and nothing was recorded. Retry-After (15s, also retryAfterSeconds in the body) gives the seconds to wait; the same request is safe to repeat.","headers":{"Retry-After":{"schema":{"type":"integer"},"description":"Seconds to wait before retrying."}},"content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}}}},"/api/relate":{"post":{"operationId":"relate","summary":"Compile and run a catalog-confirmed relate spec into one join+coverage answer (D7/D8).","x-citytap-auth-class":"anonymous-rate-limited","requestBody":{"required":true,"content":{"application/json":{"schema":{"$ref":"#/components/schemas/relate-request"}}}},"responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/relate-response"}}}},"400":{"$ref":"#/components/responses/refused"},"429":{"$ref":"#/components/responses/rate-limited"},"503":{"description":"The read plane cannot take this statement now, with Retry-After in seconds: refusal busy (saturated, shutting down, or no answer inside serve's backstop) or replica-conflict (the replica cancelled it to apply replication). A 503 with no refusal means the read plane is not answering at all: no read replica, or the query executor is unreachable.","headers":{"Retry-After":{"schema":{"type":"integer"},"description":"Seconds to wait before retrying."}},"content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}},"x-citytap-request-body-cap-bytes":8192,"x-citytap-caps":{"plan_cardinality_cap_rows":500000,"statement_timeout_ms":30000},"x-citytap-refusals":"A well-formed spec can be refused with 400 {error, refusal}; refusal names the failing check (for example binding-family-not-yet-supported, mart-absent, unsupported-measure) so a client can branch without parsing prose. The agency family's zero-overlap answers can additionally carry refusal succession-strict-no-edge-crossing in the 200 envelope: succession=strict found no shared member and the successor-mapped closure would bridge the pair. Every statement of one relate shares a single 30000ms deadline. A statement that runs past its wall-clock budget is cancelled and refused 400 with refusal timeout; narrow it or ask for fewer rows. When the read plane is saturated, or the read replica cancelled the statement to apply replication, the answer is 503 with refusal busy or replica-conflict and a Retry-After header; retry after that many seconds.","x-citytap-agency-family":"dim=agency joins on entity identity through the reviewed catalog.agency_crosswalk — the crosswalk is a joined table, values it does not resolve are excluded and counted (crosswalk_coverage on the response), and there is no raw-equality fallback. succession is mapped by default (recorded succession edges only, DoITT to OTI; containment never merges). Marts past the executor's plan cap serve from the materialized profile sidecar: provenance.{a,b}.resolution says which route answered and sampled labels those counts estimates."}},"/api/standards/evidence":{"get":{"operationId":"standards-evidence","summary":"The published bytes one conformance claim rests on: the pointer, the value at it, and how many records carry the field. An absent field is an answer, not an error — it is what disproves the ledger row above it.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/standards-evidence"}}}},"429":{"$ref":"#/components/responses/rate-limited"},"503":{"description":"refusal lake-unreachable: the artifact store is not answering, so nothing was read. Retry-After (15s, also retryAfterSeconds in the body) gives the seconds to wait; the request only reads, so repeating it is safe. A 503 with no refusal means the feed has not been generated yet in this environment.","headers":{"Retry-After":{"schema":{"type":"integer"},"description":"Seconds to wait before retrying."}},"content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}}}},"/data.json":{"get":{"operationId":"pod-feed","summary":"The registry-generated metadata feed, streamed verbatim from the lake.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/json":{}}},"429":{"$ref":"#/components/responses/rate-limited"},"503":{"description":"refusal lake-unreachable: the artifact store is not answering, so nothing was read. Retry-After (15s, also retryAfterSeconds in the body) gives the seconds to wait; the request only reads, so repeating it is safe. A 503 with no refusal means the feed has not been generated yet in this environment.","headers":{"Retry-After":{"schema":{"type":"integer"},"description":"Seconds to wait before retrying."}},"content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}}}},"/dcat.jsonld":{"get":{"operationId":"dcat-feed","summary":"The registry-generated linked-data feed, streamed verbatim from the lake.","x-citytap-auth-class":"anonymous-rate-limited","responses":{"200":{"description":"The answer.","content":{"application/ld+json":{}}},"429":{"$ref":"#/components/responses/rate-limited"},"503":{"description":"refusal lake-unreachable: the artifact store is not answering, so nothing was read. Retry-After (15s, also retryAfterSeconds in the body) gives the seconds to wait; the request only reads, so repeating it is safe. A 503 with no refusal means the feed has not been generated yet in this environment.","headers":{"Retry-After":{"schema":{"type":"integer"},"description":"Seconds to wait before retrying."}},"content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}}}}},"components":{"schemas":{"conformance":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/conformance.json","title":"GET /api/conformance","description":"The conformance ledger: one row per standard requirement, its implementing artifact, its status, and the catalog.standards_map row it traces to. The requirement TEXT is not stored in the ledger — it is joined from standards_map and standards, so conformance stays data in exactly one place and a claim here cannot drift from the crosswalk the emitters execute.","type":"object","additionalProperties":false,"required":["as_of","rows","totals"],"properties":{"as_of":{"type":"string","format":"date-time"},"totals":{"type":"object","additionalProperties":false,"required":["implemented","partial","planned","not_applicable"],"properties":{"implemented":{"type":"integer","minimum":0},"partial":{"type":"integer","minimum":0},"planned":{"type":"integer","minimum":0},"not_applicable":{"type":"integer","minimum":0}}},"rows":{"type":"array","items":{"type":"object","additionalProperties":false,"required":["canonical_field","standard_id","standard_name","standard_url","requirement_path","requirement_notes","action","direction","implementing_artifact","status","evidence_url"],"properties":{"canonical_field":{"type":"string"},"standard_id":{"type":"string"},"standard_name":{"type":"string"},"standard_url":{"type":["string","null"]},"requirement_path":{"type":["string","null"]},"requirement_notes":{"type":["string","null"]},"action":{"type":"string"},"direction":{"type":"string"},"implementing_artifact":{"type":"string","minLength":1},"status":{"enum":["implemented","partial","planned","not_applicable"]},"evidence_url":{"type":["string","null"]}}}}}},"dataset-detail":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/dataset-detail.json","title":"GET /api/datasets/{four_four}","description":"The canonical metadata record, the two INDEPENDENT freshness axes, the column manifest with its Postgres mapping, the Local Law 11 §23-502(d) attribution triple, and the lineage receipts. Every metadata field is REQUIRED to be present; a field the portal does not publish is explicitly null. Absent and null are different claims and this schema does not let them look alike.","type":"object","additionalProperties":false,"required":["metadata","freshness","columns","attribution","lineage","exports","dictionary_url"],"properties":{"metadata":{"$ref":"#/$defs/metadata"},"freshness":{"$ref":"#/$defs/freshness"},"columns":{"type":"array","items":{"$ref":"#/$defs/column"}},"attribution":{"$ref":"#/$defs/attribution"},"lineage":{"type":"array","items":{"$ref":"#/$defs/run"}},"exports":{"$ref":"#/$defs/exports"},"dictionary_url":{"type":["string","null"]}},"$defs":{"exports":{"description":"Which downloadable artifacts actually exist for this dataset. A client must not offer a download this does not report as available: the only other way to find out is to request the export and read a 404.","type":"object","additionalProperties":false,"required":["parquet"],"properties":{"parquet":{"type":"object","additionalProperties":false,"required":["available","produced_at"],"properties":{"available":{"type":"boolean"},"produced_at":{"type":["string","null"],"format":"date-time"}}}}},"metadata":{"type":"object","additionalProperties":false,"description":"The 36 canonical fields. The count is not asserted anywhere as a number — the contract test derives the expected set from THIS schema, so adding a field means editing one place and both the API and the gate follow.","required":["portal_domain","four_four","asset_type","name","description","agency_id","agency_name","attribution_raw","attribution_link","portal_category","tags","provenance","domain_key","parent_four_four","row_label","row_identifier_field","created_at","published_at","date_made_public","data_updated_at","metadata_updated_at","updated_at","update_frequency","data_change_frequency","update_frequency_details","automation","license_id","license_name","license_url","owner_display","download_count","page_views_total","size_tier","mirror_mode","sync_strategy","release_status","permalink"],"properties":{"portal_domain":{"type":"string"},"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"asset_type":{"type":"string"},"name":{"type":"string"},"description":{"type":["string","null"]},"agency_id":{"type":["string","null"]},"agency_name":{"type":["string","null"]},"attribution_raw":{"type":["string","null"]},"attribution_link":{"type":["string","null"]},"portal_category":{"type":["string","null"]},"tags":{"type":["array","null"],"items":{"type":"string"}},"provenance":{"type":"string"},"domain_key":{"type":["string","null"]},"parent_four_four":{"type":["string","null"]},"row_label":{"type":["string","null"]},"row_identifier_field":{"type":["string","null"]},"created_at":{"type":["string","null"],"format":"date-time"},"published_at":{"type":["string","null"],"format":"date-time"},"date_made_public":{"type":["string","null"]},"data_updated_at":{"type":["string","null"],"format":"date-time"},"metadata_updated_at":{"type":["string","null"],"format":"date-time"},"updated_at":{"type":["string","null"],"format":"date-time"},"update_frequency":{"type":["string","null"]},"data_change_frequency":{"type":["string","null"]},"update_frequency_details":{"type":["string","null"]},"automation":{"type":["string","null"]},"license_id":{"type":["string","null"]},"license_name":{"type":["string","null"]},"license_url":{"type":["string","null"]},"owner_display":{"type":["string","null"]},"download_count":{"type":["integer","null"]},"page_views_total":{"type":["integer","null"]},"size_tier":{"type":["string","null"]},"mirror_mode":{"type":"string"},"sync_strategy":{"type":["string","null"]},"release_status":{"type":["string","null"]},"permalink":{"type":["string","null"]}}},"freshness":{"type":"object","additionalProperties":false,"description":"Two axes, each carrying its own state, its own transition time and its own inputs. Nothing in axis_b is computed from axis_a or the reverse — they answer different questions (does the publisher keep its promise; is our copy current) and a response that derived one from the other would agree with itself while being wrong about the world.","required":["axis_a","axis_b","declared_cadence","observed_cadence"],"properties":{"axis_a":{"type":"object","additionalProperties":false,"required":["freshness_state","since","data_updated_at","effective_period_days","overdue_since","badge","fenced","fences"],"properties":{"freshness_state":{"type":"string"},"since":{"type":["string","null"],"format":"date-time"},"data_updated_at":{"type":["string","null"],"format":"date-time"},"effective_period_days":{"type":["number","null"]},"overdue_since":{"type":["string","null"],"format":"date-time"},"badge":{"enum":["green","neutral","amber","red","grey"]},"fenced":{"type":"boolean","description":"Whether the refresh machinery has fenced this dataset out of refresh selection (catalog.staleness_state.is_zombie). A SECOND axis-A boolean and never a tenth freshness_state: it answers whether anything will act on the state, not what the state is. True means the axis-A state above keeps ageing while no refresh will be attempted, so an OVERDUE dataset that is also fenced will not be repaired by waiting."},"fences":{"type":["array","null"],"description":"Every reason the automatic refresh lane will not refresh this dataset on its own, in citytap-core's REFRESH_FENCES order. The codes are computed by the same exported SQL the refresh dispatcher filters its worklist on, so a fence shown here is a fence applied. [] means the dataset was assessed and nothing fences it; null means it has no catalog.staleness_state row and was never assessed. Wider than `fenced`, which stays is_zombie alone: Wherever this is not null, `fenced` is true exactly when it holds a `zombie` entry.","items":{"type":"object","additionalProperties":false,"required":["reason","on_demand"],"properties":{"reason":{"enum":["lifecycle_closed","zombie","schema_drift","not_mirrored_by_policy","strategy_not_automatic","operator_review_hold"],"description":"lifecycle_closed: RETIRED or SUSPENDED on the portal. zombie: the dataset has stopped changing. schema_drift: a column-set change is detected or frozen. not_mirrored_by_policy: mirror_mode is not 'mirror'. strategy_not_automatic: the sync strategy is neither incremental nor full_replace (bulk_export included). operator_review_hold: flagged monster-scale and not yet cleared by an operator."},"on_demand":{"enum":["served","refused"],"description":"What a Refresh asked for on the dataset page does despite this fence: served means the request is admitted and pulled, refused means some hop declines it. Taken from citytap-core's REFRESH_FENCES, never decided here. It states what the button does, not that the load will succeed."}}}}}},"axis_b":{"type":"object","additionalProperties":false,"required":["sync_state","since","last_synced_at","local_row_count","remote_row_count","remote_row_count_at","mirror_lag_seconds"],"properties":{"sync_state":{"type":"string"},"since":{"type":["string","null"],"format":"date-time"},"last_synced_at":{"type":["string","null"],"format":"date-time"},"local_row_count":{"type":["integer","null"]},"remote_row_count":{"type":["integer","null"]},"remote_row_count_at":{"type":["string","null"],"format":"date-time","description":"When the remote count was last obtained by the reconcile lane. Distinct from `since`, which times the sync transition: the two answer different questions and a reader given only `since` cannot tell a live verification from a stale one."},"mirror_lag_seconds":{"type":["number","null"]}}},"declared_cadence":{"type":["string","null"]},"observed_cadence":{"type":"object","additionalProperties":false,"required":["median_gap_days","declared_period_days"],"properties":{"median_gap_days":{"type":["number","null"]},"declared_period_days":{"type":["number","null"]}}}}},"column":{"type":"object","additionalProperties":false,"required":["field_name","position","display_name","description","source_datatype","pg_name","pg_type","dimension_key","is_row_identifier","unit","dimension_assertion","dimension_expr","dimension_expr_arg"],"properties":{"field_name":{"type":"string"},"position":{"type":"integer","minimum":1},"display_name":{"type":["string","null"]},"description":{"type":["string","null"]},"source_datatype":{"type":"string"},"pg_name":{"type":"string"},"pg_type":{"type":"string"},"dimension_key":{"type":["string","null"]},"is_row_identifier":{"type":"boolean"},"unit":{"type":["string","null"]},"dimension_assertion":{"type":["string","null"],"enum":["declared","inferred","derived",null]},"dimension_expr":{"type":["string","null"],"enum":["identity","month_from_timestamp","month_from_year_month_pair",null]},"dimension_expr_arg":{"type":["string","null"]}}},"attribution":{"type":"object","additionalProperties":false,"description":"Local Law 11 §23-502(d): a redistributor states the SOURCE, the VERSION redistributed and the MODIFICATIONS made. All three are required and the first two are non-empty on every asset; modifications may be null where we have made none, which is itself a claim.","required":["source","version","modifications","publisher","retrieved_at"],"properties":{"source":{"type":"string","minLength":1},"version":{"type":"object","additionalProperties":false,"required":["kind","value"],"properties":{"kind":{"enum":["content-sha256","snapshot-timestamp"]},"value":{"type":"string"}}},"modifications":{"type":["string","null"]},"publisher":{"type":"string"},"retrieved_at":{"type":["string","null"],"format":"date-time"}}},"run":{"type":"object","additionalProperties":false,"required":["run_id","kind","trigger","started_at","finished_at","outcome","rows_upserted","content_sha256","producer_release"],"properties":{"run_id":{"type":"integer"},"kind":{"type":"string"},"trigger":{"type":"string"},"started_at":{"type":"string","format":"date-time"},"finished_at":{"type":["string","null"],"format":"date-time"},"outcome":{"type":["string","null"]},"rows_upserted":{"type":["integer","null"]},"content_sha256":{"type":["string","null"]},"producer_release":{"type":["string","null"]}}}}},"dataset-suggest":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/dataset-suggest.json","title":"GET /api/datasets/suggest","description":"Type-ahead for dataset address fields. `names` holds every dataset whose own name carries the query; `metadata` holds those matched only on description, tags, agency or portal category. The two arrays are never merged and `metadata` is always presented after `names` — a description hit is corroboration, not the dataset the reader named. `ranking` is `lexical`: token containment and trigram similarity, with no semantic model behind it.","type":"object","additionalProperties":false,"required":["q","names","metadata","total_names","total_metadata","limit","ranking"],"properties":{"q":{"type":"string"},"names":{"type":"array","items":{"$ref":"#/$defs/suggestion"}},"metadata":{"type":"array","items":{"$ref":"#/$defs/suggestion"}},"total_names":{"type":"integer","minimum":0},"total_metadata":{"type":"integer","minimum":0},"limit":{"type":"integer","minimum":1},"ranking":{"enum":["lexical"]}},"$defs":{"suggestion":{"type":"object","additionalProperties":false,"required":["four_four","name","agency_name","domain_key","freshness_state","matched_on","snippet","score"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"name":{"type":"string"},"agency_name":{"type":["string","null"],"description":"The publisher: the canonical agency name where the dataset resolves to one, otherwise the publisher's own attribution as published. A display string, never an identifier or a facet key."},"domain_key":{"type":["string","null"]},"freshness_state":{"type":["string","null"]},"matched_on":{"enum":["address","name","description","tags","agency","category"]},"snippet":{"type":["string","null"],"description":"The matched text in context, for a metadata hit. Null on a name or address hit: there is nothing to explain when the reader can already see the match in the name."},"score":{"type":"number"}}}}},"datasets-list":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/datasets-list.json","title":"GET /api/datasets","description":"Search and browse the catalog. `semantic` is reserved for phase 13: the field is part of the shape from v1 so a client written today does not change when pgvector ranking lands, and it is null until then. A client MUST NOT infer from its presence that semantic ranking happened — `ranking` says which one ran.","type":"object","additionalProperties":false,"required":["results","total","limit","offset","ranking","facets","semantic"],"properties":{"results":{"type":"array","items":{"$ref":"#/$defs/summary"}},"total":{"type":"integer","minimum":0},"limit":{"type":"integer","minimum":1},"offset":{"type":"integer","minimum":0},"ranking":{"enum":["lexical","recency"]},"semantic":{"type":"null"},"facets":{"type":"object","additionalProperties":false,"required":["domain","agency","cadence","freshness_state","mirror_mode"],"properties":{"domain":{"$ref":"#/$defs/facet"},"agency":{"$ref":"#/$defs/facet"},"cadence":{"$ref":"#/$defs/facet"},"freshness_state":{"$ref":"#/$defs/facet"},"mirror_mode":{"$ref":"#/$defs/facet"}}}},"$defs":{"facet":{"type":"array","items":{"type":"object","additionalProperties":false,"required":["value","count"],"properties":{"value":{"type":["string","null"]},"count":{"type":"integer","minimum":1},"label":{"type":"string","description":"Human name for an id-valued facet bucket (agency). Absent when the value is its own name or is null."}}}},"summary":{"type":"object","additionalProperties":false,"required":["four_four","name","description","agency_id","agency_name","domain_key","asset_type","update_frequency","freshness_state","sync_state","mirror_mode","data_updated_at","last_synced_at","permalink","score"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"name":{"type":"string"},"description":{"type":["string","null"]},"agency_id":{"type":["string","null"]},"agency_name":{"type":["string","null"]},"domain_key":{"type":["string","null"]},"asset_type":{"type":"string"},"update_frequency":{"type":["string","null"]},"freshness_state":{"type":["string","null"]},"sync_state":{"type":["string","null"]},"mirror_mode":{"type":"string"},"data_updated_at":{"type":["string","null"],"format":"date-time"},"last_synced_at":{"type":["string","null"],"format":"date-time"},"permalink":{"type":"string"},"score":{"type":["number","null"]}}}}},"dimensions":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/dimensions.json","title":"GET /api/dimensions/{key}","description":"One conformed dimension family: the canonical column name, the value convention, every observed alias, and the datasets carrying it. This is the public glue vocabulary — the answer to 'what do I join these two datasets on'.","type":"object","additionalProperties":false,"required":["dimension_key","title","canonical_field_name","description","value_convention","dim_table","aliases","datasets_with_column","notes"],"properties":{"dimension_key":{"type":"string","pattern":"^[a-z_]+$"},"title":{"type":"string"},"canonical_field_name":{"type":"string","description":"The family's canonical column name. On a lexically entered family this is the TSM-standard spelling and it appears among the aliases below. On a crosswalk-entered family it names the RESOLVED identifier a consumer joins on (agency_id), never a published cell — joining on the published cell is the thing that family's value_convention exists to countermand."},"description":{"type":["string","null"]},"value_convention":{"type":["string","null"]},"dim_table":{"type":["string","null"]},"notes":{"type":["string","null"]},"datasets_with_column":{"type":"integer","minimum":0},"aliases":{"type":"array","description":"Every column-name spelling that enters this family, canonical first. This list IS the lexical entry path: catalog.dimension_aliases instructs the column tagger to match dataset_columns.field_name against these strings. It is empty — deliberately, not for want of seeding — on a crosswalk-entered family (the register at $defs/crosswalk_entered_family_key), which is entered by reviewing the VALUES a column carries. An empty list on any other family is a seeding gap, and the if/then/else below refuses it.","items":{"type":"object","additionalProperties":false,"required":["alias","is_canonical","notes"],"properties":{"alias":{"type":"string"},"is_canonical":{"type":"boolean"},"notes":{"type":["string","null"]}}}}},"$defs":{"crosswalk_entered_family_key":{"description":"The family keys entered by a reviewed VALUE crosswalk rather than by column-name spelling. Membership in such a family is decided per column against the values that column carries, so one spelling can be tagged on one column and refused on another: inside agency's own reviewed cohort the spelling 'agency' is 18 tagged and 2 refused, which makes membership provably not a function of the name. A family listed here therefore publishes NO aliases at all — an alias row would instruct the column tagger to tag by spelling the very columns the review refused. Adding a key to this register is a reviewed decision, and it is the only way an empty alias list passes this schema.","enum":["agency"]}},"if":{"required":["dimension_key"],"properties":{"dimension_key":{"$ref":"#/$defs/crosswalk_entered_family_key"}}},"then":{"$comment":"Exactly zero, not 'at least zero'. The permission is narrow in both directions: seeding an alias for a crosswalk-entered family re-opens the lexical name-match path the review replaced, and this branch refuses it rather than publishing it.","properties":{"aliases":{"maxItems":0}}},"else":{"$comment":"Every other family is entered by spelling and names at least its own canonical one. A family published with no spellings at all is a forgotten dimension_aliases seed, which this branch refuses.","properties":{"aliases":{"minItems":1}}}},"error":{"type":"object","required":["error"],"properties":{"error":{"type":"string","description":"One sentence naming the refusal."},"refusal":{"type":"string","description":"The refusal's name, for a client to branch on without parsing error. The read routes (query, relate, rows) answer timeout, busy or replica-conflict; relate also names the catalog or mart check that failed. The graph routes answer a graph-* name on 503 when the graph store cannot answer: graph-credential-refused, graph-store-unreachable, graph-store-fault, graph-store-timeout, graph-generation-missing, graph-unpublished, graph-unconfigured."},"retryAfterSeconds":{"type":"integer","description":"On a retryable 503: the same number of seconds as the Retry-After header."}},"description":"Every refusal body carries error. Refusals a client should branch on also carry refusal."},"field-coverage-dataset-response":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/field-coverage-dataset.json","title":"GET /api/field-coverage/datasets/{fourFour}","description":"Every column of one dataset with its verdict, its method, and whether anybody read it. The leaf of the coverage drill and the only place on this surface where an UNANALYSED column is listed by name — a count of 49,547 undecided columns is not actionable, and the eleven in front of you are. concept_outcome is null when no live concept assertion exists; that is the unanalysed state and is deliberately not a fifth outcome string, so a client cannot confuse 'nobody looked' with a verdict. An unknown four-by-four answers 404, never an empty field list, because a dataset with no columns is a real and different state.","type":"object","additionalProperties":false,"required":["four_four","name","agency_id","domain_key","mart_declared","band","coverage","fields"],"properties":{"four_four":{"type":"string"},"name":{"type":"string"},"agency_id":{"type":["string","null"]},"domain_key":{"type":["string","null"]},"mart_declared":{"type":"boolean"},"band":{"type":"string","enum":["no-inventory","untouched","started","half","most","complete"]},"coverage":{"type":"object","additionalProperties":false,"required":["columns_total","columns_in_scope","bound","refused","insufficient","unanalysed","decided","investigated","dimension_columns","dimensions"],"properties":{"columns_total":{"type":"integer","minimum":0},"columns_in_scope":{"type":"integer","minimum":0},"bound":{"type":"integer","minimum":0},"refused":{"type":"integer","minimum":0},"insufficient":{"type":"integer","minimum":0},"unanalysed":{"type":"integer","minimum":0},"decided":{"type":"integer","minimum":0},"investigated":{"type":"integer","minimum":0},"dimension_columns":{"type":"integer","minimum":0},"dimensions":{"type":"integer","minimum":0}}},"fields":{"type":"array","description":"Every column the catalog holds for this dataset, in position order — covered and uncovered alike.","items":{"type":"object","additionalProperties":false,"required":["field_name","position","display_name","description","source_datatype","pg_name","is_computed_region","is_row_identifier","dimension_key","dimension_assertion","concept_outcome","concept_key","concept_title","concept_method","method_investigated","concept_confidence","evidence_quote","role","role_outcome","review_state"],"properties":{"field_name":{"type":"string"},"position":{"type":"integer","minimum":0},"display_name":{"type":["string","null"]},"description":{"type":["string","null"]},"source_datatype":{"type":["string","null"]},"pg_name":{"type":["string","null"]},"is_computed_region":{"type":"boolean","description":"A portal spatial-join artifact. Inside columns_total, outside columns_in_scope."},"is_row_identifier":{"type":"boolean"},"dimension_key":{"type":["string","null"]},"dimension_assertion":{"type":["string","null"]},"concept_outcome":{"type":["string","null"],"enum":["bound","refused","insufficient_evidence",null],"description":"null means NO live concept assertion — the column is unanalysed."},"concept_key":{"type":["string","null"]},"concept_title":{"type":["string","null"]},"concept_method":{"type":["string","null"]},"method_investigated":{"type":"boolean","description":"false on a bound column means the verdict is a carried lexical tag nobody read."},"concept_confidence":{"type":["number","null"]},"evidence_quote":{"type":["string","null"],"description":"The substring of the column's own description that carried the verdict, verbatim."},"role":{"type":["string","null"]},"role_outcome":{"type":["string","null"]},"review_state":{"type":["string","null"]}}}}}},"field-coverage-datasets-response":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/field-coverage-datasets.json","title":"GET /api/field-coverage/datasets","description":"Datasets ranked by remaining analysis work. Query parameters: band (one of the six), dimension (a conformed-dimension key — restricts to datasets carrying it), agency, order (undecided | decided | columns | bound | unread | name; default undecided), limit (<= 200), offset. An unknown band or order is refused with 400 rather than answered with an empty list, which would read as 'no datasets are in that band'. total counts the FILTERED population, so the page arithmetic is honest under every filter.","type":"object","additionalProperties":false,"required":["total","limit","offset","filters","datasets"],"properties":{"total":{"type":"integer","minimum":0},"limit":{"type":"integer","minimum":1,"maximum":200},"offset":{"type":"integer","minimum":0},"filters":{"type":"object","additionalProperties":false,"required":["band","dimension","agency","order"],"description":"Echoed so a caller can tell an empty page under a filter from an empty catalogue.","properties":{"band":{"type":["string","null"]},"dimension":{"type":["string","null"]},"agency":{"type":["string","null"]},"order":{"type":"string"}}},"datasets":{"type":"array","items":{"type":"object","additionalProperties":false,"required":["four_four","name","agency_id","domain_key","mart_declared","band","columns_total","columns_in_scope","bound","refused","insufficient","unanalysed","decided","investigated","dimension_columns","dimensions"],"properties":{"four_four":{"type":"string"},"name":{"type":"string"},"agency_id":{"type":["string","null"]},"domain_key":{"type":["string","null"]},"mart_declared":{"type":"boolean"},"band":{"type":"string","enum":["no-inventory","untouched","started","half","most","complete"]},"columns_total":{"type":"integer","minimum":0},"columns_in_scope":{"type":"integer","minimum":0},"bound":{"type":"integer","minimum":0},"refused":{"type":"integer","minimum":0},"insufficient":{"type":"integer","minimum":0},"unanalysed":{"type":"integer","minimum":0},"decided":{"type":"integer","minimum":0},"investigated":{"type":"integer","minimum":0,"description":"Decided columns whose method actually read something. bound minus investigated is this dataset's carried-tag debt."},"dimension_columns":{"type":"integer","minimum":0},"dimensions":{"type":"integer","minimum":0,"description":"Distinct conformed dimensions this dataset carries. Zero means it is not in the graph web."}}}}}},"field-coverage-response":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/field-coverage.json","title":"GET /api/field-coverage","description":"How much of the catalogue's field corpus has been analysed. The quadruple (bound / refused / insufficient / unanalysed) is FOUR FACTS and is never collapsed to a ratio here — 'we looked and found nothing' and 'nobody has looked' are different states, and a single percentage reports them as one number. `ledger` repeats the same four straight from semantic.investigation_coverage so a reader can check this surface against the ledger rather than trust it. A dataset with no column inventory at all is a HARVEST gap, banded no-inventory and never folded into unanalysed. Derived percentages are the client's to compute from the counts it was given.","type":"object","additionalProperties":false,"required":["as_of","environment","catalog","coverage","ledger","bands","graph_web","describes"],"properties":{"as_of":{"type":"string","format":"date-time"},"environment":{"type":"string"},"catalog":{"type":"object","additionalProperties":false,"required":["datasets","datasets_with_inventory","datasets_no_inventory","datasets_mart_declared","columns_total","columns_in_scope","columns_computed_region"],"properties":{"datasets":{"type":"integer","minimum":0},"datasets_with_inventory":{"type":"integer","minimum":0},"datasets_no_inventory":{"type":"integer","minimum":0,"description":"Datasets holding no catalog.dataset_columns rows at all. A harvest backlog, not an analysis backlog: there is nothing yet to investigate."},"datasets_mart_declared":{"type":"integer","minimum":0,"description":"Datasets declaring a pg_schema.pg_table. A CLAIM — see catalog.v_declared_marts_absent for whether the relation exists."},"columns_total":{"type":"integer","minimum":0,"description":"semantic.investigation_coverage's denominator; includes the portal's computed-region artifacts."},"columns_in_scope":{"type":"integer","minimum":0,"description":"The investigation queue's denominator; excludes computed regions. A percentage against one denominator is unreconcilable with the other, so both ship."},"columns_computed_region":{"type":"integer","minimum":0}}},"coverage":{"type":"object","additionalProperties":false,"required":["bound","refused","insufficient","unanalysed","decided","investigated"],"properties":{"bound":{"type":"integer","minimum":0},"refused":{"type":"integer","minimum":0,"description":"An investigation ran and REJECTED the candidate concept. A result, not a gap."},"insufficient":{"type":"integer","minimum":0},"unanalysed":{"type":"integer","minimum":0,"description":"No live concept assertion exists. Nobody has looked."},"decided":{"type":"integer","minimum":0},"investigated":{"type":"integer","minimum":0,"description":"Of the decided columns, how many came from a method that actually read something. A bound column from an uninvestigated method is a carried lexical tag."}}},"ledger":{"type":"object","additionalProperties":false,"required":["columns_total","bound","refused","insufficient","unanalysed"],"description":"semantic.investigation_coverage's own row. Must equal the corresponding members of catalog/coverage above; published rather than asserted so a drift is visible instead of hidden.","properties":{"columns_total":{"type":"integer","minimum":0},"bound":{"type":"integer","minimum":0},"refused":{"type":"integer","minimum":0},"insufficient":{"type":"integer","minimum":0},"unanalysed":{"type":"integer","minimum":0}}},"bands":{"type":"array","description":"The six bands, worst work first. Exhaustive and disjoint: datasets sums to catalog.datasets and each column total sums to its catalogue-wide counterpart.","items":{"type":"object","additionalProperties":false,"required":["band","datasets","columns_total","columns_in_scope","decided","bound","refused","insufficient","unanalysed","investigated","mart_declared","in_graph_web"],"properties":{"band":{"type":"string","enum":["no-inventory","untouched","started","half","most","complete"]},"datasets":{"type":"integer","minimum":0},"columns_total":{"type":"integer","minimum":0},"columns_in_scope":{"type":"integer","minimum":0},"decided":{"type":"integer","minimum":0},"bound":{"type":"integer","minimum":0},"refused":{"type":"integer","minimum":0},"insufficient":{"type":"integer","minimum":0},"unanalysed":{"type":"integer","minimum":0},"investigated":{"type":"integer","minimum":0},"mart_declared":{"type":"integer","minimum":0},"in_graph_web":{"type":"integer","minimum":0}}}},"graph_web":{"type":"object","additionalProperties":false,"required":["dimensions","datasets","datasets_without_edge","columns","edges","by_dimension"],"description":"The dataset-to-dimension web's reach. Bipartite: an edge is a dataset carrying a column tagged with a conformed dimension. Sharing a dimension is weaker than joining.","properties":{"dimensions":{"type":"integer","minimum":0},"datasets":{"type":"integer","minimum":0},"datasets_without_edge":{"type":"integer","minimum":0},"columns":{"type":"integer","minimum":0},"edges":{"type":"integer","minimum":0,"description":"One per (dimension, dataset) pair — what the web draws. Not the tagged-column count, which is larger."},"by_dimension":{"type":"array","items":{"type":"object","additionalProperties":false,"required":["dimension_key","title","datasets","columns","investigated_columns","mart_declared"],"properties":{"dimension_key":{"type":"string"},"title":{"type":["string","null"]},"datasets":{"type":"integer","minimum":0},"columns":{"type":"integer","minimum":0},"investigated_columns":{"type":"integer","minimum":0},"mart_declared":{"type":"integer","minimum":0}}}}}},"describes":{"type":"object","additionalProperties":false,"required":["coverage","investigated","bands","graph_web","columns_in_scope"],"description":"One sentence per axis naming the population it counts. Ships in the payload so a reader can disagree with a definition rather than with a number.","properties":{"coverage":{"type":"string","minLength":1},"investigated":{"type":"string","minLength":1},"bands":{"type":"string","minLength":1},"graph_web":{"type":"string","minLength":1},"columns_in_scope":{"type":"string","minLength":1}}}}},"field-coverage-web-response":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/field-coverage-web.json","title":"GET /api/field-coverage/web","description":"The dataset-to-dimension web as nodes and edges, for a rendered graph. BIPARTITE ON PURPOSE — a dimension node and a dataset node, one edge per membership. Dataset-to-dataset edges would be roughly 250,000 lines (borough alone cliques 551 datasets into 151,800 pairs), unrenderable, and a stronger claim than the catalog holds: sharing a dimension is not the same claim as joining. The dimension node states the true fact once and says WHY two datasets relate. Query parameters: dimension (restrict to one family), limit (edges, <= 5000). Dataset nodes are deduplicated across dimensions, and are built from the RETURNED edges, so every edge always has both endpoints even under truncation.","type":"object","additionalProperties":false,"required":["total_edges","returned_edges","truncated","filters","dimension_nodes","dataset_nodes","edges","describes"],"properties":{"total_edges":{"type":"integer","minimum":0,"description":"The unclipped edge count for the requested filter."},"returned_edges":{"type":"integer","minimum":0},"truncated":{"type":"boolean","description":"true means edges were clipped at the limit and every count derived from this payload is a floor, never a total."},"filters":{"type":"object","additionalProperties":false,"required":["dimension","limit"],"properties":{"dimension":{"type":["string","null"]},"limit":{"type":"integer","minimum":1}}},"dimension_nodes":{"type":"array","description":"One per conformed dimension present in the returned edges, biggest first.","items":{"type":"object","additionalProperties":false,"required":["dimension_key","datasets","columns"],"properties":{"dimension_key":{"type":"string"},"datasets":{"type":"integer","minimum":0},"columns":{"type":"integer","minimum":0}}}},"dataset_nodes":{"type":"array","description":"One per dataset present in the returned edges, deduplicated across dimensions — a dataset carrying borough AND zip is one node with two edges, which is the fact the web is drawn to show. Each carries its own coverage so a renderer can colour a node by how analysed it is.","items":{"type":"object","additionalProperties":false,"required":["four_four","name","agency_id","domain_key","mart_declared","band","columns_total","decided","bound","unanalysed","dimensions"],"properties":{"four_four":{"type":"string"},"name":{"type":"string"},"agency_id":{"type":["string","null"]},"domain_key":{"type":["string","null"]},"mart_declared":{"type":"boolean"},"band":{"type":"string","enum":["no-inventory","untouched","started","half","most","complete"]},"columns_total":{"type":"integer","minimum":0},"decided":{"type":"integer","minimum":0},"bound":{"type":"integer","minimum":0},"unanalysed":{"type":"integer","minimum":0},"dimensions":{"type":"integer","minimum":0,"description":"The node's TRUE degree across the whole web, which under a dimension filter or a truncated response exceeds the edges present here."}}}},"edges":{"type":"array","items":{"type":"object","additionalProperties":false,"required":["dimension_key","four_four","columns","investigated_columns","field_names"],"properties":{"dimension_key":{"type":"string"},"four_four":{"type":"string"},"columns":{"type":"integer","minimum":0},"investigated_columns":{"type":"integer","minimum":0,"description":"Zero means the edge rests entirely on carried lexical tags — a weaker claim than one somebody read."},"field_names":{"type":"array","description":"The columns that make the membership, so a reader who clicks an edge sees WHICH field it is.","items":{"type":"string"}}}}},"describes":{"type":"object","additionalProperties":false,"required":["shape","not_a_join_claim","investigated_columns","truncated"],"properties":{"shape":{"type":"string","minLength":1},"not_a_join_claim":{"type":"string","minLength":1},"investigated_columns":{"type":"string","minLength":1},"truncated":{"type":"string","minLength":1}}}}},"freshness":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/freshness.json","title":"GET /api/freshness","description":"The overdue list, against OTI's own 2x standard. `standard` is in the payload because an overdue claim about a public agency has to say whose rule it is being measured by.","type":"object","additionalProperties":false,"required":["as_of","standard","counts_by_state","overdue","total_overdue"],"properties":{"as_of":{"type":"string","format":"date-time"},"standard":{"type":"object","additionalProperties":false,"required":["rule","authority","standard_id"],"properties":{"rule":{"type":"string","minLength":1},"authority":{"type":"string","minLength":1},"standard_id":{"type":"string","minLength":1}}},"counts_by_state":{"type":"array","items":{"type":"object","additionalProperties":false,"required":["freshness_state","count"],"properties":{"freshness_state":{"type":"string"},"count":{"type":"integer","minimum":1}}}},"total_overdue":{"type":"integer","minimum":0},"overdue":{"type":"array","items":{"type":"object","additionalProperties":false,"required":["four_four","name","agency_id","update_frequency","effective_period_days","data_updated_at","days_since_update","overdue_factor","permalink"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"name":{"type":"string"},"agency_id":{"type":["string","null"]},"update_frequency":{"type":["string","null"]},"effective_period_days":{"type":["number","null"]},"data_updated_at":{"type":["string","null"],"format":"date-time"},"days_since_update":{"type":["number","null"]},"overdue_factor":{"type":["number","null"]},"permalink":{"type":"string"}}}}}},"graph-expand-request":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/graph-expand-request.json","title":"POST /api/graph/expand — request body","description":"One node's expansion through the instance's own explore-graph configuration. The whole request body is capped at 16384 bytes; past it the route answers 400 {\"error\":\"request body too large\"}. A missing or non-string node answers 400 {\"error\":\"name the node IRI to expand\"}. Extra keys in the body are ignored, never refused. Answers carry truncated and true_degree; when truncated is true, the link set is a clipped, deterministic subset and true_degree is the node's real edge count.","type":"object","required":["node"],"properties":{"node":{"type":"string","description":"A node IRI in this graph, or a Socrata four-by-four, which resolves to its dataset node."},"limit":{"type":"integer","minimum":1,"maximum":200,"default":50,"description":"How many edges to return. A non-number reads as the default of 50; the registry bounds the argument to 1–200."}}},"graph-query-request":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/graph-query-request.json","title":"POST /api/graph/query — request body","description":"A named query from the published registry, with typed argument values. There is no path from a request to command text: the id selects a committed command and the args travel as bound values. The whole request body is capped at 16384 bytes; past it the route answers 400 {\"error\":\"request body too large\"}. A missing or non-string id answers 400 {\"error\":\"name a query id from the published registry\"}. An id outside the registry answers 400 in the shape {\"error\":\"no named query \\\"not-a-query\\\" — the public class reaches the registry and nothing else\"}. Extra keys in the body are ignored, never refused. Answers carry truncated, true_count and row_limit; when truncated is true, every count in that answer is a floor computed over the clipped row set, never a total, true_count is null because the registry publishes no count form of a named query, and row_limit names the cap that clipped it.","type":"object","required":["id"],"properties":{"id":{"type":"string","description":"A registered query id. The composed manifest carries the enum, derived from the registry's own executable entries — the registry export is the single source of this set.","enum":["agency-datasets","bridges-dimension","bridges-field","expand-node","neighbors","node-basics","path-search","search-concepts","shared-fields","theme-datasets"]},"args":{"type":"object","description":"The query's typed arguments by name, as declared by the registry entry: IRIs and four-by-fours as strings, bounded integers for limits and depths. A value outside its declared type or range is refused with 400."}}},"graph-query-response":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/graph-query-response.json","title":"POST /api/graph/query — response body","description":"A named registry query's answer: the query's own payload fields spread first, then the five contract fields every graph answer carries (state, query_id, snapshot_pin, produced_at, queue) — a payload key of the same name can never replace one of the five. TRUNCATION RULE: when truncated is true, every count in that answer is a FLOOR computed over the clipped row set, never a total. true_count and row_limit say how far a floor is from the total: true_count is the size of the WHOLE result set and row_limit names the cap whether or not it bit. Row shapes vary per query id (the request schema's id enum is derived from the committed registry); executable queries answer rows/count/elapsedMs/truncated/true_count/row_limit.","type":"object","required":["state","query_id","snapshot_pin","produced_at","queue"],"properties":{"rows":{"type":"array","description":"The query's rows, capped by the registry entry's own row cap. Row keys are the named query's, fixed per id.","items":{"type":"object"}},"count":{"type":"integer","minimum":0,"description":"How many rows this answer carries. With truncated: true this is a floor over the clipped set, never the true total."},"true_count":{"type":["integer","null"],"minimum":0,"description":"The size of the WHOLE result set, on the discipline /api/graph/expand's true_degree already follows. A number on an unclipped answer, because the rows carried ARE all of them. NULL on a clipped answer (truncated: true), because the registry publishes no count form of a named query — so null means \"nothing measured this\", never zero. A measured denominator for a clipped answer needs a count sibling per registry entry, which is registry work rather than a wire change; until then null is the truthful value."},"row_limit":{"type":["integer","null"],"minimum":0,"description":"The registry entry's row cap for this query, reported whether or not it bit. It exists so a caller can tell rows.length landing on a round number apart from rows.length landing on the ceiling. null when the entry declares no cap."},"elapsedMs":{"type":"number","description":"Store-side execution time for this answer, in milliseconds."},"truncated":{"type":"boolean","description":"true means the row set was clipped at the query's cap — every count in the answer is then a floor, and any ranking computed from it ran over the clipped rows."},"state":{"type":"string","description":"How this answer was produced — \"live\" when computed against the store now; other values name a cached or queue-degraded posture."},"query_id":{"type":"string","description":"The registry id that ran. Always one of the committed registry's ids, never caller text."},"snapshot_pin":{"type":"string","description":"WHICH graph generation answered (e.g. run-121.e21ad5fbe37b3096) — what makes two answers comparable. It is a generation name, never an address into the store."},"produced_at":{"type":"string","format":"date-time","description":"When this answer was produced."},"queue":{"type":"object","additionalProperties":false,"required":["depth_at_entry","waited_ms","deep"],"description":"The cost plane's admission record for this answer.","properties":{"depth_at_entry":{"type":"integer","minimum":0,"description":"How many requests were queued ahead of this one."},"waited_ms":{"type":"number","description":"How long this request waited for admission, in milliseconds."},"deep":{"type":"boolean","description":"true means the queue was deep at entry and the answer may have been served from a degraded path."}}}}},"ontology-response":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/ontology-response.json","title":"GET /api/ontology — response body","description":"The shared ontology bundle for a scope: the concept scheme, one binding row per column in scope (bound or not), and the coverage quadruple. TWO ID SPACES, AND THE RULE THAT BRIDGES THEM: concept keys (e.g. place.school_district) name MEANING — what a column is about; conformed-dimension FAMILY keys (e.g. school_district) name the ANALYTICAL JOIN KEYS — what /api/relate can join on. A concept maps to AT MOST ONE family, via concepts[].dimension_key (null when the concept is not joinable); a family key is never itself a concept key. Nothing in this bundle authorises a join by itself — catalog.dataset_columns.dimension_key does, and its per-column tag tier (the closed vocabulary declared | inferred | derived | investigated) is surfaced by the dimension routes (assertions on GET /api/dimensions/{key}/datasets, a/b assertion on pair bindings), while THIS bundle's bindings carry the CONCEPT-assertion axes (outcome/method/investigated/confidence). Scoping: name up to 25 datasets (?four_four=, repeatable or comma-separated), a ?concept= or a ?dimension=; more than 25 is refused rather than truncated, and a ?concept= naming no concept in catalog.semantic_concepts is refused with 404 rather than answered with an empty bundle indistinguishable from a real but unbound concept. UNSCOPED requests get the vocabulary and the catalog-wide coverage with bindings: [] — every column is never dumped. Bindings are capped at 5000 rows; past the cap scope.truncated is true, the bundle carries the first 5000 in catalog order, and the coverage counts describe ONLY the rows present — every count is then a floor for the selection, never its total.","type":"object","additionalProperties":false,"required":["title","concept_scheme","as_of","scope","coverage","catalog_coverage","concepts","bindings","not_measured"],"properties":{"title":{"type":"string"},"concept_scheme":{"type":"string","description":"The concept namespace this bundle serialises. citytap runs no triple store and no SPARQL endpoint — this document SERIALISES a graph, it does not host one."},"as_of":{"type":"string","format":"date-time"},"scope":{"type":"object","additionalProperties":false,"required":["four_fours","concept","dimension","scoped","truncated"],"properties":{"four_fours":{"type":"array","items":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"description":"The datasets named, deduplicated. At most 25 per bundle."},"concept":{"type":["string","null"],"description":"The concept key named, from the meaning id space (e.g. place.school_district)."},"dimension":{"type":["string","null"],"description":"The family key named, from the join-key id space (e.g. school_district)."},"scoped":{"type":"boolean","description":"false means no scope was named: the bundle carries the vocabulary and bindings is empty."},"truncated":{"type":"boolean","description":"true means the selection exceeded the 5000-binding cap and every coverage count is a floor over the rows present."}}},"coverage":{"type":"object","additionalProperties":false,"required":["describes","bound","refused","insufficient","unanalysed","not_computable"],"description":"The four-way split over the population `describes` names — four different facts, never a percentage. WHICH POPULATION DEPENDS ON THE SCOPE, and describes always says which: with a ?four_four= or ?dimension= scope it is the columns in the selection and the four counts sum to bindings.length; UNSCOPED it is the whole catalog (bindings is empty and the counts do NOT sum to it, because the scope is the catalog rather than the rows carried); with a ?concept= scope it is the columns whose live concept assertion names that key. Counts are floors when scope.truncated is true. A bucket that cannot be computed at the requested scope is null, never 0 — see not_computable.","properties":{"describes":{"type":"string","description":"The population these four counts cover, in one sentence. Read it before comparing them with catalog_coverage, which counts something else."},"bound":{"type":["integer","null"],"minimum":0},"refused":{"type":["integer","null"],"minimum":0},"insufficient":{"type":["integer","null"],"minimum":0,"description":"null under a ?concept= scope: an insufficient_evidence assertion carries no concept_key (the ledger's CHECK forbids one), so no column is insufficient FOR a concept."},"unanalysed":{"type":["integer","null"],"minimum":0,"description":"The one bucket that means \"nobody looked\" — a different fact from having been investigated and found to carry no concept. null under a ?concept= scope: a column with no assertion names no concept."},"not_computable":{"type":["object","null"],"additionalProperties":{"type":"string"},"description":"Present as an object exactly when some bucket above is null: one entry per null bucket, keyed by its name, whose value says why the number cannot be produced at this scope. null when all four counts are numbers. This is what separates \"we looked and found none\" from \"this scope cannot express that number\"."}}},"catalog_coverage":{"type":["object","null"],"additionalProperties":false,"required":["describes","columns_total","bound","refused","insufficient","unanalysed"],"description":"The same quadruple over the WHOLE catalog, plus its column total — present whatever the scope, so an agent can size the investigated surface without dumping it. NOT the selection: compare it with coverage only after reading both describes fields. null where semantic.investigation_coverage answers no row.","properties":{"describes":{"type":"string","description":"The population these counts cover, in one sentence — always the whole catalog, stated so it cannot be mistaken for the scoped coverage beside it."},"columns_total":{"type":"integer","minimum":0,"description":"Every column in the catalog. The live denominator for the four counts here; do not carry a copy of it in your own prose."},"bound":{"type":"integer","minimum":0},"refused":{"type":"integer","minimum":0},"insufficient":{"type":"integer","minimum":0},"unanalysed":{"type":"integer","minimum":0}}},"concepts":{"type":"array","description":"The whole concept vocabulary, in every bundle — the meaning id space.","items":{"type":"object","additionalProperties":false,"required":["key","kind","title","definition","dimension_key","broader_key","is_world_time","counts_concept","population","observed_outcome","notes"],"properties":{"key":{"type":"string","description":"The concept id (e.g. place.school_district, actor.agency) — meaning, never a join key."},"kind":{"type":"string"},"title":{"type":"string"},"definition":{"type":["string","null"]},"dimension_key":{"type":["string","null"],"description":"THE BRIDGE between the id spaces: the one conformed-dimension family this concept joins on (e.g. school_district), or null when the concept names meaning with no conformed join key. A concept maps to at most one family."},"broader_key":{"type":["string","null"],"description":"The parent concept, for walking the scheme upward."},"is_world_time":{"type":["boolean","null"]},"counts_concept":{"type":["string","null"]},"population":{"type":["string","null"]},"observed_outcome":{"type":["string","null"]},"notes":{"type":["string","null"]}}}},"bindings":{"type":"array","description":"One row per column in scope, BOUND OR NOT — an investigated column that carries no concept ships with a null concept and a stated reason instead of being dropped. At most 5000 rows per bundle (scope.truncated says when the cap clipped). Each row names the per-dataset column (four_four + field_name) and the concept-assertion axes.","items":{"type":"object","additionalProperties":false,"required":["four_four","dataset_name","field_name","display_name","description","source_datatype","pg_type","concept_key","outcome","method","investigated","confidence","evidence_quote","unbound_reason","role","role_outcome","role_method","role_confidence","unit_key","null_sentinels","review_state"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"dataset_name":{"type":["string","null"]},"field_name":{"type":"string","description":"The dataset's own column, as the catalog names it — the per-dataset half of a binding."},"display_name":{"type":["string","null"]},"description":{"type":["string","null"]},"source_datatype":{"type":["string","null"]},"pg_type":{"type":["string","null"]},"concept_key":{"type":["string","null"],"description":"The bound concept (meaning id space), only when outcome is bound — a refusal never carries its candidate concept forward."},"outcome":{"type":"string","enum":["bound","refused","insufficient","unanalysed"],"description":"The concept-assertion outcome for this column."},"method":{"type":["string","null"],"description":"How the assertion was made (e.g. operator_cohort)."},"investigated":{"type":["boolean","null"],"description":"false means the binding was derived from the column's name or type, not from reading its own evidence — filter on this to set your own trust threshold."},"confidence":{"type":["number","null"]},"evidence_quote":{"type":["string","null"]},"unbound_reason":{"type":["string","null"],"description":"Why a column carries no concept, in an actionable sentence; null on a bound row."},"role":{"type":["string","null"],"description":"The role axis rides along: identifier says the values are a surrogate key that joins to nothing outside its own dataset, even when the concept matches."},"role_outcome":{"type":["string","null"]},"role_method":{"type":["string","null"]},"role_confidence":{"type":["number","null"]},"unit_key":{"type":["string","null"]},"null_sentinels":{"description":"Values this column uses to mean null, or null.","type":["array","object","string","null"]},"review_state":{"type":["string","null"]}}}},"joinFamily":{"type":"object","additionalProperties":false,"required":["family","coverage","spine","crosswalk","joinOutcome"],"description":"THE AGENCY JOIN FAMILY, present exactly when the scope names it (?dimension=agency, ?concept=actor.agency, or a concept whose dimension_key bridges agency). The spine is the governance-corroborated catalog.agencies mirror; the crosswalk is the hand-reviewed variant→spine-id record with the METHOD on every mapped row (filter on it to set your own bar) and the refusal reason on every looked-at-unresolved one; coverage is THREE SEPARATE COUNTS per column and for the family — never a ratio, and the lowercase sentinel \"not_computable\" appears where the coverage view cannot answer (an environment still behind the substrate migrations), never 0. Before the disposal migration lands, counts read zero, crosswalk and joinOutcome read empty, and the spine already answers — assert shape, not values.","properties":{"family":{"const":"agency"},"coverage":{"type":"object","additionalProperties":false,"required":["family","perColumn"],"properties":{"family":{"type":"object","additionalProperties":false,"required":["resolved","looked_at_unresolved","not_yet_looked_at","profiled_max_at"],"description":"The whole family's counts in match-key units — the grand-total row of semantic.v_agency_coverage, always present once the view exists (honest zeros mean nothing measured yet).","properties":{"resolved":{"anyOf":[{"type":"integer","minimum":0},{"const":"not_computable"}]},"looked_at_unresolved":{"anyOf":[{"type":"integer","minimum":0},{"const":"not_computable"}]},"not_yet_looked_at":{"anyOf":[{"type":"integer","minimum":0},{"const":"not_computable"}]},"profiled_max_at":{"type":["string","null"],"description":"ISO-8601 of the newest profile row anywhere in the family; null when nothing is profiled."}}},"perColumn":{"type":"array","description":"One row per column the coverage view measures — empty until the profile sweep or a review round lands.","items":{"type":"object","additionalProperties":false,"required":["four_four","field","resolved","looked_at_unresolved","not_yet_looked_at"],"properties":{"four_four":{"type":"string"},"field":{"type":"string"},"resolved":{"type":"integer","minimum":0},"looked_at_unresolved":{"type":"integer","minimum":0},"not_yet_looked_at":{"type":"integer","minimum":0}}}}}},"spine":{"type":"array","description":"catalog.agencies whole, ordered by agency_id: the join key (agency_id), the governance-authority corroboration (goid), the OMB 3-digit alternate key (omb_code — corroboration, never the join key), and the validity dates a citable instrument backs.","items":{"type":"object","additionalProperties":false,"required":["agency_id","name","acronym","omb_code","goid","valid_from","valid_to"],"properties":{"agency_id":{"type":"string"},"name":{"type":"string"},"acronym":{"type":["string","null"]},"omb_code":{"type":["string","null"]},"goid":{"type":["string","null"]},"valid_from":{"type":["string","null"],"description":"YYYY-MM-DD, or null when no citable instrument dates it."},"valid_to":{"type":["string","null"],"description":"YYYY-MM-DD — the succession boundary where one exists (doitt: 2022-01-19), else null."}}}},"crosswalk":{"type":"array","description":"The reviewed per-column value crosswalk, whole (its size is bounded by human review, and a clipped crosswalk is a wrong crosswalk). Each row either maps a verbatim value to a spine agency WITH its method, or refuses it with a stated reason — the refusal rows ARE the looked-at-unresolved record.","items":{"type":"object","additionalProperties":false,"required":["four_four","field","raw_value","spine_agency_id","method","refusal_reason"],"properties":{"four_four":{"type":"string"},"field":{"type":"string"},"raw_value":{"type":"string","description":"The published value VERBATIM — case, padding and all (provenance); matching happens on lower(trim(raw_value)), the one normalization boundary."},"spine_agency_id":{"type":["string","null"]},"method":{"type":["string","null"],"enum":["exact_code","exact_name","alias_table","manual_review","exact_goid",null],"description":"How the mapping was established, on every mapped row; null exactly on refusal rows. exact_goid is the strongest: the published value IS the spine agency's goid, the governance authority's own record id, matched character for character, so no judgement about which entity a rendering denotes was made. Every other method resolves a human-readable rendering and can be wrong about the entity it names."},"refusal_reason":{"type":["string","null"],"description":"Why a looked-at value did not resolve; present exactly when spine_agency_id is null."}}}},"joinOutcome":{"type":"array","description":"The per-column join disposition, and every row of it traces to a verdict a reviewed migration STATED — never to the shape of the data. Two records dispose of a column and the population is their union: the scalar agency tag on catalog.dataset_columns (joinable — the tag is what authorises the compiled join, and a tagged column stays joinable even where a few of its cells pack several agencies, because the join reads the reviewed crosswalk on the match key and each packed cell's own crosswalk row is a refusal that the spine-id predicate drops), and a live refused actor.agency assertion that superseded a bound one (not_joinable_multivalued where the column's cells pack several agencies and resolve through catalog.agency_value_bridges, excluded otherwise). A non-joinable row always carries the operator's own stated reason. A column whose disposition is not yet decided does not appear — pending is not an outcome.","items":{"type":"object","additionalProperties":false,"required":["four_four","field","outcome","reason"],"properties":{"four_four":{"type":"string"},"field":{"type":"string"},"outcome":{"type":"string","enum":["joinable","not_joinable_multivalued","excluded"]},"reason":{"type":["string","null"],"description":"null on joinable; the stated reason otherwise."}}}}}},"not_measured":{"type":"array","items":{"type":"string"},"description":"What this bundle deliberately does not claim, as readable sentences — including why an unscoped bundle carries no bindings and why the cap clipped a large one."}}},"partners":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/partners.json","title":"GET /api/datasets/{four_four}/partners","description":"Related datasets in three tiers. `partners` is the confirmed shared-meaning ranking. `catalog_partners` carries the two facts the catalog holds for every dataset — who published it and what subject it covers — so a dataset with no tagged column is still placed among its neighbours. Neither catalog tier is join evidence.","type":"object","additionalProperties":false,"required":["four_four","total","limit","offset","partners","catalog_partners"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"total":{"type":"integer","minimum":0},"limit":{"type":"integer","minimum":1},"offset":{"type":"integer","minimum":0},"partners":{"type":"array","items":{"$ref":"#/$defs/semanticPartner"}},"catalog_partners":{"type":"object","additionalProperties":false,"required":["agency","theme"],"properties":{"agency":{"type":["object","null"],"additionalProperties":false,"required":["agency_id","agency_name","total","datasets"],"properties":{"agency_id":{"type":"string"},"agency_name":{"type":["string","null"]},"total":{"type":"integer","minimum":0},"datasets":{"type":"array","items":{"$ref":"#/$defs/catalogPartner"}}}},"theme":{"type":["object","null"],"additionalProperties":false,"required":["domain_key","domain_title","total","same_agency_excluded","datasets"],"properties":{"domain_key":{"type":"string","pattern":"^[a-z0-9-]+$"},"domain_title":{"type":["string","null"]},"total":{"type":"integer","minimum":0},"same_agency_excluded":{"type":"integer","minimum":0},"datasets":{"type":"array","items":{"$ref":"#/$defs/catalogPartner"}}}}}}},"$defs":{"semanticPartner":{"type":"object","additionalProperties":false,"required":["four_four","name","domain_key","score","shared_dimension_keys","same_theme"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"name":{"type":"string"},"domain_key":{"type":["string","null"]},"score":{"type":"number"},"shared_dimension_keys":{"type":"array","items":{"type":"string"}},"same_theme":{"type":"boolean"}}},"catalogPartner":{"type":"object","additionalProperties":false,"required":["four_four","name","asset_type","agency_id","agency_name","domain_key","page_views_total"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"name":{"type":"string"},"asset_type":{"type":["string","null"]},"agency_id":{"type":["string","null"]},"agency_name":{"type":["string","null"]},"domain_key":{"type":["string","null"]},"page_views_total":{"type":["integer","null"],"minimum":0}}}}},"query-request":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/query-request.json","title":"POST /api/query — request body","description":"Read-only SQL, relayed verbatim to the sandboxed query-executor. Exactly one read-only SELECT statement is permitted, over the pg_ro catalog of mirrored marts, using allowlisted functions only. The whole request body is capped at 25000 bytes; past it the route answers 400 {\"error\":\"request body exceeds the 25000-byte cap\"}. A missing, empty or whitespace-only sql answers 400 {\"error\":\"sql is required\"}. A statement that is not a single read-only SELECT answers 400 whose error begins \"only a single read-only SELECT statement is permitted\" and carries the parser's own reason in parentheses. A function off the allowlist answers 400 in the shape {\"error\":\"function \\\"btrim\\\" is not on the allowlist; add it deliberately with a reason and a test\"}. Extra keys in the body are ignored, never refused.","type":"object","required":["sql"],"properties":{"sql":{"type":"string","minLength":1,"description":"A single read-only SELECT statement. Relations are addressed as pg_ro.<schema>.<table>. Hard caps apply executor-side: a plan-estimate gate, a returned-row cap and a per-statement wall-clock budget — see the operation's x-citytap-caps in the composed manifest."}}},"query-response":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/query-response.json","title":"POST /api/query — response body","description":"The sandboxed executor's answer, relayed verbatim: the rows the statement produced, the routing record, and the planner's own row estimate. Caps (src/query/executor-config.ts defaults): a plan estimated past 5000000 rows is refused before execution, at most 50000 rows are returned, and the statement's wall-clock budget is 30000 ms — each shrinks (to a tenth, a tenth and a third) while the executor runs under a degraded read-host binding. estimated_rows is the planner's estimate, never a count of what came back.","type":"object","additionalProperties":false,"required":["rows","routing","estimated_rows"],"properties":{"rows":{"type":"array","description":"One object per result row, keyed by the statement's own output column names. At most 50000 rows (5000 degraded).","items":{"type":"object"}},"routing":{"type":"object","additionalProperties":false,"required":["decisions","read_host_degraded"],"properties":{"decisions":{"type":"array","description":"The executor's per-relation routing record: which source (pg, lake) served each relation the statement read. Empty for a statement that touched no routed relation.","items":{"type":"object","properties":{"source":{"type":"string","description":"Which plane served the relation: pg (the read-only Postgres attach) or lake (a materialized curated parquet, which additionally carries its stamp)."}}}},"read_host_degraded":{"type":"boolean","description":"true means the executor is bound below its preferred read replica and this answer ran under the reduced caps. This is an ENVIRONMENT POSTURE, not an error — normal and permanent in single-instance environments (staging answers true on every query). Do not branch on it as a failure signal."}}},"estimated_rows":{"type":"number","description":"The planner's cost-gate estimate for the statement (the number checked against the 5000000-row gate), not the returned-row count."}}},"relate-request":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/relate-request.json","title":"POST /api/relate — request body","description":"A relate spec: which two datasets, which conformed-dimension family, which binding columns and which measures. The catalog tag is the authorization to join; every identifier in the compiled statement comes from the catalog, never from this body. The whole request body is capped at 8192 bytes; past it the route answers 400 {\"error\":\"request body exceeds the 8192-byte cap\"}. A body that does not parse to this shape answers 400 {\"error\":\"malformed relate spec — expected {\\\"dim\\\",\\\"a\\\":{\\\"fourFour\\\",\\\"bindingColumn\\\",\\\"measure\\\"},\\\"b\\\":{…}}\"}. A well-formed spec can still be refused with 400 {error, refusal}, where refusal names the failing check (for example binding-family-not-yet-supported, mart-absent, unsupported-measure); an unknown dim's error names the currently compiled family set. Extra keys in the body are ignored, never refused.","type":"object","required":["dim","a","b"],"properties":{"dim":{"type":"string","description":"A conformed-dimension family the relate compiler supports. The refusal for an unknown family lists the live compiled set — as probed: month, borough, zip, community_district, council_district, precinct, school_district, nta, agency. The agency family joins on entity identity through the reviewed catalog.agency_crosswalk (a joined table, never generated SQL): values the crosswalk does not resolve are excluded from the join and counted on the response's crosswalk_coverage — there is no raw-equality fallback."},"succession":{"type":"string","description":"Agency family only (ignored by every other family, which has no succession edges). mapped — the default when absent — carries each side's resolved spine id to its terminal successor through the closure over catalog.agency_succession ONLY (recorded rename/absorption edges, e.g. DoITT to OTI; containment such as HRA/DHS under DSS never merges in any mode). strict joins the ids exactly as the crosswalk resolved them; a strict zero overlap that the mapped closure would bridge answers refusal succession-strict-no-edge-crossing in the 200 envelope. Any other value is refused 400. The composed manifest carries the enum, derived from the compiler's own SUCCESSION_MODES export.","enum":["mapped","strict"]},"a":{"$ref":"#/$defs/side"},"b":{"$ref":"#/$defs/side"}},"$defs":{"side":{"type":"object","required":["fourFour","bindingColumn","measure"],"properties":{"fourFour":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$","description":"The dataset's Socrata four-by-four."},"bindingColumn":{"type":"string","description":"The catalog field_name the caller claims carries the family tag on this dataset. Validated against the catalog and the mirrored mart, never trusted."},"measure":{"$ref":"#/$defs/measure"}}},"measure":{"type":"object","required":["family"],"properties":{"family":{"type":"string","description":"One of the closed measure-family vocabulary. count needs no column; every other family needs one; share_of_rows additionally needs values. The composed manifest carries the enum, derived from the compiler's own MEASURE_FAMILIES export.","enum":["avg","count","max","median","min","share_of_rows","sum"]},"column":{"type":"string","description":"The catalog field_name the aggregate reads. Required for every family except count."},"values":{"type":"array","items":{"type":"string"},"description":"share_of_rows only: the value set, validated against what the mart actually holds."}}}}},"relate-response":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/relate-response.json","title":"POST /api/relate — response body","description":"One catalog-confirmed join, answered as aggregate rows plus a coverage claim — or a NAMED REFUSAL in the same 200 envelope. WHAT relatable MEANS: relatable: true on GET /api/datasets/{fourFour}/bindings/{partner} means the pair compiles AND its family's measurement upheld both binding columns — the nightly value-conformance probe for every family but agency (share ≥ 0.99 over a ≤200k-row sample; each side's probe carries the measurement and its checked_at), and measured crosswalk coverage (semantic.v_agency_coverage resolved > 0 on both sides) for the agency family, whose columns carry no nightly probes yet; the listing's provenance field says which instrument decided. A pair with no measurement yet is relatable: false with refusal: \"probe-unavailable\", and a pair whose column failed the probe is refusal: \"value-conformance-refuted\" with the observed counterexamples. Execution additionally needs BOTH mirrored marts to physically hold the tagged columns; a missing one is refused 400 {error, refusal: \"mart-absent\"} and the error names the missing mart (\"<fourFour>.<column> is in the catalog's column manifest but not in the mirrored mart at <schema>.<table>\"). The full 400 refusal vocabulary a client can branch on without parsing prose: no-confirmed-binding, mart-absent, binding-family-not-yet-supported, census-tract-vintage-unresolved, unsupported-measure, share-of-rows-value-not-observed, value-conformance-refuted, succession-cycle. TWO RESULT MODES, AND result_mode SAYS WHICH ONE ANSWERED. result_mode: \"series\" — every conformed-dimension family (month, borough, zip, community_district, council_district, precinct, school_district, nta, agency): rows ARE the answer, one aggregate row per member. result_mode: \"overlap\" — the IDENTITY-KEY families (dim: \"bin\", \"bbl\"), which name one building out of 1.08M or one tax lot out of 858K rather than one of a handful of members: the answer is coverage's three exact counts over every matched key, and rows is a CAPPED, DETERMINISTIC SAMPLE of the matched keys (rows_cap states the ceiling, rows_sampled: true says the list was clipped). A client that charts an overlap answer's rows would draw at most rows_cap bars and imply they were all of them — branch on result_mode. An identity-key join excludes the borough placeholder sentinels (BIN 1000000-5000000; an all-zero BBL block or lot) from BOTH sides, because they are well-formed by construction and would otherwise match thousands of unrelated records. Only the count measure is admissible on an identity key; anything else is refused 400 unsupported-measure. Identity-key bindings are advertised by the SHAPE of the column's sampled values, not by the nightly probe (these families have no member vocabulary to probe against) — GET /api/datasets/{fourFour}/bindings/{partner} stamps those listings provenance: \"identity-key-shape\", and a column whose sampled values are not keys of the family's shape is refused value-conformance-refuted exactly as a refuted probe is. ZERO OVERLAP IS A REFUSAL FOR A DIMENSION FAMILY AND A FINDING FOR AN IDENTITY KEY: a dimension join whose coverage.overlap_count is 0 answers 200 with finding: null and refusal: \"no-shared-member-values\" regardless of probe state — two columns tagged with one family that share no member is a defect signal (a side's values do not conform, or nobody has probed them), and the statement names both sides' most frequent values and any unprobed side. An identity-key join that matches nothing is the OPPOSITE: two datasets naming different buildings is a true answer, so it carries finding and refusal: null; the identity-key refusal (\"value-conformance-refuted\") is reserved for a side that produced no well-formed key at all. THE AGENCY FAMILY (dim: \"agency\") joins on entity identity through the reviewed catalog.agency_crosswalk: members are catalog.agencies ids, values the crosswalk does not resolve are EXCLUDED from the join and COUNTED on crosswalk_coverage (never raw-matched, never silently dropped), succession applies at join time per spec.succession (mapped default; a strict zero overlap the mapped closure would bridge answers refusal \"succession-strict-no-edge-crossing\" in the 200 envelope), and provenance.{a,b}.resolution says whether a side was read directly or through the materialized-profile sidecar (sampled: true labels those counts estimates). On a finding, refusal is null and finding repeats the statement. Caps: the request body is capped at 8192 bytes, a plan estimated past 500000 rows is refused, and the statement budget is 30000 ms. Note the casing asymmetry: the request spec takes camelCase sides ({\"fourFour\",\"bindingColumn\"}); this response echoes snake_case (spec.a.four_four).","type":"object","additionalProperties":false,"required":["spec","result_mode","rows","coverage","statement","finding","refusal","probe","provenance"],"properties":{"spec":{"type":"object","additionalProperties":false,"required":["dim","a","b"],"description":"The validated spec, echoed so the caller can read the member key (spec.dim) and both dataset identities off the answer itself.","properties":{"dim":{"type":"string","description":"The conformed-dimension FAMILY KEY the join ran on (e.g. school_district) — the analytical join-key id space, distinct from concept keys like place.school_district, which name meaning (see the ontology bundle's concepts[].dimension_key for the mapping)."},"succession":{"type":"string","enum":["mapped","strict"],"description":"Agency family only: the succession mode the join ACTUALLY ran under — the default made explicit, so a caller who sent nothing can see which semantics answered. Absent for every other family."},"a":{"$ref":"#/$defs/specSide"},"b":{"$ref":"#/$defs/specSide"}}},"result_mode":{"type":"string","enum":["series","overlap"],"description":"WHICH KIND OF ANSWER THIS IS, and therefore what rows means. \"series\": rows ARE the answer — one aggregate row per conformed member (every dimension family). \"overlap\": coverage's three counts are the answer and rows is a capped sample of the matched keys (the identity-key families bin and bbl). Branch on this before rendering; see the top-level description."},"rows_cap":{"type":"integer","minimum":1,"description":"Overlap mode only (absent in series mode): the ceiling on how many matched keys rows can list. The counts on coverage are exact regardless."},"rows_sampled":{"type":"boolean","description":"Overlap mode only (absent in series mode): true when more keys matched than rows lists — the list is clipped, the counts are not."},"rows":{"type":"array","description":"In series mode, one aggregate row per conformed member value, from a FULL OUTER JOIN — a member held by only one side appears with the other side's value null. Each row carries the member under the FAMILY'S OWN KEY, spec.dim (a borough relate's rows carry \"borough\", an nta relate's carry \"nta\"): branch on spec.dim, never sniff the row. The member is the rendered label (\"Brooklyn\", not the stored code \"BK\" or 3), so it is NOT usable as an equality filter against the source column.","items":{"type":"object","required":["a_value","b_value"],"properties":{"a_value":{"type":["number","string","null"],"description":"Side a's measure for this member; null where the member appears only in b."},"b_value":{"type":["number","string","null"],"description":"Side b's measure for this member; null where the member appears only in a."}}}},"coverage":{"type":"object","additionalProperties":false,"required":["overlap_count","range_start","range_end","a_only_count","b_only_count"],"description":"The join+coverage facts, counted over the returned aggregate rows.","properties":{"overlap_count":{"type":"integer","minimum":0,"description":"Members both sides hold."},"range_start":{"type":["string","null"],"description":"YYYY-MM, for the temporal (month) family only — a range is a claim about order. Every categorical family answers null."},"range_end":{"type":["string","null"],"description":"YYYY-MM, month family only; null for categorical families."},"a_only_count":{"type":"integer","minimum":0,"description":"Members only side a holds."},"b_only_count":{"type":"integer","minimum":0,"description":"Members only side b holds."}}},"statement":{"type":"string","description":"The one sentence a reader-facing surface prints. On a finding: a deterministic join+coverage sentence naming both dataset titles and the family — never a statistical or causal claim. On a refusal: the refusal's own prose, naming both sides' most frequent values and any side the probe has not upheld."},"finding":{"type":["string","null"],"description":"The join+coverage claim, or null when the answer is a refusal. Branch on this (or on refusal), never on rows.length: a non-empty rows array can still be a refusal when every row is one-sided."},"crosswalk_coverage":{"type":"object","additionalProperties":false,"required":["a","b"],"description":"Agency family only (absent for every other family): each side's measured crosswalk coverage from semantic.v_agency_coverage — THREE SEPARATE COUNTS in match-key units, never a ratio, so the values the join excluded are counted on the answer itself rather than silently dropped. null per side when nothing is measured for that column yet (nothing profiled and no reviewed crosswalk row — including the family's empty-crosswalk birth state).","properties":{"a":{"$ref":"#/$defs/crosswalkCoverage"},"b":{"$ref":"#/$defs/crosswalkCoverage"}}},"refusal":{"type":["string","null"],"enum":["no-shared-member-values","succession-strict-no-edge-crossing","value-conformance-refuted",null],"description":"null on a finding. \"value-conformance-refuted\" (overlap mode only): one side produced no well-formed key of the family's shape over its whole mart, which refutes that column's own tag — distinct from a zero overlap, which in overlap mode is a finding. \"no-shared-member-values\": coverage.overlap_count is 0 — a defect signal, not a finding (see the top-level description). \"succession-strict-no-edge-crossing\" (agency family, succession=strict only): the strict join holds no shared member AND a second, measured run of the successor-mapped variant does — the miss is the withheld edge crossing, stated by name rather than served as a silent zero."},"probe":{"type":"object","additionalProperties":false,"required":["a","b"],"description":"Each side's latest value-conformance probe from the nightly falsifier sweep (semantic.assertion_checks, falsifier_kind value_domain), or null when the column has never been probed. A refuted probe never reaches this envelope — it is refused 400 value-conformance-refuted.","properties":{"a":{"$ref":"#/$defs/probe"},"b":{"$ref":"#/$defs/probe"}}},"provenance":{"type":"object","additionalProperties":false,"required":["a","b","routing"],"properties":{"a":{"$ref":"#/$defs/provenanceSide"},"b":{"$ref":"#/$defs/provenanceSide"},"routing":{"type":"array","description":"The executor's per-relation routing record for the compiled statement — which source served each side's mart, with the planner's per-relation row estimate.","items":{"type":"object"}}}}},"$defs":{"probe":{"type":["object","null"],"additionalProperties":false,"required":["checked_at","outcome","share","reason","counterexamples"],"properties":{"checked_at":{"type":"string","description":"ISO-8601 — when the probe ran."},"outcome":{"type":"string","enum":["upheld","inconclusive"],"description":"upheld: the column's values conform to the family's members at the declared coverage. inconclusive: the sweep could not decide (not a soft refutation). refuted never appears here — see the top-level description."},"share":{"type":["number","null"],"minimum":0,"maximum":1,"description":"The conforming share the sweep measured; null when the check recorded none."},"reason":{"type":["string","null"],"description":"The sweep's own reason on an inconclusive check; null on an upheld one."},"counterexamples":{"type":"array","items":{"type":"string"},"maxItems":5,"description":"Non-conforming values the sweep observed, as text; empty on an upheld check."}}},"specSide":{"type":"object","additionalProperties":false,"required":["four_four"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$","description":"The side's Socrata four-by-four (snake_case here; the request took camelCase fourFour)."}}},"provenanceSide":{"type":"object","additionalProperties":false,"required":["four_four","title"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"title":{"type":"string","description":"The dataset's catalog name — the same word the statement uses."},"resolution":{"type":"string","enum":["direct","materialized-profile"],"description":"Agency family only: HOW this side was resolved. direct — the mirrored mart was read through the crosswalk join. materialized-profile — the mart exceeds the executor's plan cap, so the join ran over semantic.agency_value_profiles ⋈ catalog.agency_crosswalk and the raw mart was never scanned. Absent for every other family."},"sampled":{"type":"boolean","description":"Agency family only: true exactly on the materialized-profile route — profiles for large marts are TABLESAMPLE-derived, so this side's counts are labelled estimates; presence in the profile still proves >0 true rows. Absent for every other family."}}},"crosswalkCoverage":{"type":["object","null"],"additionalProperties":false,"required":["resolved","looked_at_unresolved","not_yet_looked_at","last_profiled_at"],"description":"One column's coverage in catalog.agency_match_key units: resolved (a reviewed row names a spine agency), looked_at_unresolved (a reviewed row states a refusal reason — the value is counted, never dropped), not_yet_looked_at (a profiled value no reviewed row covers). null when the coverage view holds no row for the column.","properties":{"resolved":{"type":"integer","minimum":0},"looked_at_unresolved":{"type":"integer","minimum":0},"not_yet_looked_at":{"type":"integer","minimum":0},"last_profiled_at":{"type":["string","null"],"description":"ISO-8601 of the column's newest profile row; null when only crosswalk rows exist."}}}}},"scorecard":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/scorecard.json","title":"GET /api/scorecard","description":"The city's own published KPI definitions (updated-on-time = met declared frequency, periodic assets only; overdue = past 2x declared cadence), plus mirror_lag — the measurement the city's dashboard cannot make, because it is about OUR copy rather than theirs. `definitions` ships in the payload so a reader can check the arithmetic against the rule rather than trusting the number.","type":"object","additionalProperties":false,"required":["as_of","environment","datasets_total","rows_total","pct_updated_on_time","overdue_count","no_regular_updates","datasets_total_catalog","no_cadence_catalog","overdue_catalog","fenced_total","fenced_overdue","freshness_observed_at","declared_marts_absent","alert_deliveries_enqueued","alert_deliveries_abandoned","mirror_lag","definitions"],"properties":{"as_of":{"type":"string","format":"date-time"},"environment":{"type":"string"},"datasets_total":{"type":"integer","minimum":0},"rows_total":{"type":["integer","null"],"minimum":0},"pct_updated_on_time":{"type":["number","null"]},"overdue_count":{"type":"integer","minimum":0},"no_regular_updates":{"type":"integer","minimum":0},"datasets_total_catalog":{"type":"integer","minimum":0,"description":"Every tracked dataset in catalog.staleness_state. The denominator the five counts beside it are drawn over, published so a consumer states its numbers over the population they came from rather than over the mirrored subset mirror_lag reports."},"no_cadence_catalog":{"type":"integer","minimum":0,"description":"Tracked datasets declaring no update cadence, over datasets_total_catalog. no_regular_updates answers the same question over the city's own asset_type='dataset' filter; this one is drawn over the same population as its denominator."},"overdue_catalog":{"type":"integer","minimum":0,"description":"Tracked datasets in freshness_state OVERDUE, over datasets_total_catalog."},"fenced_total":{"type":"integer","minimum":0,"description":"Tracked datasets fenced out of refresh selection (catalog.staleness_state.is_zombie)."},"fenced_overdue":{"type":"integer","minimum":0,"description":"Fenced datasets whose axis-A state is OVERDUE, STALE or DUE — the subset carrying a state that says 'late' beside a fence that says 'nothing will act on it'. Waiting does not repair these."},"freshness_observed_at":{"type":["string","null"],"format":"date-time","description":"When the freshness machine last wrote the OLDEST row of catalog.staleness_state — min(record_updated_at). The observation clock for every freshness count in this payload, and deliberately not `as_of`: as_of is the request clock over a cron-written table and asserts a recency these counts do not have. The minimum rather than the maximum because these are counts over the whole table, so only the oldest row's stamp bounds every row in them. NOT catalog.staleness_state.last_probe_at, which is declared and never written (0 of 3,019 rows on prod, 2026-09-19)."},"declared_marts_absent":{"type":"integer","minimum":0,"description":"Datasets declaring a pg_schema.pg_table whose relation is not in the database (catalog.v_declared_marts_absent). These sit INSIDE the mart-declared set, so without this count the not-queryable total derived from datasets - datasets_mart_declared is a floor rather than a number."},"alert_deliveries_enqueued":{"type":"integer","minimum":0,"description":"Alert deliveries claimed in the last 30 days. The window is a design choice: a lifetime count cannot return to zero, so one historical failure would leave an abandoned-delivery signal reading dirty forever."},"alert_deliveries_abandoned":{"type":"integer","minimum":0,"description":"Of alert_deliveries_enqueued, those whose 8-attempt budget ran out or whose destination is undeliverable. Normally 0 — a measured-clean number, not an absent one."},"mirror_lag":{"type":"object","additionalProperties":false,"required":["mirrored_datasets","datasets_behind","p50_seconds","p95_seconds","max_seconds"],"properties":{"mirrored_datasets":{"type":"integer","minimum":0},"datasets_behind":{"type":"integer","minimum":0},"p50_seconds":{"type":["number","null"]},"p95_seconds":{"type":["number","null"]},"max_seconds":{"type":["number","null"]}}},"definitions":{"type":"object","additionalProperties":false,"required":["updated_on_time","overdue","mirror_lag","source"],"properties":{"updated_on_time":{"type":"string","minLength":1},"overdue":{"type":"string","minLength":1},"mirror_lag":{"type":"string","minLength":1},"source":{"type":"string","minLength":1}}}}},"standards-evidence":{"$schema":"https://json-schema.org/draft/2020-12/schema","$id":"https://citytap.codenow.nyc/schema/standards-evidence.json","title":"GET /api/standards/evidence","description":"The published bytes ONE conformance claim rests on. Takes ?standard=<standard_id>&field=<canonical_field> — the two columns every /api/conformance row carries — and optionally &four_four=<4x4> to ask about one dataset's record. The pointer comes from catalog.standards_map.emit_pointer, which the registry derives from the same crosswalk rows its emitters execute, so this route holds no knowledge of what any standard calls any field. present=false with records_with_key=0 is a real answer and the reason the route exists: it is what disproves a ledger row claiming a requirement is implemented.","type":"object","additionalProperties":false,"required":["standard_id","canonical_field","requirement_path","action","direction","notes","artifact","artifact_etag","pointer","subject","present","value","records_with_key","records_total"],"properties":{"standard_id":{"type":"string","minLength":1},"canonical_field":{"type":"string","minLength":1},"requirement_path":{"type":["string","null"],"description":"The field path inside ONE record, as the crosswalk states it (e.g. bureauCode, theme[0], dct:title)."},"action":{"enum":["pull","parse","compute","hardcode","absent","emit"]},"direction":{"enum":["harvest","emit","both"]},"notes":{"type":["string","null"]},"artifact":{"type":"string","minLength":1,"description":"The lake key of the generation quoted, date-partitioned per environment. A consumer citing this answer cites this key."},"artifact_etag":{"type":["string","null"]},"pointer":{"type":"string","pattern":"^(/[^/]+)+$","description":"RFC 6901, RESOLVED: the record index is substituted. With four_four it is that dataset's record; without one it is the first record carrying the field, and subject is null to say so. Where no record carries it, the stored template with its literal {i} is returned."},"subject":{"type":["object","null"],"additionalProperties":false,"required":["four_four","title"],"properties":{"four_four":{"type":"string","pattern":"^[a-z0-9]{4}-[a-z0-9]{4}$"},"title":{"type":["string","null"]}}},"present":{"type":"boolean","description":"Whether the pointer resolves. With four_four it is about that record; without one it is whether any record carries the field."},"value":{"description":"The value at the pointer, verbatim from the artifact, or null where it resolves to nothing. Any JSON type: the crosswalk maps arrays and objects as well as scalars."},"records_with_key":{"type":"integer","minimum":0,"description":"How many records of the artifact carry this field. Zero against a ledger row claiming 'implemented' is the falsification."},"records_total":{"type":"integer","minimum":0,"description":"Entries in the array the pointer indexes. For a document-level pointer, 1 — the document itself."}}}},"responses":{"rate-limited":{"description":"The caller's window is spent. Retry-After says how many seconds to wait; the operation's x-citytap-auth-class names which window.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}},"refused":{"description":"The request was refused: a malformed body, a value outside its declared bounds, or a claim a catalog or mart check disproved. The request schema's description quotes the refusal texts verbatim.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/error"}}}}}},"x-citytap-rate-classes":{"anonymous-rate-limited":"No credential. One shared fixed window per derived client address covers every route in this class; past it the answer is 429 with Retry-After in seconds.","anonymous-query-rate-limited":"No credential. POST /api/query draws on its own fixed window per derived client address, separate from the shared anonymous one, so sandboxed-SQL traffic and page reads never spend each other's budget; past it the answer is 429 with Retry-After in seconds.","anonymous-manifest-rate-limited":"No credential. GET /api/openapi.json draws on its own generous fixed window per derived client address, separate from the shared anonymous one, so manifest polling and page reads never spend each other's budget; past it the answer is 429 with Retry-After in seconds.","anonymous-suggest-rate-limited":"No credential. The type-ahead draws on its own fixed window per derived client address, so a keystroke stream pauses itself and nothing else.","anonymous-refresh-rate-limited":"No credential required. A bearer is OPTIONAL: supply one and the request uses the signed-in refresh lane, omit one and it is treated as anonymous. An INVALID bearer is refused with 401 rather than treated as anonymous. The route draws on its own small fixed window per derived client address; separately, an anonymous request is granted only while the environment's hourly portal allowance is above the share reserved for the catalog's own sweeps, and is answered 429 when it is not.","anonymous-session-labelled":"No credential. The same per-address window as anonymous-rate-limited, with a per-passcode fairness bucket beneath it. The passcode is caller-chosen and authorises nothing.","unauthenticated-probe":"Liveness and readiness probes. Never rate-limited."},"x-citytap-authenticated":[{"method":"GET","pattern":"/api/alerts","summary":"The caller's own watches, with each one's condition, cadence and last evaluation."},{"method":"POST","pattern":"/api/alerts","summary":"Turn one saved view into a watch, evaluated on its source datasets' declared cadence."}],"x-citytap-refusals":{"graph-credential-refused":{"status":503,"retryAfterSeconds":60},"graph-generation-missing":{"status":503,"retryAfterSeconds":15},"graph-store-fault":{"status":503,"retryAfterSeconds":15},"graph-store-timeout":{"status":503,"retryAfterSeconds":15},"graph-store-unreachable":{"status":503,"retryAfterSeconds":15},"graph-unconfigured":{"status":503,"retryAfterSeconds":3600},"graph-unpublished":{"status":503,"retryAfterSeconds":300},"lake-unreachable":{"status":503,"retryAfterSeconds":15},"refresh-gatekeeper-uninterpretable":{"status":500,"retryAfterSeconds":null},"refresh-gatekeeper-unreachable":{"status":503,"retryAfterSeconds":15}}}