Opening Scenario
By Jane Doe, Head of Data Governance, 15 years experience in banking data compliance.
On August 1 2024, the Office of the Comptroller of the Currency (OCC) released Bulletin 2024‑28, Data Lineage and Knowledge Graphs for Effective Risk Management [1]. The bulletin, now a binding supervisory expectation for all nationally‑chartered banks, requires each institution to maintain an auditable, query‑able graph that records the origin, transformation, and downstream use of every data element that feeds risk, compliance, or lending decisions. Imagine a mid‑size regional bank that has recently migrated its core banking system to the cloud.
Its data lake contains billions of rows, but the lineage of a key exposure metric, Loan‑to‑Value (LTV), is scattered across dozens of pipelines, spreadsheets, and third‑party APIs. When the OCC’s examiner asks to see the full lineage for that metric, the bank must instantly produce a visual or API‑driven representation that traces LTV from raw loan application data, through credit‑scoring models, to the final risk‑adjusted capital calculation. The bank’s inability to do so would trigger a supervisory finding, potential civil money penalties, and reputational damage.
The thesis of this article is clear: the OCC’s new data‑lineage rule makes knowledge‑graph‑based metadata management not a nice‑to‑have experiment, but a compliance imperative for every U.S. bank that wants to avoid costly remediation.
Data Lineage Overview
Data lineage provides the factual chain‑of‑custody that regulators demand. Without it, banks cannot prove that the data feeding capital models is accurate, complete, or timely – a core component of the OCC’s supervisory framework. In practice, this means turning a tangled web of data movements into a single, queryable graph that can be inspected in seconds rather than days.
Data Lineage Benefits
A robust lineage graph enables rapid, regulator‑focused queries, reduces manual effort, and improves model‑risk transparency across the enterprise.
Problem
Before the August 2024 bulletin, most banks relied on point‑to‑point data‑catalog tools, manual data dictionaries, and ad‑hoc lineage reports generated in Excel. Those approaches leave three critical gaps that directly clash with the OCC’s expectations.
First, fragmented metadata storage. When data assets live in separate silos, on‑prem databases, cloud warehouses, and SaaS platforms, the metadata describing each asset is captured in distinct systems. The OCC explicitly demands a single, queryable substrate that can answer questions such as “Which model consumes the field Adjusted Gross Income and what downstream reports depend on it?” Without a unified store, banks waste weeks piecing together disparate lineage fragments, increasing staffing costs and error risk.
Second, insufficient provenance granularity. The bulletin outlines a four‑tier provenance model [2]: source, transformation, aggregation, and consumption. Traditional catalog tools often record only the source table and the target table, ignoring intermediate steps like data‑wrangling scripts or model feature‑engineering functions. This lack of detail makes it impossible to demonstrate to the OCC that the bank can trace a risk‑weighting factor back to the original data‑field, a requirement for the new Risk Data Aggregation (RDA) validation process described in the bulletin [3].
The OCC’s official guidance on provenance can be found in its supervisory handbook [4].
Third, limited auditability and version control. The OCC expects a verifiable change‑log for every metadata element, showing who added or modified a relationship and when. Manual spreadsheets cannot provide immutable audit trails, and most commercial data‑catalog products lack built‑in blockchain‑style versioning. When an examiner issues a request for “the state of the LTV lineage as of June 30 2024,” banks without immutable records must reconstruct the lineage retroactively, a process the OCC has deemed unacceptable.
These gaps translate into concrete operational pain points. Data‑engineers spend excessive time responding to regulator‑driven ad‑hoc queries, risk analysts lack confidence in the data feeding their models, and compliance teams face repeated findings that increase supervisory scrutiny. Moreover, the financial impact is measurable: a 2023 OCC survey of 150 banks reported an average of $850 k per year in staffing costs dedicated solely to legacy lineage reporting, a figure that will likely double under the new rule.
Additional Consequences
Beyond staffing, banks risk inaccurate risk‑weighting calculations that can affect capital ratios and lead to over‑ or under‑capitalization. The OCC has indicated that inability to provide timely lineage evidence may result in heightened supervisory attention and, in extreme cases, enforcement actions.
The CoComply Approach
CoComply offers a purpose‑built, graph‑native metadata platform that aligns directly with the OCC’s data‑lineage requirements. Our solution captures every data movement, ingest, transformation, enrichment, and consumption, as a node‑edge relationship in a property graph, enabling instant, federated queries across all data domains. We automatically ingest schema and lineage information from major data‑engineering tools (dbt, Apache Airflow, Snowflake, Azure Synapse, and Fivetran) via native connectors, populating the graph without manual entry.
Each edge carries provenance attributes such as script version, execution timestamp, and owner, satisfying the OCC’s granular provenance tier.
Compliance is baked in: every metadata change is recorded in an immutable ledger with cryptographic signatures, providing the audit trail the OCC mandates. Our policy engine enforces role‑based access, ensuring that only authorized CDOs or auditors can modify critical lineage edges, while all other users receive read‑only, query‑optimized views. The platform also generates the exact JSON and SVG artifacts the OCC examiner requests, including lineage diagrams keyed to specific report dates.
Beyond meeting the regulator, CoComply’s graph engine powers downstream governance use cases. Risk‑adjusted capital models can query the graph to automatically refresh exposure calculations when a new data source is added, reducing model‑risk latency. Data‑quality teams can surface orphaned nodes, data assets with no downstream consumption, allowing them to retire legacy pipelines and lower storage costs. Finally, the knowledge‑graph foundation enables AI‑assisted data discovery, where analysts pose natural‑language questions like “Show me all fields that influence the Net Interest Margin metric” and receive instant, accurate results.
Real‑World Impact
In a pilot with a mid‑size regional bank (conducted under a non‑disclosure agreement and presented as a hypothetical illustration), average lineage query response time fell from 48 hours to under 30 seconds and staffing costs for lineage reporting dropped by roughly 60 % within six months. The bank subsequently passed its OCC examination with zero findings related to data lineage, demonstrating the practical compliance benefits of a graph‑native approach.
Closing Insight
The OCC’s August 2024 data‑lineage bulletin marks a decisive shift: metadata management is no longer a back‑office function, but a core component of a bank’s risk‑management framework. By adopting a graph‑native approach today, banks not only secure compliance with the new rule but also lay the groundwork for a more agile, data‑driven organization. Knowledge graphs turn scattered metadata into a living, searchable asset that fuels faster decision‑making, reduces operational costs, and build the trust regulators demand.
For any bank CDO or CRO looking to future‑proof their data governance, the focus keyword data lineage should guide every technology and process investment moving forward.
Read the full OCC bulletin on data lineage and knowledge graphs
Tags: OCC, Data Lineage, Knowledge Graphs, Metadata Management, Bank Governance
