Test reusable reference-data product priorities
Repeated copies of the same data can signal either healthy adaptation or avoidable fragmentation.
Standardization helps only when it preserves the context that downstream work actually needs.
“I spend too much time proving whose supplier mapping is the one we should trust.”
Paloma Cruz · Procurement Data Scientist
Builds supplier-risk models using reference data that is prepared differently by procurement, finance, and operations teams.
What pulls against what
- semantic fit vs. standardization
- sampled lineage vs. confident conclusions
- compute savings vs. consumer adoption
- domain autonomy vs. shared reuse
What is at stake
The right intervention could reduce duplicated compute and reconcile drifting reference logic; the wrong one could impose a dataset nobody adopts
Why Databricks
In shared analytics environments, reuse tends to depend as much on usable semantics as on technical availability.
Written for
This is the setup. The work is inside.
Running it puts you in the room: the full situation and its constraints, stakeholders who push back in their own words, and the decisions that are yours to make. What you produce becomes a Day One Plan — work you can show someone instead of describing.