Databricks
Data EngineerStrategicAug 6, 2026

Test reusable reference-data product priorities

Repeated copies of the same data can signal either healthy adaptation or avoidable fragmentation.

Standardization helps only when it preserves the context that downstream work actually needs.

I spend too much time proving whose supplier mapping is the one we should trust.

Paloma Cruz · Procurement Data Scientist

Builds supplier-risk models using reference data that is prepared differently by procurement, finance, and operations teams.

What pulls against what

  • semantic fit vs. standardization
  • sampled lineage vs. confident conclusions
  • compute savings vs. consumer adoption
  • domain autonomy vs. shared reuse

What is at stake

The right intervention could reduce duplicated compute and reconcile drifting reference logic; the wrong one could impose a dataset nobody adopts

Why Databricks

In shared analytics environments, reuse tends to depend as much on usable semantics as on technical availability.

Written for

Data-product-oriented engineerLineage-driven investigatorExperiment-minded platform engineer

This is the setup. The work is inside.

Running it puts you in the room: the full situation and its constraints, stakeholders who push back in their own words, and the decisions that are yours to make. What you produce becomes a Day One Plan — work you can show someone instead of describing.