A recovery design becomes harder to change once state begins moving across regions.
Shorter downtime and stronger correctness guardrails can demand different choices at the same cutover.
Who you’d be doing this for
“If we fail over, I need to know the numbers are right—not just that jobs are green.”
Esteban Teixeira · Director of Data Engineering
Owns data products that support regulated reporting and operational decision-making during outages.
What is at stake
Rare failback simulations show ordering conflicts across stateful pipeline metadata before a committed production cutover. You have to weigh a six-hour recovery objective against proof that recovery never silently loses or duplicates data.
Why it isn’t already fixed
Every obvious fix costs something else. That’s the part you’d have to decide.
- six-hour recovery vs proof of correctness
- availability commitments vs least-privilege controls
- automated findings vs verified evidence
- broad coverage vs safe initial cohort
Why Databricks
Lakeflow pipelines carry checkpoints, source offsets, stateful operator state, table versions, transaction metadata, schedules, and dataflow dependencies that must recover across region outages.
Written with these in mind
Not your kind of problem? 34 more at Databricks, or browse every organization.
This is the setup. The work is inside.
Running it puts you in the room: the full situation and its constraints, stakeholders who push back in their own words, and the decisions that are yours to make. What you produce becomes a Day One Plan — work you can show someone instead of describing.