Director of Engineering (Data Infrastructure)
Databricks
Skill Required
Key highlights
- 14+ years of distributed systems engineering experience required
- 6+ years leading infrastructure organizations and 4+ years managing managers
- BS in Computer Science or Engineering required
- MS or Ph.D. preferred
- Experience building 99.999%+ reliable systems
- Comprehensive benefits and perks offered
Role overview
Databricks processes petabytes of data and billions of transaction events daily - every cluster launch, every query executed, every dollar billed flows through infrastructure that must never fail. When we process billions in billing transactions with 99.999% accuracy requirements, when we ingest terabytes per second across 100+ regions, when a five‑minute outage costs millions in revenue and customer trust - infrastructure isn't just important, it's existential. In this leadership opportunity, you will build the data infrastructure organization that makes Databricks' continued growth possible. You'll establish foundational teams in Bengaluru owning the bedrock systems that guarantee billing correctness, operational resilience, and zero‑downtime recovery across our entire monetization stack, alongside multi‑region data ingestion, developer platforms, and deployment automation that eliminate friction at petabyte scale.
Responsibilities
- Deliver the infrastructure vision for systems processing billions in daily billing transactions with zero tolerance for error, building disaster recovery that's provably reliable, testing frameworks that catch what production sees, correctness systems that make billing errors structurally impossible, and observability that predicts failures before they happen
- Build Bengaluru's data infrastructure organization by establishing it as the destination for India's top infrastructure talent, hiring multiple engineering managers who become force multipliers, and creating a culture where solving hard distributed systems problems at scale is the daily work
- Own business-critical systems operating 24/7/365 across 100+ regions where even 99.9% uptime means hours of customer pain, driving reliability improvements that prevent millions in revenue loss while eliminating operational toil through frameworks that make systems self‑healing, self‑tuning, and self‑documenting
- Ship platforms that compound engineering leverage across Databricks: correctness frameworks that catch billing errors before customers do, deployment automation that makes regional expansion push‑button, data integration systems that process petabyte‑scale flows without human intervention, and testing infrastructure where comprehensive coverage is automatic, not heroic
- Position infrastructure as product by treating internal engineering teams as customers with SLAs, measuring adoption and satisfaction, iterating based on feedback, and demonstrating that every dollar invested in infrastructure returns multiplicative gains in product velocity, reliability improvements, or cost reductions
Requirements
- 14+ years in distributed systems engineering with 6+ years leading infrastructure organizations and 4+ years managing managers at companies where infrastructure failures meant immediate revenue impact, customer escalations, or regulatory consequences - and you built the systems and teams that made those failures rare
- Technical depth across petabyte‑scale data pipelines and distributed systems reliability where you can engage from "how should we architect multi‑region disaster recovery" to "why is this Kafka cluster exhibiting this latency pattern" while knowing when to coach versus when to decide
- Track record defining multi‑year infrastructure vision and translating it into sequential deliverables that show value quarterly while building toward architectural end states, positioning infrastructure investments as business enablers rather than cost centers, and making build‑vs‑buy decisions that compound over time
- Experience building 99.999%+ reliable systems with established practices for SLOs/SLIs, chaos engineering, disaster recovery, and sophisticated observability that predicts failures before they happen
- Proven ability to scale infrastructure organizations in high‑growth environments where you've doubled engineering while maintaining quality bar, developed engineering managers, and created teams where retention is high because the problems are interesting and the culture is strong
- Communication skills to make complex infrastructure decisions legible to executives (translating technical investments into business outcomes), influence cross‑functional partners without authority, build trust across global teams in different timezones with different working styles, and represent Databricks' technical brand externally
- BS in Computer Science or Engineering
Nice to have
- MS or Ph.D. preferred
- Experience with Apache Spark, Delta Lake, large‑scale data infrastructure, fintech/billing systems, or leading infrastructure through hypergrowth strongly preferred
Benefits
- Comprehensive benefits and perks that meet the needs of all employees (specific details available per region)
Additional details
- (P-1490) job identifier
- About Databricks: Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog.
- Our Commitment to Diversity and Inclusion: Databricks is committed to fostering a diverse and inclusive culture where everyone can excel. Hiring practices are inclusive and meet equal employment opportunity standards, considering individuals without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio‑economic status, veteran status, and other protected characteristics.
- Compliance: If access to export‑controlled technology or source code is required for performance of job duties, it is at the employer's discretion whether to apply for a U.S. government license, and the employer may decline to proceed with an applicant on this basis alone.
- Engineering culture is born from Apache Spark and open source, where technical depth matters and infrastructure engineers are celebrated as craftspeople.
- For specific details on the benefits offered in your region click here.