Senior Software Engineer - Ingestion
Databricks
Bengaluru, IndiaPosted 1 month ago
Skill Required
Engineering - PipelineSoftware EngineerSoftware DeveloperSystem Designrelated fieldSQL ServerPrototypingEngineeringDatabricksSalesforceServiceNowSharePointdebuggingbuildingWorkdayPythonOracledesignC++ScalaUnityCloudJavaSQLAI
Key highlights
- 5+ years of production-level experience required
- BS (or higher) in Computer Science or related field required
- Comprehensive benefits package
- Work on large-scale distributed systems processing hundreds of TB of data daily
Role overview
At Databricks, we are passionate about enabling data teams to solve the world's toughest problems by building and running the world's best data and AI infrastructure platform. Ingesting data into the Lakehouse is a strategic area of investment, and Lakeflow Connect aims to provide ready-to-use, point-and-click connectors for a wide variety of sources, including enterprise applications, databases, cloud storage, message queues, and local files. This role involves working closely with other products to embed Connect into Databricks surfaces and requires engineers with experience in core Database internals to extract data from OLTP systems efficiently while minimizing load on production systems.
Responsibilities
- Solve real business needs at large scale by applying your software engineering.
- Deliver a highly scalable, available, and fault-tolerant engine processing hundreds of TB of data daily across thousands of customers
- Low level systems debugging, performance measurement & optimization on large production clusters.
- Build architecture design, influence product roadmap, and take ownership and responsibility over new projects
- Use your deep experience to help prevent and investigate production issues.
- Plan and lead complicated technical projects that work with several teams within the company.
- Break down complex problems quickly into potential solutions, knowns, and unknowns, and de-risk (through prototyping/validation).
Requirements
- BS (or higher) in Computer Science, or a related field
- 5+ years of production level experience in one of: Python, Java, Scala, C++, or similar language.
- Experience developing large-scale distributed systems from scratch
- Hands-on experience in developing and operating backend systems.
- Ability to contribute effectively throughout all project phases, from initial design and development to implementation and ongoing operations, with guidance from senior team members.
Nice to have
- Experience in areas like Database replication, backup, transaction recovery at one of the major database vendors (Microsoft SQL Server, Oracle, IBM etc)
Benefits
- Comprehensive benefits and perks that meet the needs of all employees (region-specific details available via provided link).
Additional details
- P-1403
- Lakeflow Connect is a key platform capability, and every surface in Databricks (Dashboards, Notebooks, SQL, AI) requires ingestion capabilities.
- A key part of Connect is to extract data from OLTP systems while imposing minimal load on production systems, using techniques such as incremental data capture, log parsing, etc.
- Databricks is the Data and AI company, with over 20,000 organizations worldwide relying on its platform. Headquartered in San Francisco with 30+ offices globally, it offers a unified platform including Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog.
- Databricks is committed to fostering a diverse and inclusive culture where everyone can excel, with hiring practices that meet equal employment opportunity standards.
- If access to export-controlled technology or source code is required for job duties, Databricks may apply for a U.S. government license at its discretion and may decline to proceed with an applicant on this basis alone.