Cloud Infrastructure & Platform Specialist (Databricks + AWS + Azure) (Mid-Level
Irth
Skill Required
Key highlights
- Experience: 3–5 years
- Location: Remote – India
- Department: Insights (AI/ML) – Platform Engineering
- Key Responsibilities: Infrastructure provisioning, IAM, Databricks administration, networking, security, and platform onboarding for multi-cloud data estate
- Nice-to-Have: Experience with asset integrity, GIS/geospatial data, pipeline data, or utility/energy/oil & gas compliance requirements
- Benefits include: 401(k) with company match, PTO, company-paid holidays, on-call compensation, and flexible work options
Role overview
Irth Solutions is a leading provider of cloud-based SaaS software for damage prevention, asset integrity, stakeholder engagement and land management, helping energy, utility, telecom, and infrastructure companies protect their critical network infrastructure. With nearly three decades of industry experience, Irth serves customers across North America and continues to expand its platform with new data-driven and AI-powered capabilities. Irth is building a governed, multi-cloud data estate on Databricks to power cross-product insights across Damage Prevention, Asset Integrity Pipeline, Land Management, and Stakeholder Management. We are looking for a Cloud Infrastructure & Platform Specialist to own and operate the underlying cloud and Databricks infrastructure that enables our Data Engineering, Data Science, and MLOps teams to build and deliver data and AI solutions. You will be the go-to engineer for infrastructure provisioning, identity and access management, Databricks administration, networking, security, and platform onboarding. You will help provision workspaces and catalogs, manage IAM roles and administrative access, establish secure platform guardrails, and resolve infrastructure bottlenecks so engineering teams can onboard and ship efficiently. This role is ideal for a mid-level cloud/platform engineer who enjoys solving infrastructure problems, automating repetitive tasks, and serving as the technical enabler for other engineering teams.
Responsibilities
- Provision, configure, and manage core cloud infrastructure across AWS and Azure supporting Irth's Databricks-based data estate, including VPCs/VNets, subnets, route tables, S3 buckets/Azure Storage Accounts, compute resources, and private endpoints
- Establish and maintain secure DEV, QA, and PROD environments with appropriate isolation and platform guardrails; support cloud landing-zone implementation and ensure environments follow established security, governance, and cost policies
- Use Infrastructure-as-Code (IaC) technologies such as Terraform, Azure ARM/Bicep, and AWS CloudFormation; version, review, test, and deploy infrastructure changes through controlled engineering workflows
- Troubleshoot infrastructure, networking, and cloud-service issues affecting data and ML workloads
- Provision and manage IAM roles, service principals, managed identities, groups, and access policies across AWS, Azure, and Databricks; implement least-privilege access and role-based access control (RBAC) aligned with enterprise security standards
- Configure and support identity federation and automated identity lifecycle management using SSO, SCIM, SAML, and OAuth; manage access across cloud resources, Databricks workspaces, Unity Catalog, storage, compute, and other platform services
- Review and process elevated or administrative access requests through controlled and auditable procedures; balance engineering-team productivity with security, compliance, and least-privilege requirements; participate in periodic access reviews and remediation
- Create, configure, and administer Databricks workspaces and associated platform resources; configure and manage Unity Catalog metastores, catalogs, schemas, storage credentials, external locations, and permissions
- Support onboarding of Data Engineering, Data Science, and MLOps projects by establishing appropriate platform resources and access; troubleshoot onboarding blockers including missing administrative privileges, cluster-policy conflicts, workspace configuration issues, catalog and schema permission problems, storage-access failures, and identity and authentication issues
- Define and maintain Databricks cluster policies to control compute configurations, security, and cost; configure instance profiles, managed identities, storage credentials, and other mechanisms required for secure data access; establish workspace-level guardrails for compute, networking, storage, and user access
- Partner with Data Engineering, Data Science, and MLOps teams to ensure new projects can be onboarded efficiently with appropriate catalogs, schemas, permissions, and infrastructure already available; act as the primary escalation point for Databricks platform and infrastructure issues
- Implement secure private connectivity between Databricks, cloud services, and enterprise resources; configure and troubleshoot AWS PrivateLink/VPC endpoints, Azure Private Link/private endpoints, VNet/VPC peering, secure cluster connectivity, network security groups, firewall controls, and egress restrictions
- Implement encryption and key-management practices using AWS KMS, AWS Secrets Manager, Azure Key Vault, and equivalent enterprise security services; establish appropriate network segmentation and isolation between environments; support security reviews and enterprise compliance requirements
- Maintain audit logging and provide evidence for access reviews, administrative activity, infrastructure changes, data access, and retention controls; help ensure infrastructure and platform configurations align with applicable regulatory and organizational requirements
- Implement and enforce cloud and Databricks tagging standards across infrastructure and workloads; monitor infrastructure and Databricks spend and identify opportunities for cost optimization; support showback/chargeback initiatives and cost reporting; implement resource policies and controls that prevent unnecessary or unapproved cloud and Databricks spend
- Automate recurring infrastructure and access-management activities, including workspace provisioning, catalog and schema setup, IAM role assignment, access provisioning, environment configuration, and standard platform onboarding; build reusable automation that reduces manual provisioning and improves onboarding turnaround time
- Monitor platform health and respond to infrastructure and access-related incidents; troubleshoot cloud, networking, identity, and Databricks issues affecting production workloads; maintain clear operational runbooks, architecture documentation, access procedures, troubleshooting guides, and onboarding documentation
Requirements
- 3–5 years of experience in cloud infrastructure, platform engineering, cloud administration, or a closely related role
- Hands-on experience working with both AWS and Microsoft Azure, including IAM and role provisioning, cloud networking, storage services, and compute and platform resources
- Hands-on experience administering Databricks, including workspace provisioning and configuration, Unity Catalog, catalogs and schemas, metastores, cluster policies, and workspace and data-access controls
- Working knowledge of Infrastructure-as-Code (IaC) tools, with Terraform preferred; experience with Azure ARM/Bicep or AWS CloudFormation is also valuable
- Strong understanding of Identity and Access Management (IAM) concepts, including RBAC, least-privilege access, service principals, managed identities, SSO, SCIM, and SAML
- Familiarity with cloud networking and security fundamentals, including VPC/VNet architecture, subnets and routing, private connectivity, network segmentation, secrets management, encryption, and key management
- Ability to quickly troubleshoot infrastructure, networking, identity, permission, and platform-access issues that may block engineering teams
- Strong communication skills, with the ability to explain infrastructure and access issues clearly to Data Engineering, Data Science, MLOps, and other technical teams
- Familiarity with enterprise compliance and security frameworks, including SOC 2, ISO 27001, and GDPR
- Understanding of audit logging, access reviews, administrative activity tracking, and evidence collection
Nice to have
- Experience supporting multiple engineering teams—including Data Engineering, Data Science, and MLOps—with workspace, catalog, schema, compute, and access provisioning
- Experience with infrastructure and platform CI/CD, including GitHub Actions, Databricks Asset Bundles (DABs), infrastructure deployment automation, and DEV → QA → PROD environment promotion
- Knowledge of FinOps and cloud cost-management practices, including resource tagging, budgets, cost monitoring, showback/chargeback, and cost anomaly detection and alerting
- Experience implementing cloud and Databricks policies that balance security, reliability, developer productivity, and cost efficiency
- Relevant industry certifications, such as AWS Solutions Architect, Microsoft Azure Solutions Architect, Azure Administrator, Databricks Platform Administrator, or equivalent cloud or platform certifications
- Experience onboarding Asset Integrity Pipeline data, including inspection, risk, maintenance, and compliance datasets, into a governed enterprise data estate
- Experience designing catalogs, schemas, storage structures, and access controls for asset-integrity or infrastructure datasets
- Experience provisioning infrastructure and access for utility, energy, oil & gas, pipeline, or asset-integrity data sources with elevated security or compliance requirements
- Familiarity with GIS and geospatial data storage, networking, access, and platform-integration patterns relevant to pipeline and infrastructure datasets
- Understanding of regulatory and compliance reporting requirements associated with pipeline, utility, energy, or asset-integrity data
- Experience working with infrastructure and data platforms where data residency, auditability, security, and controlled access are critical requirements
Benefits
- Competitive Salary – A competitive compensation package based on experience and qualifications
- Medical, Dental, and Vision Insurance – Comprehensive insurance coverage to support you and your family
- 401(k) Plan with Company Match
- Generous Paid Time Off (PTO) – Time off to support work-life balance and personal needs
- Company-Paid Holidays – Paid holidays throughout the year
- Flexible Work Options – Work-from-home opportunities are available, depending on role and business needs
- On-Call Compensation – Additional pay for eligible on-call shifts
Additional details
- Location: Remote – India
- Department: Insights (AI/ML) – Platform Engineering
- Reports to: Data Platform & Analytics Manager