Staff SRE for Cloud Network Infrastructure Team (Managing edge in multi-cloud(AWS & GCP), mTLS, Global Routing, AI Automation)
Okta
Bengaluru, IndiaPosted 1 month ago
Skill Required
SW Eng - Infrastructure-672AI EngineerSite Reliability EngineerCloud EngineerInfrastructure EngineerNetwork EngineerCloudAWSGCPAIBackup and RecoveryEngineeringKubernetesPrometheusnetworkingObservabilityUnreal EngineTerraformWiresharkdesigningdebuggingsecuritybuildingDatadogPythonDockerTCP/IPdesignArgoCDNginxAPIs
Key highlights
- 8+ years of production Cloud Infrastructure Engineering experience
- 5+ years of hands‑on experience with enterprise cloud networking constructs (AWS TGW, Direct Connect, GCP networking)
- Immersive in-person onboarding experience
- Bachelor’s degree in Computer Science or related field (or equivalent professional experience)
Role overview
The Staff Software Engineer - Infrastructure at Okta will architect, evolve, and secure Okta’s global core network fabric and multi‑cloud platform layers. The role focuses on building highly resilient edge management in AWS and GCP, executing enterprise‑wide cloud migrations, enforcing Zero‑Trust cryptographic controls, integrating AI‑driven automation, and leading DDoS mitigation and platform automation while maintaining high‑availability SLAs and mentoring a small local team.
Responsibilities
- Design and engineer Next-Gen traffic entryways using Cloudflare for SaaS, driving migration to decentralized edge infrastructure, eliminating legacy hardware routing bottlenecks, shifting compute to Edge Workers, and lowering latency for millions of global requests.
- Build, scale, and maintain Okta’s multi‑cloud backbone across AWS and GCP, architecting hub‑and‑spoke models with AWS Transit Gateway, AWS Global Accelerator, and GCP Network Connectivity Center to unify converged enterprise product networks (Harmony).
- Implement and manage strict cryptographic security controls at scale, enforcing TLS 1.3 externally and deep mTLS mesh architectures internally across microservice proxies (Envoy/Nginx).
- Innovate on platform operations by designing deterministic playbooks and automated telemetry ingestion pipelines using AI models (LLMs/Agents), moving from reactive debugging to autonomous anomaly detection and self‑healing systems.
- Provide first‑line technical triage during high‑volume volumetric and application‑layer (L7) DDoS attacks during India Data Center hours, analyzing real‑time proxy/VPC flow logs and implementing rapid WAF mitigations to preserve cell isolation and tenant availability.
- Champion Infrastructure‑as‑Code (IaC) principles, continuously identifying and eliminating network configuration drifts across heterogeneous cloud landing zones by maintaining high‑quality Terraform and automation modules.
- Create detailed global architecture blueprints, technical documentation, and disaster recovery runbooks; mentor a lean, high‑performing 3‑member local team while collaborating with US‑time engineering leaders.
Requirements
- 8+ years of experience in production Cloud Infrastructure Engineering, Systems Engineering, or Core Cloud Networking roles.
- 5+ years of hands‑on experience with deep enterprise cloud networking constructs: AWS Transit Gateway, VPC peering, AWS Global Accelerator, Direct Connect, and GCP equivalents (Shared VPCs, Cloud Routers).
- Strong working knowledge of Edge Delivery networks and edge compute stacks (Cloudflare for SaaS, Cloudflare Tunnels, Custom Hostname routing, Edge Workers).
- Deep understanding of network layers and security protocols: HTTP/S routing, TCP/IP stack debugging (tcpdump, Wireshark), Public Key Infrastructure (PKI), certificate rotation pipelines, and proxy layers (Envoy, Nginx, API Gateways).
- Advanced proficiency with Infrastructure‑as‑Code (IaC) tools, specifically Terraform, for managing complex, multi‑account and multi‑region landing zones.
- Robust scripting and software development skills in Python or Go for building infrastructure tools, log parsers, and platform automation frameworks.
- Expertise with large‑scale telemetry aggregation, logging, and monitoring systems (Splunk, Datadog, Prometheus, Grafana, AWS CloudWatch, VPC Flow Logs).
- Familiarity with containerized infrastructure elements (Docker, basic Kubernetes deployments) as they interface with public network load balancers.
- Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field (or equivalent professional experience).
Nice to have
- Practical experience or strong architectural interest in implementing AI‑assisted automation layers (LLM APIs, agentic orchestration frameworks) for parsing infrastructure metrics or automating repetitive on‑call tasks.
- Prior experience supporting high‑availability, multi‑tenant SaaS platforms with strict cell‑based architectures and zero‑downtime requirements.
- AWS Certified Advanced Networking - Specialty, Google Cloud Certified Professional Cloud Network Engineer, or equivalent industry networking credentials.
Benefits
- Immersive, in-person onboarding experience designed to accelerate impact and connect to mission from day one.
- Global community spanning over 20 offices worldwide, fostering connection and collaboration.
- Equal Opportunity Employer commitment to inclusive hiring.
Additional details
- Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential of AI.
- Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era.
- This work requires a relentless drive to solve complex challenges with real-world stakes.
- We are looking for builders and owners who operate with speed and urgency and execute with excellence.
- This is an opportunity to do career-defining work.
- We're all in on this mission.
- Workforce Identity Cloud provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers.
- If you like to be challenged and have a passion for solving large-scale automation, global traffic routing, and resilient cloud infrastructure problems, we would love to hear from you.
- The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self‑educate on new network topologies, multi‑cloud ecosystems, and agentic engineering models.
- Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.
- If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation.
- Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.
- P24489_3505296
- #LI_Hybrid