We're looking for a Cloud & DevOps Engineer who is as comfortable reading and writing code as they are provisioning infrastructure. You'll own how our web platform is built, shipped, run, and monitored; and when something breaks, you'll be able to dive into the codebase, find the root cause, and fix it yourself rather than just handing it back to a developer. This is a hands-on, ownership-heavy role. You'll be the person who keeps deployments smooth, environments healthy, and the on-call pager quiet, and who makes the whole team faster by automating the boring parts.
Responsibilities
- Design, build, and maintain CI/CD pipelines that take code from commit to production safely and repeatably.
- Own our cloud infrastructure end to end.
- Containerize and orchestrate services, and keep local, staging, and production environments consistent.
- Set up and maintain observability: logging, metrics, alerting, and dashboards so issues are caught before users notice.
- Debug production incidents across the stack that includes reading application code, reproducing the bug, and shipping the fix.
- Harden the platform: manage secrets, access control, backups, and disaster recovery.
- Write scripts and internal tooling to automate repetitive operational work.
- Partner with developers on release planning, rollbacks, and performance tuning.
Requirements
- 2+ years in DevOps, SRE, platform, or cloud engineering.
- Proven experience running production workloads on a major cloud provider.
- Strong CI/CD, containerization, and Linux fundamentals.
- The ability to independently trace a bug into application code and fix it, then push the change through your own pipeline.
- Solid grasp of security basics: secrets management, least-privilege access, and safe handling of sensitive data.
- Clear communication and calm, methodical incident handling.
Nice to have
- Experience with autoscaling, blue-green or canary deploys, and cost optimization.
- Familiarity with API-driven, multi-service backends and asynchronous/background job processing.
- Prior on-call or incident-response experience with postmortem discipline.
- Exposure to third-party API and payment integrations.
Additional details
- Tech stack you'll work with: You don't need deep expertise in every item, but you should be productive across most of these and able to learn the rest quickly:
- Cloud: AWS (compute, storage, IAM, networking, managed logging/monitoring).
- CI/CD: Pipeline tooling (e.g. GitHub Actions, CodePipeline/CodeDeploy, or similar), automated builds, tests, and deploys.
- Containers: Docker and Docker Compose; container orchestration experience welcome.
- Web serving & networking: Nginx or a similar reverse proxy, TLS/HTTPS, DNS, load balancing.
- Backend runtime: Python (FastAPI or similar ASGI frameworks), you should know enough to read, debug, and patch application code.
- Frontend awareness: A JavaScript/TypeScript build toolchain (Node, Vite/React or similar), same criteria, you should know enough to build, deploy, and troubleshoot the client app.
- Data stores: A document/NoSQL database (e.g. MongoDB) and an in-memory store/cache (e.g. Redis).
- OS & scripting: Linux administration and shell scripting; comfort on the command line.
- Version control: Git and github that includes branching, PR workflows, resolving conflicts, and clean commit hygiene.
- Observability: Centralized logging, metrics, and error-tracking tooling.