Site Reliability Engineer
CXM
WorldwideremotePosted 3 days ago
Skill Required
Site-Reliability-EngineeringSREPlatform-EngineeringProduction-EngineeringReliability-EngineeringSite-Reliability-EngineerSite-Reliability-Operations-EngineerSite-Reliability-Engineering-JobsDevOps-Site-Reliability-EngineerSite-Reliability-Engineering-(SRE)Site-Reliability-Engineer-IISite Reliability EngineerSystem DesignSOCWindows ServertroubleshootingMicroservicesObservabilityCD pipelinesEngineeringPostgreSQLKubernetesPrometheusautomationShell ScriptingTerraformdebuggingbuildingPythonDockerdesignDevOps.NETCI/CDCloudDesign PatternsC#Fulltime
Key highlights
- Mid-Level (3–5 years)
- Remote (Americas, LatAm preferred)
- Americas time zones (UTC-3 to UTC-8)
- Rotation aligned with the London trading day
- Full-time, Permanent
- .NET/C#, Windows Server, AWS, Aurora PostgreSQL, Prometheus, Grafana, Terraform
Role overview
Join our Platform & Production Reliability team and help ensure the reliability, performance, and availability of our mission-critical trading systems. As an Application Site Reliability Engineer (SRE), you will own the day-to-day reliability of our .NET/C# services running on Windows, starting with our in-house liquidity bridge that connects MetaTrader trading servers to external liquidity providers. Over time, you will expand your impact across related trading and back-office services. This is a hands-on role for an engineer who enjoys solving production challenges, improving observability, automating operations, and building resilient systems where uptime directly impacts customer experience.
Responsibilities
- Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
- Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.
- Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
- Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
- Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Improve deployment safety, release automation, and rollback strategies.
- Partner with developers to improve application operability, resilience, and fault isolation.
- Automate operational tasks through scripting and infrastructure automation.
- Create and maintain runbooks, operational documentation, and incident response procedures.
- Continuously improve monitoring, alert quality, automation, and platform reliability.
- Debug and support .NET/C# applications in production.
- Support Windows Server environments.
- Script and automate operational tasks using PowerShell, Python, or Bash.
- Work with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
- Work with metrics, logging, tracing, and alerting best practices.
- Work with modern CI/CD pipelines.
- Knowledge of deployment strategies, release automation, and rollback mechanisms.
- Work with AWS infrastructure.
- Use Terraform or other Infrastructure as Code (IaC) tools.
- Troubleshoot and support Aurora PostgreSQL or other relational database platforms.
- Work with SLIs & SLOs, error budgets, incident response, root cause analysis (RCA), and alert design.
- Participate in production operations.
Requirements
- Strong experience debugging and supporting .NET/C# applications in production.
- Hands-on experience with Windows Server environments.
- Strong PowerShell scripting skills.
- Experience with Python or Bash.
- Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
- Solid understanding of metrics, logging, tracing, and alerting best practices.
- Experience with modern CI/CD pipelines.
- Knowledge of deployment strategies, release automation, and rollback mechanisms.
- Experience working with AWS.
- Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.
- Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.
- Understanding of SLIs & SLOs, error budgets, incident response, root cause analysis (RCA), and alert design.
Nice to have
- Experience supporting high-availability or low-latency financial or trading systems.
- Familiarity with MetaTrader environments or financial technology platforms.
- Experience with distributed systems and microservices.
- Knowledge of OpenTelemetry or similar observability frameworks.
- Exposure to Docker, Kubernetes, or containerized environments.
Additional details
- Team: Platform & Production Reliability
- Location: Remote (Americas, LatAm preferred)
- Working Hours: Americas time zones (UTC-3 to UTC-8)
- Employment Type: Full-time, Permanent
- Experience Level: Mid-Level (3–5 years)
- Technology Stack: .NET/C#, Windows Server, AWS, Aurora PostgreSQL, Prometheus, Grafana, Terraform
- On-call: Rotation aligned with the London trading day