Production Support Engineer
Luma Financial Technologies
Skill Required
Key highlights
- Required experience: 3+ years in Production Support, SRE, or Application Support
- Notable requirement: Willingness to participate in a rotational 24/7 on-call schedule
- Key skill: Strong SQL and Linux command-line proficiency
Role overview
Luma Financial Technologies is a fintech software platform used by broker/dealer firms, RIA offices, and private banks to research, purchase, and manage alternative investments and annuities. Headquartered in Cincinnati, OH, with offices in New York, Miami, Zurich, and Lisbon, Luma provides a suite of solutions including education resources, custom structured product pricing, electronic order entry, and post-trade management. Luma is seeking a proactive and skilled Production Support Engineer to join their IT operations team. This role is responsible for overseeing and managing incidents related to engineering and production environments, ensuring system stability and reliability by coordinating with Software Engineers, Infrastructure, and Support teams to triage, diagnose, and resolve technical issues in real time.
Responsibilities
- L2/L3 Incident Response: Triage, investigate, and resolve production bugs, system outages, and customer-impacting issues within SLA guidelines.
- Proactive Monitoring: Monitor engineering and production system performance, application metrics, and server logs using tools like Datadog.
- Root Cause Analysis (RCA): Lead post-incident write-ups and collaborate with developers to design long-term fixes to prevent recurring incidents.
- Deployment & Releases: Assist with blue-green deployments, hotfixes, and scheduled production maintenance windows.
- Environment Management: Maintain high availability across production, staging, and disaster-recovery (DR) environments.
- Database & Data Fixes: Execute safely managed SQL queries, data migrations, or manual record reconciliations when required.
- Automation: Write scripts (Bash, Python, PowerShell) to automate repetitive operational tasks, log parsing, and monitoring alerts.
- Runbook Documentation: Create, update, and standardize operational runbooks, escalation matrixes, and troubleshooting guides.
Requirements
- Education: Bachelor’s degree in Computer Science, Information Technology, or a related discipline (or equivalent practical experience).
- Experience: 3+ years in a Production Support, Site Reliability Engineering (SRE), or Application Support role.
- Operating Systems: Solid command-line skills in Linux/Unix environments.
- Database / SQL: Strong ability to write complex SQL queries, analyze slow queries, and understand database schemas (MongoDB, MySQL, or MS SQL).
- Scripting: Proficiency in at least one scripting language (Python, Bash, Shell, or PowerShell) for automation.
- Monitoring Tools: Hands-on experience with log aggregation and monitoring platforms (e.g., Datadog).
- Protocols & Architecture: Understanding of REST APIs, microservices architecture, network fundamentals (TCP/IP, DNS, HTTP statuses), and web servers (Nginx, Apache).
- Soft Skills: Excellent analytical skills to troubleshoot complex system failures under pressure, alongside the communication skills required to translate technical ideas to non-technical stakeholders.
- On-Call Availability: Willingness to participate in a rotational on-call schedule (24/7 support coverage).
Nice to have
- Platform experience: Experience with cloud platforms (AWS or Azure).
- Familiarity with containerization and orchestration (Kubernetes).
- Experience with ticketing systems like Jira.
- Exposure to CI/CD tools (GitHub Actions, GitLab CI).
- Industry Exposure: Prior Exposure to fintech, insurance or broader finance services organizations.
Additional details
- Originally posted on Himalayas