Posted today · be early
Sr. Staff Engineer - Payments & Commerce Platform
HighLevel
IndiaremotePosted 1 day ago
Skill Required
EngineeringMobile-CommercePayments-EngineeringBackend-EngineeringDistributed-Systems-EngineeringSenior-Payments-EngineerSenior-Platform-EngineerSenior-Platform-EngineeringPayments-Backend-EngineerPayment-Systems-EngineerPayments-Software-EngineerPayments-EngineerSenior-Developer-Platform-EngineerStaff-EngineerPlatform EngineerMicroservicesObservabilityautomationdesigningFirebasesecuritybuildinggRPCMongoDBTestNGdesignRedisCI/CDGCPandFirewallGoCICDAIFulltime
Key highlights
- Targeting $1T+ in annual transactions
- 10+ years backend experience required; 5+ years in Go required
- No direct reports; org-level IC role
- Opportunity to design core commerce systems at planet scale
Role overview
HighLevel is an AI-powered business operating system supporting SMBs across 150+ countries, helping businesses generate over $7 billion in ecosystem value. With 2,000+ team members globally, they're building a global commerce platform to power $1T+ in annual transactions. This role is for a Principal Engineer focused on designing core commerce systems at planet scale with significant organizational influence but no direct reports.
Responsibilities
- Architect and ship multi-tenant, planet-scale services (checkout, subscriptions, payments orchestration, invoicing, tax hooks) with clear domain boundaries (DDD) and hard SLOs
- Be the custodian of API & schema design: own protobuf/ConnectRPC conventions, versioning policy, deprecation playbooks, and Buf breaking-change checks
- Guarantee resilience & availability of core payment paths: timeouts, retries with jitter, circuit breakers, idempotency keys, outbox/Saga patterns, hedged requests, and graceful degradation
- Ensure complete auditability: append-only double-entry ledger, immutable event streams, trace-linked entities (OTel trace/span IDs), tamper-evident trails, and reconciliations that tie out to the cent
- Own error boundaries end-to-end: enumerate failure domains (PSP, network, data, concurrency, quota, browser, device); design uniform error contracts; implement compensations/backfills and automated replay
- Keep track of every deployed thing: services, workers, triggers, cron, subscriptions—own the service catalog and scorecards (owners, SLOs, runbooks, PDBs, HPA/VPA, budgets, quotas, timeouts)
- Configuration & limits stewardship: enforce sane defaults across GKE, Pub/Sub, Redis, Firestore/Mongo, ClickHouse—connection pools, ack deadlines, batch sizes, TTLs, memory/FD limits, and GCP quotas
- Observability as a product: pervasive OpenTelemetry, RED/USE metrics, exemplars, trace sampling, SLO dashboards, and alerting that wakes humans only for user-impacting issues
- Production excellence: canary/blue-green rollouts, automated rollbacks, chaos drills, DR playbooks (RPO/RTO), multi-region failover strategies, and incident command on rotation
- Security & compliance by design: PCI scope minimization, tokenization/vaulting, secrets/KMS hygiene, data retention/archival, and privacy controls—embed checks in CI/CD
- Developer acceleration: pave golden paths (service templates, ADR/RFC process, linting/formatting, contract tests, ephemeral envs, load/perf harnesses) to make the right thing the easy thing
- Steward core domain evolution: orchestration → ledger → reconciliation flows with crisp invariants and consistency guarantees (read-your-writes where needed, eventual where appropriate)
- Own reliability strategy: SLIs/SLOs, error budgets, capacity planning, cost/FinOps guardrails, multi-region posture, and DR exercises
- Lead API & data governance: canonical models, schema lifecycle (compatibility matrix, migrations), data lifecycle (retention, archival, compliance)
- Practice leadership for HighLevel: design reviews, postmortems, technical strategy, coding standards, and mentorship across teams—raise the bar for the org
- Participate in hiring & team growth: help us hire, scale, and train the right team; shape interview loops, rubrics, onboarding, and ongoing learning (brown bags, reviews, pair design)
- Collaborate with cross-functional partners: Product/Marketing/Support to translate platform capabilities and constraints into roadmaps, GTM narratives, and reliable customer outcomes
- Maintain risk & roadmap: maintain a technical risk register, make build-vs-buy calls, and propose simplifications or deprecations that meaningfully reduce complexity and MTTR
Requirements
- 10+ years building and operating backend systems (at least 5+ years in Go)
- 2–3+ years acting as a Staff/Principal-level IC or Tech Lead for critical paths
- Deep proficiency with protobuf + ConnectRPC/gRPC and API lifecycle management (versioning, compatibility, contract testing, Buf)
- Distributed systems fundamentals: idempotency, exactly-once-ish via dedupe/outbox, ordering, consensus basics, backpressure, concurrency control
- Event-driven architectures on GCP (Pub/Sub), plus Redis for fast paths; strong schema design in MongoDB/Firestore and analytics/reporting patterns on ClickHouse
- Kubernetes/GKE operations at scale: autoscaling (HPA/VPA), PDBs, resource limits/requests, multi-region topologies, CI/CD, canary/blue-green
- Reliability engineering: SLIs/SLOs, error budgets, capacity & load testing, incident management, DR/BCP
- Security & compliance: secrets/KMS best practices, PCI basics (scope reduction, key rotation), and data governance (retention/archival)
- Testing discipline: unit, integration, contract, property-based, performance; test data management and deterministic environments
- Frontend collaboration: solid understanding of Vue.js + TanStack Query to shape clean API surfaces and performance budgets across the boundary
- Exceptional technical writing & communication: design docs, ADRs/RFCs, postmortems, and stakeholder updates
Nice to have
- Hands-on integrations with major PSPs/local rails (e.g., UPI, wallets, BNPL, cards/3DS2) and reconciliation at scale
- Experience with active-active or multi-region designs; chaos engineering; traffic management
- Observability leadership with OpenTelemetry at org scale (tail-based sampling, exemplars)
- FinOps experience: cost baselining, quotas, budget alarms, and workload right-sizing
- Familiarity with regulatory frameworks (PCI DSS, SOC 2/ISO 27001) and privacy laws relevant to our markets
Additional details
- Tech Stack: Backend: Go, ConnectRPC | Databases: MongoDB, Firestore, ClickHouse | Cloud: GCP (GKE), Pub/Sub, Redis, OpenTelemetry
- HighLevel supports SMBs across 150+ countries, fueling community-driven growth rooted in real customer outcomes
- Businesses operating on HighLevel have generated over $7 billion in ecosystem value
- HighLevel powers more than 4 billion API hits and 2.5 billion message events daily
- With 250 terabytes of distributed data, 250+ microservices and over 1 million domain names supported
- HighLevel enables more than 1.5 billion messages, 200 million leads and 20 million conversations monthly for the more than 1 million businesses we support
- HighLevel operates as a global, remote-first organization built for speed and ownership
- Equal Opportunity Employer; demographic information collected voluntarily for compliance purposes only
- This is an IC role with org-level influence (no direct reports), focused on designing systems, shaping standards, and growing engineers