Senior Engineer - Platform & Data Infrastructure
Ampd Energy
IndiaremotePosted 1 month ago
Skill Required
Platform-EngineeringData-Infrastructure-EngineeringBackend-EngineeringCloud-Infrastructure-EngineeringSenior-Data-Platform-EngineerSenior-Data-Infrastructure-EngineerSenior-Platform-EngineerSenior-Data-Management-Platform-EngineerLead-Data-and-AI-Platform-EngineerSenior-Cloud-Data-Platform-EngineerSenior-Data-EngineeringSoftware-EngineerData EngineerPlatform EngineerInfrastructure EngineerSOCETLObservabilityEngineeringTypeScriptnetworkingTerraformdesigningNode.jsFirmwarebuildingGraphQLPrometheusdesignExpressKafkaCloudjQueryAPIsUnityVueSQLAWSGCPandFulltime
Key highlights
- 5–10 years backend/streaming experience required
- Competitive benefits and professional growth opportunities
- Hands-on role scaling a high-throughput telemetry pipeline
- Collaborative, multinational work environment
Role overview
Join Ampd Energy’s Software Solutions team to build and scale Enernet, the cloud platform powering fleet, device monitoring, alarms, scheduling, reporting, and mobile apps for technicians. This hands-on role focuses on engineering a fast, reliable, and cost-efficient platform to handle growing telemetry data (from ~400 to 1,000+ devices) while ensuring stability, low latency, and scalability. The position is critical for laying the foundation for higher-value product work by stabilizing the data pipeline and optimizing performance as the fleet expands.
Responsibilities
- Own the data pipeline flow from AWS IoT Core, Kafka into TimescaleDB, and legacy GCP Pub/Sub, ensuring throughput, correctness, latency, and cost targets are met.
- Provide vision and improvements for pipeline stability and scalability.
- Evolve the GraphQL contract serving web and mobile clients, designing secure, low-latency APIs (GraphQL and REST) while keeping the BFF layer clean as the schema and product grow.
- Drive down infrastructure and observability costs per device, cut latency, and remove throughput bottlenecks to prevent degradation as the fleet scales.
- Ensure monitoring spend remains proportionate to what it monitors.
- Implement Kafka partitioning, consumer groups, batched writes, backpressure, and hot-path isolation to prevent misbehaving devices from degrading the fleet.
- Use the Grafana stack to identify and track slow queries, monitor the pipeline, and maintain proactive alerts to prevent incidents.
- Evolve the live system safely with staged, reversible changes, including parallel runs, shadow validation, and rollback to avoid disruptions.
- Review frontend and mobile work, understand the GraphQL contract, and trace bugs from a Vue chart to a Timescale query to a Kafka consumer.
- Stabilize and cost-optimize the telemetry path to meet latency and throughput targets as the fleet grows, with observability to prove performance.
- Extend the GraphQL contract to support new product features without regressing latency or breaking clients.
- Act as a senior voice in incident response and reduce support load that currently leaks into sprint time.
- Collaborate cross-functionally with firmware, service, and commercial teams to turn raw telemetry into reliable, actionable tooling.
Requirements
- 5–10 years of experience building and operating production backend systems, including meaningful time on high-throughput streaming or time-series data.
- Deep, practical knowledge of Kafka (partitioning, consumer groups, rebalancing, ordering, idempotency) and schema evolution.
- Strong SQL and time-series database skills (TimescaleDB, ClickHouse, InfluxDB, or equivalent).
- Experience designing secure, low-latency APIs (GraphQL and REST) for web and mobile clients.
- Proven ability to migrate a live system without downtime, using parallel writes, shadow validation, staged cutover, and rollback.
- Fluency in TypeScript/Node.js (services are largely NestJS) and AWS in production (ECS, IoT Core, networking).
- Working knowledge of Terraform or other Infrastructure as Code (IaC) tools.
- Ability to write clear technical documentation for internal and external stakeholders.
- Operator’s mindset with first-hand experience managing live systems during production incidents, informing architecture and engineering decisions.
- Strong communication skills to explain systems clearly to diverse audiences (e.g., firmware engineers, service technicians, CEOs).
- Ability to set technical direction, review rigorously, and improve the engineers around you.
- Bias toward simplification, including deleting code and tackling legacy technical debt.
Nice to have
- Prior experience with Go (services are moving to Go for latency-sensitive paths; productivity in Go is expected quickly).
Benefits
- Competitive benefits package.
- Strong opportunities for professional growth.
- Collaborative, multinational work environment.
Additional details
- Ampd Energy is transforming construction with emission-free, advanced battery energy storage systems (BESS) to replace diesel generators, combining hardware with software and AI to reduce emissions, noise, and operational complexity.
- The company envisions an emission-free future and develops robust, versatile, and easy-to-use clean energy products.
- Ampd Energy seeks individuals with entrepreneurial drive, adaptability, and enthusiasm who thrive in continuous learning environments.
- All information provided will be treated in strict confidence and used solely for recruitment purposes.
- Ampd Energy is an equal-opportunity employer; all candidates are assessed on merit without regard to age, race, gender, sexual orientation, religion, nationality, marital status, political affiliation, or any other protected factor.