AI Platform Engineer
PairSoft
IndiaremotePosted 3 days ago
Skill Required
AI-Platform-EngineeringAI-ML-InfrastructurePlatform-EngineeringLLMOpsBackend-EngineeringAI-Platform-EngineerPlatform-AI-EngineerAI-ML-Platform-EngineerLead-AI-Platform-EngineerSenior-AI-Platform-EngineerAI-Agent-Platform-EngineerAI-Data-Platform-EngineerAI-Platform-ArchitectMachine-Learning-Platform-EngineerAI EngineerPlatform EngineerAIsoftware engineeringSystem DesignMicrosoft DynamicsGenerative AIManual TestingSOCRAGtroubleshootingObservabilityIstioEngineeringKubernetesautomationLangChainTerraformdesigningsecuritybuildingTestNGPythonOracledesignDevOpsOpenAPIFulltime
Key highlights
- 6+ years building production distributed systems
- 5+ years of professional software engineering experience
- 5+ years in MLOps, LLMOps, or ML platform engineering
- 2+ years hands-on production experience with LLM-based systems
Role overview
PairSoft is a global team transforming financial data management through automation technology that integrates with existing ERP systems. They are seeking an experienced MLOps/LLMOps Engineer to join their team and build the central AI platform, focusing on retrieval, prompt, and guardrail systems to support their procure-to-pay platform for mid-market and enterprise clients.
Responsibilities
- Design, build, and operate services in the central AI platform, ensuring clear API contracts, versioning, and SLOs.
- Write production Python for AI services and contribute to shared libraries, SDKs, and integration patterns.
- Instrument services with cost tagging, latency and error metrics, quality signals, and audit logs.
- Own on-call rotation for platform services, authoring and improving runbooks.
- Collaborate with product engineering leads to onboard AI features onto the central platform.
- Provide technical support, integration guidance, and troubleshooting to internal platform consumers.
- Contribute to Architecture Decision Records, documenting tradeoffs and pushing back on decisions when necessary.
- Set the operational bar regarding observability, alerting, incident response, and post-incident reviews.
- Own vendor evaluation for specialized tools, conducting bakeoffs and making cost, quality, and reliability tradeoffs explicit.
- Contribute to AI security posture, including PII handling, tenant isolation, prompt injection defense, and audit logging.
- Build retrieval, prompt, and guardrail systems to ensure LLM output quality.
- Develop RAG-as-a-service platform components: ingestion, chunking, embedding, retrieval quality, and hybrid search.
- Manage prompt engineering at scale including templates, evaluation, versioning, and per-tenant customization.
- Implement guardrails and content safety features like input filtering, output validation, PII redaction, and tool-use sandboxing.
- Develop agent frameworks and tool-use patterns for production workflows.
- Conduct domain-specific fine-tuning experiments and quality benchmarking.
- Manage a multi-provider model gateway with routing, fallback, retry, and rate-limit logic.
- Operate a prompt registry with versioning and rollout controls (canary, feature flags).
- Manage tenant isolation architecture for safe customer data flow.
- Implement cost attribution and budget enforcement at the gateway layer.
- Develop and maintain the observability platform: prompt and response tracing, cost per request, quality signals, and drift detection.
- Build evaluation infrastructure: golden datasets, offline evals, LLM-as-judge patterns, and regression testing.
- Create model deployment pipelines, including fine-tuned models.
- Maintain an alerting and SLO framework where quality regression is a first-class alert.
- Execute fine-tuning and RLHF pipelines when product-specific tuning is justified.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or equivalent.
- 5+ years of professional software engineering experience with a strong production track record.
- 6+ years building production distributed systems, ideally including internal developer platforms or API gateways at scale.
- 5+ years in MLOps, LLMOps, ML platform engineering, or a hybrid DevOps plus ML role at production scale.
- Hands-on experience with observability tools for LLM systems (e.g., LangSmith, Langfuse, Braintrust, Arize).
- Working knowledge of evaluation methodology for LLM systems (benchmark design, LLM-as-judge, human review).
- Working fluency in the modern LLM ecosystem: OpenAI or Anthropic APIs, at least one orchestration framework (LangChain, LlamaIndex), at least one vector database, and one observability tool.
- 2+ years of hands-on production experience with LLM-based systems (prompt engineering, RAG, evaluation, or LLM infrastructure).
- Strong Python and one of Go or Java; comfortable with async patterns, backpressure, and rate limiting.
- Experience designing multi-tenant systems with hard isolation guarantees.
- Cloud-native depth on Azure or AWS: Kubernetes, service mesh, IaC (Terraform), CI/CD.
- Experience shipping model updates safely in production using canaries, shadow evaluation, and rollback triggers.
- Comfort with the full ML lifecycle: training pipelines, serving infra, monitoring, and cost management.
- Strong grasp of AI security fundamentals: PII handling, tenant isolation, prompt injection basics.
- Ability to communicate technical decisions clearly in async writing.
- Fluent English language skills.
Nice to have
- Domain experience in procure-to-pay, ERP integration, accounts payable, procurement, or adjacent finance and operations software.
- Experience at a product company or PE-backed B2B SaaS, ideally on an internal platform team.
- Contributions to open-source AI/ML infrastructure projects.
- Experience with agent frameworks (LangGraph, AutoGen, CrewAI, or custom orchestration) in production.
- Prior experience on a founding platform team where you shipped v1 of a service.
Additional details
- PairSoft is an equal opportunity workplace.
- This role is distributed across time zones.
- Candidate Data Privacy Notice link provided.
- Originally posted on Himalayas.