Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Chennai, Tamil Nadu, India
  • Mumbai, Maharashtra, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / Together AI

Senior Software Engineer — Infra Agent Systems Remote India

Together AI

IndiaremotePosted 3 days ago
Together AI logo

Skill Required

EngineeringSoftware EngineerSoftware DeveloperSystem DesignObservabilityTypeScriptKubernetesPrometheusSalesforceautomationITILbuildingTestNGPythonArgoCDdesignKafkaCloudRustAPIsRAGandGenerative AIGoAI

Key highlights

  • Remote based in India
  • 5+ years of experience required
  • Go, TypeScript, Python, or Rust proficiency
  • Kubernetes and GitOps experience required
  • GPU/Bare-metal infrastructure experience (preferred)

Role overview

Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact.

Responsibilities

  • Work across Infrastructure Agent Systems: Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack.
  • Work across Core Agent Platform: Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve.
  • Deliver software and also operate and support it in production.
  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.
  • Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
  • Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.
  • Turn what agents learn in production into reliable, reviewed software and automation.

Requirements

  • 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms.
  • Strong systems design skills and experience owning significant systems from design through production.
  • Depth in at least one of the following: AI agent systems, orchestration, tool use, evaluation, or grounding; Knowledge graphs or graph data modeling; Search, retrieval, ranking, RAG, or semantic search systems.
  • Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.
  • Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.
  • Comfortable working across languages such as Go, TypeScript, Python, or Rust.

Nice to have

  • Experience in GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers.
  • Experience in graph databases.
  • Experience in event-driven systems and messaging platforms such as NATS or Kafka.
  • Experience in observability platforms such as Prometheus and Grafana.
  • Experience in building evaluation frameworks or improving the quality and reliability of LLM-powered systems.

Additional details

  • We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure.
  • There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure.
  • You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure.
  • Remote based in India
  • Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
  • We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama.
  • We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.
  • Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
  • Originally posted on Himalayas

Similar jobs open now

Twilio logo
remote
Digital Marketing Manager
Twilio
Remote - India
View details
ABB logo
Senior Project Engineer -SIS
ABB
Bangalore, Karnataka, India
View details
Cognyte logo
Product Manager
Cognyte
Pune, Maharashtra, India
View details
IBM logo
Data Engineer-Data Platforms-AWS
IBM
Pune, Maharashtra, India
View details
Apply now
LocationIndia
TypeFulltime
Posted9/3/2026
Apply byOpen

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
Twilio logo
Digital Marketing Manager
Twilio
ABB logo
Senior Project Engineer -SIS
ABB
Cognyte logo
Product Manager
Cognyte
Apply now