Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Chennai, Tamil Nadu, India
  • Mumbai, Maharashtra, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / Mindrift

Freelance Agent Evaluation Engineer

Mindrift

IndiaremotePosted 8 days ago
Mindrift logo

Skill Required

AI-Evaluation-EngineerQA-Automation-EngineerSoftware-Testing-EngineerAI-Assessment-SpecialistAI-TesterAgent-Evaluation-EngineerAgent-Evaluation-SpecialistEvaluation-EngineerNLPsoftware engineeringGenerative AIMachine LearningCybersecurityEngineeringJavaScriptTypeScriptautomationanalyticsbuildingPostgreSQLFastAPITestNGPythonDockerdesignReactKafkaRedisandAIContract

Key highlights

  • Up to $30/hr equivalent compensation
  • Project-based freelance work, part-time and remote
  • 5+ years software development experience required
  • English proficiency B2+ required
  • Master's degree preferred (bachelor's accepted with 5 years experience)
  • 3+ years professional experience in QA-automation/testing or cybersecurity roles

Role overview

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. The company is building a dataset to evaluate AI coding agents—specifically how well models handle real-world developer tasks. Participants will create challenging tasks and evaluation criteria within realistic simulated environments, iterating based on QA feedback to ensure fair and robust evaluation. This is freelance, project-based work, not permanent employment.

Responsibilities

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+
  • Educational qualifications: A Master's Degree in Computer Science, Software Engineering, Data Science / Data Analytics, Artificial Intelligence / Machine Learning, Computational Linguistics / Natural Language Processing (NLP), Information Systems or other related fields
  • Bachelor's degree is accepted if candidate has 5 years of experience in the field
  • Candidates should have a minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles

Benefits

  • Paid per accepted task - rate depends on qualification tier reached and task completion efficiency, up to equivalent of $30/hr
  • Take part in a part-time, remote, freelance project that fits around your primary professional or academic commitments
  • Work on advanced AI projects and gain valuable experience that enhances your portfolio
  • Influence how future AI models understand and communicate in your field of expertise

Additional details

  • This is NOT: Not data labeling, Not prompt engineering, Not writing code from scratch - the agent writes most of the code; you guide and evaluate
  • Why this is hard: Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

Similar jobs open now

Twilio logo
remote
Digital Marketing Manager
Twilio
Remote - India
View details
ABB logo
Senior Project Engineer -SIS
ABB
Bangalore, Karnataka, India
View details
Cognyte logo
Product Manager
Cognyte
Pune, Maharashtra, India
View details
IBM logo
Data Engineer-Data Platforms-AWS
IBM
Pune, Maharashtra, India
View details
Apply now
LocationIndia
TypeContract
Posted8/29/2026
Apply by10/27/2026

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
Twilio logo
Digital Marketing Manager
Twilio
ABB logo
Senior Project Engineer -SIS
ABB
Cognyte logo
Product Manager
Cognyte
Apply now