Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Chennai, Tamil Nadu, India
  • Mumbai, Maharashtra, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / Mindrift

Software Engineering Evaluation Specialist

Mindrift

IndiaremotePosted 7 days ago
Mindrift logo

Skill Required

Software-Engineering-Evaluation-SpecialistAI-Testing-SpecialistQA-EngineerSoftware-Evaluation-SpecialistAI-Agent-EvaluationTechnical-Evaluation-SpecialistEvaluation-EngineerSoftware-Engineering-SpecialistQA-Evaluation-SpecialistDeveloper-Evaluation-SpecialistEvaluation-SpecialistEngineering-Assessment-SpecialistAI-Evaluation-SpecialistSoftware EngineerSoftware Developersoftware engineeringEngineeringSystem AdministrationGenerative AIComputer VisionautomationdebuggingNode.jssecurityPyTorchwrittenTestNGPythonDockerPytestdesignDevOpsGitNumPyNginxLinuxJavaRustJiraShell ScriptingParttime

Key highlights

  • Up to $35/hour
  • Project-based, not permanent employment
  • Create AI evaluation tasks using Docker and pytest
  • 20-60% solve rate calibration target
  • 90-minute sample-task screen for qualification
  • 8-20 hours weekly realistic load

Role overview

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This is a project-based, not permanent employment, role where you'll design coding tasks that challenge frontier AI coding agents. Each task involves creating a self-contained Docker environment with a broken piece of software that an AI agent attempts to fix, with automated tests verifying the outcome. The deliverable is a complete task package: broken code, tests, instructions, and a reference solution proving the task is solvable.

Responsibilities

  • Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem.
  • Build a reproducible Docker environment with pinned dependencies.
  • Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.
  • Write an instruction.md that reads like a Jira ticket a developer would receive.
  • Write a reference solve.sh proving the task is solvable.
  • Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.
  • Iterate based on feedback from expert QA reviewers.
  • Later: review other authors' tasks as a QA reviewer.

Requirements

  • 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.
  • Python + pytest fluency — required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.
  • Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
  • Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
  • AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.
  • English — B2+ written.

Nice to have

  • Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.
  • Modern Python tooling (uv, poetry, pyproject.toml).
  • Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov).
  • Fuzzing or property-based testing (Hypothesis).
  • Prior contribution to agent-evaluation benchmarks or related frameworks.

Benefits

  • Paid contributions, rates up to $35/hour*
  • Task-based compensation equivalent to hourly rate, depending on performance and volume.
  • Some projects include incentive payments.

Additional details

  • Participation is project-based, not permanent employment.
  • Not in scope: Data labeling, prompt engineering. Production code to ship — you design problems and verification for AI agents. Leetcode puzzles — scenarios must look like real developer work. Not every candidate task ships — quality over quantity.
  • Not a fit: Data Science, ML, or Computer Vision engineers without backend-engineering output. Manual QA testers without automation or test authoring. Frontend-only, low-code / no-code, IT Support, or Business Analysts. Engineers who have never written pytest from scratch. Junior, intern, or assistant as the most recent role.
  • Apply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid.
  • Onboarding: ~10 hours per first task. Steady state: ~5 hours per task, 2–4 parallel tasks per author. Realistic weekly load: 8–20 hours. Higher volume available for top performers.
  • You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.
  • *Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.
  • Submit your CV via the Mindrift platform. Indicate your English level, note this role (Software Engineering Evaluation Specialist — Terminal Bench), and include a GitHub profile link if available.

Similar jobs open now

ABB logo
Senior Project Engineer -SIS
ABB
Bangalore, Karnataka, India
View details
IBM logo
Data Engineer-Data Platforms-AWS
IBM
Pune, Maharashtra, India
View details
IBM logo
Technical Consultant-AI Integration
IBM
Pune, Maharashtra, India
View details
C
Java Backend Developer
creospan private limited
Pune, Maharashtra, India
View details
Apply now
LocationIndia
TypeParttime
SalaryUSD 35 - 35
Posted8/30/2026
Apply by10/29/2026

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
ABB logo
Senior Project Engineer -SIS
ABB
IBM logo
Data Engineer-Data Platforms-AWS
IBM
IBM logo
Technical Consultant-AI Integration
IBM
Apply now