Posted today · be early

ML Annotation QA Engineer

Gather AI

IndiaremotePosted 1 day ago
Gather AI logo

Skill Required

Quality-AssuranceML-Model-QA-AnnotationAnnotation-QualityData-QualityMachine-LearningAI-QA-EngineerData-Annotation-QAAI-QA-AnalystAI-Quality-Assurance-EngineerAI-ML-Testing-EngineerML-Validation-EngineerAI-Quality-Assurance-AnalystAI-Data-Annotation-AnalystData-Annotation-AnalystAI-Data-Annotation-SpecialistMachine Learning EngineerQA EngineerMachine LearningComputer VisionEngineeringStatisticsanalyticalEmbedded CwrittenPythonCloudjQueryJiraSQLandAIFulltime

Key highlights

  • Experience: 2–5 years in ML QA, annotation quality, or analytics at an AI/ML company
  • Education: BS in Computer Science/Engineering, Electrical Engineering, or equivalent
  • Key Technologies: Python, SQL, and Jira
  • Project Focus: Warehouse forklift vision, barcode readability, and localisation analysis

Role overview

We are looking for an ML Annotation QA Engineer to own the quality of annotated data across our computer vision and machine learning programs. This role is responsible for the judgment-heavy analysis that cannot be reliably outsourced, for the decision rules behind it, and for turning annotation output into an ongoing read on how our systems are actually performing in the field. We work with an external annotation partner at production volume, and that continues. What we need in house is someone who can analyse the annotated data, make and defend the calls the vendor cannot make consistently, and build the aggregate view that shows which facilities and equipment are degrading and why. Annotation drift, a model regression, a tool bug, and genuine field degradation all look similar in a chart and require completely different responses — telling them apart is the core of the job.

Responsibilities

  • Own the judgment-heavy quality analysis on annotated data that cannot be reliably outsourced — working the daily review queue and producing verdicts and root cause in house.
  • Own, version, and refine the verdict taxonomy, decision rules, and quality guidelines for the categories you cover.
  • Build and maintain performance trackers over annotated data — error rates by facility, site, equipment, and data format over time, against an agreed baseline.
  • Detect anomalies against that baseline and flag them the day they appear rather than weeks later.
  • Run root cause analysis on flagged anomalies, distinguishing annotation error from model or system error from genuine degradation in the field.
  • Report findings to engineering and ML with reproducible evidence and stated confidence, fast enough that the issue is still observable.
  • Identify systematic failure patterns rather than one-off misses, and maintain a documented pattern library others can use.
  • Query and analyse annotation data directly with Python and SQL to test hypotheses, without waiting on extracts from anyone.
  • Feed annotation-quality findings back as concrete SOP and instruction changes when the root cause is labeling rather than system behaviour.
  • Specify annotation tool improvements — what the tool should surface so analysis stops requiring manual work — and validate the fixes.
  • Stand up quality analysis and reporting for new annotation programs as they come online.
  • Track work across Jira and contribute to pre-release validation for the behaviours you cover.
  • Take over barcode and location root cause analysis from the annotation vendor (First 90 Days).
  • Move from supervised review to owning the daily queue at the agreed review threshold, inside the expected time budget (First 90 Days).
  • Become the primary owner of verdicts and root cause for the categories you cover (First 90 Days).
  • Develop working expertise in the data and the domain behind it: what is captured, at what grain, and where the known quality limits are (First 90 Days).
  • Develop expertise in racks, locations, levels, bins and bin types; LPN, SKU, AWB and other code formats; exceptions and exception types; OCR versus barcode and the failure modes of each (First 90 Days).
  • Stand up a facility performance tracker with an agreed baseline, thresholds, and reporting cadence (First 90 Days).
  • Establish anomaly detection against the performance tracker, and get the team to the point of trusting and using it (First 90 Days).
  • Investigate flagged anomalies independently, with a root cause hypothesis and stated confidence (First 90 Days).
  • Take ownership of barcode and location verdict taxonomy and decision rules (First 90 Days).
  • Take ownership of the pattern library of known failure modes (First 90 Days).
  • Handle version control, closing documentation gaps, adding edge-case guidance, and feeding changes back to the annotation vendor’s SOPs where the root cause is labeling (First 90 Days).
  • Establish yourself as the primary owner of annotation quality analysis and root cause (First 6 Months).
  • Maintain a performance tracker the team relies on, running on a regular cadence (First 6 Months).
  • Close the reporting cycle to same day, so issues are raised while still observable (First 6 Months).
  • Standardize the verdict taxonomy and decision rules for the categories you cover (First 6 Months).
  • Build a documented pattern library of known failure modes that others can use (First 6 Months).
  • Improve the annotation tool in at least one way that measurably reduces manual effort (First 6 Months).
  • Give ML and engineering a clearer picture of what error patterns mean for model accuracy and field performance (First 6 Months).
  • Create a repeatable method for standing up annotation quality on the next program (First 6 Months).

Requirements

  • BS in Computer Science/Engineering, Electrical Engineering, or equivalent experience.
  • Experience working with annotated datasets for CV/ML, including assessing label quality.
  • Strong understanding of statistics, and able to work with data.
  • Root cause analysis skills — generating competing hypotheses and naming the evidence that separates them.
  • Familiarity with Python and SQL, or equivalent, for querying and analysing data independently.
  • Experience writing quality guidelines, decision rules, and labeling taxonomies.
  • Excellent documentation, communication, and collaboration skills.
  • 2–5 years of experience in ML QA, annotation quality, data quality, or analytics at an AI/ML company.
  • Hands-on experience with annotated ML datasets, including assessing label quality.
  • Demonstrated experience finding, diagnosing, and reporting data anomalies to a technical audience.
  • Able to get to a defensible answer from unfamiliar data without someone preparing it first.
  • Understanding of data privacy and confidentiality requirements when working with customer operational data.
  • Python
  • SQL, or an equivalent query language
  • Jira

Nice to have

  • 2+ years in quality assurance, data quality, or product support.
  • 1+ years experience with enterprise-grade ticketing systems (e.g. Jira).
  • Computer vision annotation experience as a reviewer or auditor — video event labeling, bounding boxes, polygon segmentation, counting, or classification.
  • Annotation quality methodology — gold-set validation, inter-annotator agreement, sampling design.
  • Strong spatial and geometric reasoning — relevant wherever labels describe position in physical space.
  • Experience in warehouse automation, robotics, or computer vision applications.
  • Dashboarding or BI tooling for recurring reports.
  • Small team experience.
  • Experience with annotation platforms and quality tooling.
  • Experience authoring annotation guidelines or standard operating procedures.

Additional details

  • Job Title: ML Annotation QA Engineer
  • About Gather AI: We’ve developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone.
  • Team Context: Our engineering organization spans autonomy, computer vision and machine learning, embedded and hardware systems, full-stack, and cloud, all working in parallel across multiple active product lines.
  • Collaboration: You will work closely with Machine Learning Engineers, QA, and Engineering.
  • Immediate Assignment: Warehouse forklift vision program, where barcode readability and localisation analysis are the immediate need.
  • Success Traits - Analytical: able to tell a real pattern from noise and to say honestly when the data will not settle a question.
  • Success Traits - Detail-oriented: a verdict or a tracker that is subtly wrong is worse than none at all.
  • Success Traits - Comfortable with ambiguity: willing to present competing hypotheses rather than forcing a single answer.
  • Success Traits - Persistent: in chasing a cause across sites, programs, and builds.
  • Success Traits - Resourceful: comfortable building processes and reports where they don’t yet exist.
  • Success Traits - Strong written communicator: the output of this role is reports other people act on.
  • Success Traits - Collaborative: respectful, and an excellent cross-functional partner.
  • Originally posted on Himalayas
Apply now