C# Senior Engineer - AI Interaction Evaluator
G2i
WorldwideremotePosted 10 days ago
Skill Required
Senior-AI-Software-EngineerSenior-Software-AI-EngineerSenior-C#-.NET-Software-EngineerAI-Interaction-EngineerAI-Evaluation-EngineerSenior-AI-EngineerSenior-AI-Software-DeveloperSenior-AI-DeveloperSenior-C#-DeveloperAI EngineerC#AIEngineeringJavaScriptTypeScriptPythondesignOpenAPIandContract
Key highlights
- Compensation: $100–$200/hour
- Level Required: Staff / Principal-level engineer (or equivalent)
- Key benefit: 10–20 hours/week flexible contract with possible extension through early May
- Notable requirement: No production code writing; evaluation-focused role assessing AI engineering judgment
- Notable requirement: Must answer questions about whether AI behavior 'feels like something a strong engineer would actually say'
- Notable requirement: Comfortable making subjective but rigorous judgments about AI interactions
Role overview
Senior AI Interaction Evaluator (Codex / Claude Code) — a contract role evaluating the quality of AI coding agent interactions. The position focuses on assessing whether models like OpenAI Codex and Claude Code demonstrate strong engineering judgment, useful reasoning, and appropriate interaction patterns — not traditional software engineering or production code writing. Ideal candidates have Staff/Principal-level experience with TypeScript/JavaScript or Python, deep familiarity with AI-assisted development workflows, and demonstrated ability to make subjective but rigorous engineering evaluations.
Responsibilities
- Evaluate AI-generated coding interactions end-to-end
- Judge whether outputs are useful, correct at a high level, and aligned with how a strong engineer would think
- Assess the quality of explanations and reasoning, not just code
- Distinguish between different levels of response quality (e.g., what makes something a 2 vs 4)
- Provide clear, opinionated feedback on what worked, what didn't, and what felt 'off' or misleading
- Help define what great looks like when interacting with tools like Cursor
- Make subjective but rigorous judgments about AI engineering interactions
- Assess whether responses reflect strong engineering judgment and appropriate interaction patterns
Requirements
- Staff / Principal-level engineer (or equivalent experience)
- Strong background in one of the below: TypeScript / JavaScript, Python
- Hands-on experience using: OpenAI Codex, Claude Code, Cursor
- Deep familiarity with modern AI-assisted dev workflows
- Able to evaluate code without needing to fully execute or deeply review every line
- Comfortable giving direct, opinionated feedback
- High bar for what 'good engineering' looks like
Nice to have
- Experience with tools like Cursor or similar AI-first IDEs
- Prior exposure to prompt design or evaluation workflows
- Experience mentoring senior engineers or defining engineering standards
Benefits
- Contract role with flexible hours (10–20 hrs/week)
- Duration through early May with possible extension
- Project-based evaluation-focused engagement rather than traditional engineering role
- Work involving direct influence on AI model evaluation standards
Additional details
- This is not a traditional engineering role — no production code writing required
- Focus is on 'engineering taste' — whether the model thinks like a great engineer
- Position specifically evaluates interactions with modern coding agents (OpenAI Codex, Claude Code, Cursor)
- Evaluation criteria include: whether response makes sense, whether preamble and reasoning are useful, whether output reflects strong engineering judgment, whether interaction feels right to an experienced developer
- Engagement Details: Rate $100–$200/hour, Hours ~10–20 hours/week, Duration Through early May (with possible extension), Start ASAP
- Process: Take-home evaluation exercise, One behavioral interview
- Check out this Loom video for more details!
- Originally posted on Himalayas
- Role involves assessing interaction quality end-to-end, not just code correctness or syntax