⚡ Caught by PythonWatcher · 2h ago

Evaluation Engineer

location Remotelevel senior
E
Posted on LinkedIndirect hiring post
About the Role We are seeking an experienced Evaluation Engineer with a strong background in AI evaluation and quality assurance. This role focuses on developing robust evaluation frameworks and automating AI quality metrics to ensure high-quality, reliable AI systems. Key Responsibilities • Build evaluation frameworks for AI and LLM applications • Develop automated evaluation pipelines and AI quality metrics • Analyze AI model behavior and production drift • Design adversarial and red-team evaluation scenarios • Support AI governance and release validation Requirements • Strong experience with Python (Pandas, SQL, PyTest) • Hands-on experience with TypeScript and Playwright • Expertise in Trajectory-based and Trace-based Evaluation • Experience with LLM-as-a-Judge and Agent-as-a-Judge methodologies • Knowledge of pass^k evaluation, planning metrics, and error analysis • Experience with OpenTelemetry and Production Drift Monitoring • Strong understanding of Prompt Injection testing and OWASP security practices • Experience with Synthetic Ground Truth generation and qTest Good to Have • Experience with Braintrust, Promptfoo, DeepEval, Ragas, Arize Phoenix, or Langfuse • Commercial Real Estate (CRE) domain knowledge
SKILLS MENTIONED
PythonPandasSQLPyTestTypeScriptPlaywrightTrajectory-based EvaluationTrace-based Evaluation
Your application is already draftedCover letter + CV tailored to this post. Review it, then send — nothing goes out without your click.
Apply to this role

Your watcher also caught recently

AI Engineercaught 2h agoGenAI Engineercaught 2h agoStaff Software Engineercaught 2h agoSenior AI Engineercaught 2h agoAI Quality & Automation Engineercaught 2h agoLead Forward Deployment Engineercaught 2h ago