Meet your recruiter Viktoriia Tochylova viktoriia.tochylova@intellias.com https://www.linkedin.com... Vacancy details AI/ML Engineering Machine Learning Operations Engineer Strong Middle Bulgaria, Canada, Croatia, Cyprus, Egypt, Germany, India, Japan, Malta, Poland, Portugal, Saudi Arabia, Spain, Ukraine, United Arab Emirates Remote
We are looking for an Agent Evaluation Engineer to design and build enterprise-grade evaluation frameworks for agentic AI systems. In this role, you will be responsible for creating automated evaluation pipelines, deployment quality gates, and reliability measurements that ensure safe and predictable agent behavior before production release. You will work closely with AI, platform, and engineering teams to establish robust testing methodologies covering reasoning quality, tool usage, multi-turn interactions, and production feedback integration.
What project we have for you
Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
Intellia’s mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.
What you will do
- Design and implement automated evaluation frameworks for LangGraph-based agent workflows and orchestration pipelines.
- Develop build-time evaluation suites covering agent behavior, tool selection accuracy, reasoning quality, and final output quality.
- Create evaluation methodologies combining deterministic grading approaches with LLM-as-judge evaluation techniques.
- Build test harnesses for LangGraph graphs, nodes, state transitions, and agent execution paths.
- Design multi-turn conversation simulations to validate context retention, memory utilization, and workflow consistency.
- Define and maintain reliability metrics, including pass@k and pass^k methodologies, to measure both success rates and behavioral consistency.
- Implement CI/CD deployment gates that automatically block releases when evaluation thresholds are not met.
- Develop validation processes for staging environments, shadow-mode comparisons, and controlled production rollouts.
- Integrate AWS AgentCore Evaluations (on-demand and online evaluation modes) into continuous testing and quality assurance workflows.
- Transform production incidents, failures, and unexpected agent behavior into regression test cases.
- Establish evaluation metrics, quality benchmarks, and acceptance criteria for agent releases.
- Collaborate with AI engineers, platform teams, and product stakeholders to improve agent reliability and performance.
- Define observability and feedback mechanisms that connect production behavior with test framework improvements.
- Produce documentation for evaluation strategies, scoring methodologies, deployment gates, and quality standards.
What you need for this
Skills:
-
LangGraph agent evaluation instrumentation (build-time test harness for graphs and nodes)
-
Test primitives design combining deterministic graders and LLM-as-judge graders
-
Three-layer evaluation design covering tool selection/trajectory accuracy, reasoning quality, and output quality
-
Multi-turn conversation simulation and context-retention scoring
-
Multi-trial reliability methodology (pass@k for at-least-once success, pass^k for consistent success across consecutive trials)
-
CI/CD deployment gate design covering staging validation, shadow-mode traffic comparison, and A/B rollout
-
AWS AgentCore Evaluations (on-demand and online modes) as production feedback into the build-time suite
Experience:
-
4+ years building automated test or evaluation frameworks for ML, LLM, or agentic systems
-
Hands-on design of multi-layer test suites combining deterministic and LLM-as-judge grading
-
Experience implementing CI/CD quality gates that block deployment on metric thresholds
-
LangGraph or comparable agent orchestration framework experience
Nice to have:
-
AWS AgentCore Evaluations hands-on experience (CreateEvaluation, custom evaluators)
-
Shadow-mode or canary deployment experience for ML systems
-
Experience turning production incidents into regression test cases (feedback-loop automation)
What it’s like to work at Intellias
At Intellias, where technology takes center stage, people always come before processes. By creating a comfortable atmosphere in our team, we empower individuals to unlock their true potential and achieve extraordinary results. That’s why we offer a range of benefits that support your well-being and charge your professional growth.
We are committed to fostering equity, diversity, and inclusion as an equal opportunity employer. All applicants will be considered for employment without discrimination based on race, color, religion, age, gender, nationality, disability, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law.
We welcome and celebrate the uniqueness of every individual. Join Intellias for a career where your perspectives and contributions are vital to our shared success.