--- name: ak-dev-new-evaluator-provider description: > Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV). Use this skill when you need to give the test framework's pluggable AKEvaluator interface a new first-party scoring/judge backend addressable by a short config name (e.g. "trulens"), not a one-off bring-your-own evaluator. Covers implementing score-based and LLM-as-judge evaluation, factory registration, configuration, optional dependencies, and testing. license: Apache-2.0 metadata: author: yaalalabs category: developer --- # Adding a New Test Evaluator Provider This guide walks through adding a new **built-in** evaluator provider to Agent Kernel's test framework. Use the existing DeepEval implementation (`ak-py/src/agentkernel/test/core/evaluator/deepeval.py`) as reference. Before starting, check whether you actually need this skill: if the evaluator only needs to exist for *your own* project (not addressable by every AK user via a short built-in name), you don't need any of the steps below — just subclass `AKEvaluator` anywhere importable and point `test-config.yaml`'s `evaluator:` at its dotted path. That's the "bring your own evaluator" path described in the user-facing `ak-test` skill and [`docs/docs/testing/cli-testing.md`](../../../docs/docs/testing/cli-testing.md#bring-your-own-evaluator); `examples/cli/custom-evaluator/` is a complete worked example of it. This skill is only for adding a **first-party, in-repo** provider that ships with AK and gets its own short `type` name. ## Existing Providers | Provider | Short name | Scoring mode | LLM-judge mode | Extra | |---|---|---|---|---| | DeepEval | `deepeval` | `Scorer.quasi_exact_match_score` (whole-string, normalised) | `GEval` LLM-as-judge metric | `agentkernel[test]` | | Opik | `opik` | `LevenshteinRatio` (fuzzy string similarity) | `GEval` LLM-as-judge metric | `agentkernel[opik]` | | JEV | `jev` | — (`AKMetricNotSupported`) | TypeSafe Noul (yes/no probability) | `agentkernel[jev]` | ## Architecture Overview - **`AKEvaluator`** (`ak-py/src/agentkernel/test/core/evaluator/base.py`) is the abstract base every evaluator — built-in or bring-your-own — implements. It has exactly two abstract methods: - `evaluate_by_score(case: AKEvaluationCase) -> AKEvaluationResult` — deterministic scoring, no LLM call - `evaluate_by_llm(case: AKEvaluationCase) -> AKEvaluationResult` — LLM-as-judge scoring - **`AKEvaluationCase`** carries the comparison inputs (`user_input`, `actual`, `expected`, `threshold`, `context`, `criteria`); **`AKEvaluationResult`** carries the outcome (`score`, `passed`, `metric`, `evaluator`, `reason`, `cost`, `attempts`, `metadata`). - **Error contract** (every evaluator must honor this, not just DeepEval): - Raise `AKMissingInput` when a field the requested metric needs (e.g. `case.expected`) wasn't supplied. - Raise `AKMetricNotSupported` from whichever of the two methods your backend structurally cannot implement (e.g. a pure LLM-judge service has no offline scoring mode). - Raise `AKEvaluationError` when a configured backend fails to produce a score (missing credentials, transport error, unparseable judge output). Never raise `AssertionError` and never silently return a `0.0` to stand in for a failure — `0.0` must only ever mean "scored zero", not "couldn't be scored". `Test.compare` is the only place that decides pass/fail fatality; evaluators only ever set `result.passed`. - **`Test._resolve_evaluator_class`** in `ak-py/src/agentkernel/test/test.py` is the factory. It shares the same pluggable-backend shape as guardrails, sandbox providers, and trace backends (`core/util/factory.py`'s `resolve_dotted`/`require_extra`/`AKConfigError`): an `if`-per-built-in branch with the SDK import wrapped in `require_extra` (actionable `ImportError` naming the pip extra if missing), then a dotted-path bring-your-own fallback for anything else. - Evaluator instances are cached per-process on `Test._evaluator` (keyed by the configured value), guarded by `Test._evaluator_lock` — construction happens once per distinct `evaluator:` config value, not once per `Test.compare` call. ## Step-by-Step ### 1. Create the Evaluator Provider File Create `ak-py/src/agentkernel/test/core/evaluator/.py`. Keep the provider's SDK imports inside this file only — `test/core/evaluator/__init__.py` and `base.py` stay pure Python with no optional-dependency imports at module level, so importing the `AKEvaluator` interface never requires your provider's SDK to be installed. ```python # ak-py/src/agentkernel/test/core/evaluator/.py from agentkernel.test.config import AKTestConfig from .base import AKEvaluationCase, AKEvaluationError, AKEvaluationResult, AKEvaluator, AKMissingInput class AKEvaluator(AKEvaluator): def __init__(self, config: AKTestConfig) -> None: super().__init__(config) # Lazy-init any client/model here only if evaluate_by_score never needs it # (mirrors DeepevalAKEvaluator's lazy LiteLLMModel, built only on first evaluate_by_llm call). def evaluate_by_score(self, case: AKEvaluationCase) -> AKEvaluationResult: if not case.expected: raise AKMissingInput("evaluate_by_score requires AKEvaluationCase.expected") # Deterministic, offline scoring logic here. score = ... # float return AKEvaluationResult( metric="", evaluator="", score=score, passed=score >= case.threshold, ) def evaluate_by_llm(self, case: AKEvaluationCase) -> AKEvaluationResult: if not case.expected: raise AKMissingInput("evaluate_by_llm requires AKEvaluationCase.expected") try: score = ... # call the judge except Exception as exc: raise AKEvaluationError(f" llm-based evaluation failed: {exc}") from exc return AKEvaluationResult( metric="", evaluator="", score=score, reason=..., # judge's explanation, if the backend provides one passed=score is not None and score >= case.threshold, ) ``` If a mode genuinely doesn't apply to your backend (e.g. a provider that is LLM-judge-only), raise `AKMetricNotSupported` from that method instead of faking a result. `Test.compare` does not catch it — in `fallback` mode it propagates out of `evaluate_by_score` before `evaluate_by_llm` runs — so document that users of your provider must set the matching `mode` (e.g. JEV requires `mode: llm`). ### 2. Register with the Factory Add the short name to `_BUILTIN_EVALUATORS` and a branch in `Test._resolve_evaluator_class`, both in `ak-py/src/agentkernel/test/test.py`: ```python _BUILTIN_EVALUATORS = ["deepeval", "opik", "jev", ""] # ADD THIS class Test: ... @classmethod def _resolve_evaluator_class(cls, configured: str) -> type[AKEvaluator]: if configured == "deepeval": with require_extra("test", "evaluator: deepeval"): from .core.evaluator.deepeval import DeepevalAKEvaluator return DeepevalAKEvaluator if configured == "opik": with require_extra("opik", "evaluator: opik"): from .core.evaluator.opik import OpikAKEvaluator return OpikAKEvaluator if configured == "jev": with require_extra("jev", "evaluator: jev"): from .core.evaluator.jev import JevAKEvaluator return JevAKEvaluator if configured == "": # ADD THIS with require_extra("", "evaluator: "): from .core.evaluator. import AKEvaluator return AKEvaluator if "." not in configured: raise AKConfigError( f"unknown evaluator '{configured}'; expected one of {_BUILTIN_EVALUATORS} or a dotted path to an AKEvaluator subclass" ) return resolve_dotted(configured, base=AKEvaluator) ``` A dotted `evaluator:` value (e.g. `myorg.evaluators.CustomEvaluator`) resolves via `resolve_dotted` without any factory edit at all — only add an `if` branch here for a first-party, in-repo provider you want addressable by a short name. ### 3. Add Optional Dependencies Add a new extras group to `ak-py/pyproject.toml` for the provider's SDK — don't fold it into the existing `test` extra (that one stays DeepEval's, since every test user already needs it for the framework itself). Follow the pattern of the `opik` extra, the first provider added on top of the original DeepEval-only `test` extra: ```toml [project.optional-dependencies] = [ "provider-sdk>=x.y.z", ] ``` ### 4. Add Configuration Docs `evaluator:` in `test-config.yaml` is already a free-form string on `AKTestConfig` (built-in short name or dotted path) — no config schema change is needed for a new built-in, since it's just a new value the same field accepts: ```yaml mode: fallback evaluator: ``` If your provider needs extra config fields (e.g. an API key env var name, a judge model override), read them from `AKTestConfig` the same way `DeepevalAKEvaluator` reads `self._config.llm` — don't invent a parallel config path. ### 5. Add Tests Add `ak-py/tests/test_evaluator_.py`, following the shape of `ak-py/tests/test_evaluator_deepeval.py`: exercise `evaluate_by_score` for real (offline, no network) where possible, and mock the judge call in `evaluate_by_llm` so the suite stays network-free. At minimum cover: - `evaluate_by_score`: exact/mismatch cases, threshold boundary, `AKMissingInput` when `expected` is absent - `evaluate_by_llm`: success, failure wrapped as `AKEvaluationError`, `AKMissingInput` when `expected` is absent - The factory branch: `Test._resolve_evaluator_class("")` resolves to your class, and (if the SDK is optional) the `require_extra` `ImportError` path when it's missing — see `test_resolve_evaluator_class_deepeval_missing_extra_raises_import_error` in `ak-py/tests/test_cli_tester.py` for the pattern (patching `builtins.__import__`, since a cached submodule import can otherwise mask the missing dependency). ### 6. Add an Example Add `examples/cli/-evaluator/`, following the shape of `examples/cli/opik-evaluator/` (a minimal agent, a `demo_test.py` exercising the new evaluator, and a `test-config.yaml` pointing `evaluator:` at the new short name). Register it in `.github/test-config.yaml`'s e2e matrix so it runs in CI, the way every other `examples/cli/*` entry does. ### 7. Add Documentation Neither doc page carries a literal "evaluator backend table" — both describe the built-ins in prose next to the `score`/`llm`/`fallback` mode explanations. Update every prose mention that enumerates the built-ins by name, not just one page: - [`docs/docs/core-concepts/configuration.md`](../../../docs/docs/core-concepts/configuration.md) and [`docs/docs/testing/cli-testing.md`](../../../docs/docs/testing/cli-testing.md) — the `evaluator:` field description and the score/llm mode explanations. - [`docs/docs/testing/automated-testing.md`](../../../docs/docs/testing/automated-testing.md) and [`docs/docs/testing/overview.md`](../../../docs/docs/testing/overview.md) — same prose pattern, duplicated across these pages. - [`docs/docs/agent-skills.md`](../../../docs/docs/agent-skills.md) — the skill directory rows for this skill and for `ak-dev-testing-conventions`. - `.agents/skills/ak-dev-testing-conventions/SKILL.md` — the evaluator config/mode section. - `ak-py/README.md` — the Test Configuration reference (`evaluator` field) and the test-config walkthrough section. - The user-facing `ak-test` skill (`ak-py/src/agentkernel/skills/ak-test/SKILL.md`) and its `evals/evals.json`. - Landing page inventories (`docs/src/components/*/data.tsx`): a tile in the **Observability, safety & testing** row of `IntegrationsMarquee/data.tsx` (role `Evaluator`, `href` to the automated testing page, logo or `react-icons/si` glyph), and the provider in the **Pluggable Evaluators** card's `tags` and `description` under the Observe tab in `FeatureExplorer/data.tsx`. Logo sourcing and the build check are in `ak-dev-sync-docs-from-branch`, *Docs-Site Landing and Features Pages*. - The features page (`docs/src/pages/features.tsx`): the `approaches` entry for Pluggable Evaluators under Testing & Evaluation names every built-in. ## Checklist - [ ] `ak-py/src/agentkernel/test/core/evaluator/.py` implementing `AKEvaluator` - [ ] Factory registration in `Test._resolve_evaluator_class` (`ak-py/src/agentkernel/test/test.py`) and `_BUILTIN_EVALUATORS` - [ ] Optional dependency extra in `ak-py/pyproject.toml` - [ ] Unit tests in `ak-py/tests/test_evaluator_.py` - [ ] Example in `examples/cli/-evaluator/`, registered in `.github/test-config.yaml` - [ ] Documentation updated: `docs/docs/core-concepts/configuration.md`, `docs/docs/testing/cli-testing.md`, `docs/docs/testing/automated-testing.md`, `docs/docs/testing/overview.md`, `docs/docs/agent-skills.md`, `.agents/skills/ak-dev-testing-conventions/SKILL.md`, `ak-py/README.md`, the `ak-test` skill and its `evals/evals.json` - [ ] Landing page inventories: marquee tile (`IntegrationsMarquee/data.tsx`), Pluggable Evaluators card tags (`FeatureExplorer/data.tsx`); the Pluggable Evaluators `approaches` entry in `features.tsx`