npx skills add https://github.com/nvidia/skills
promptingcompany/nv-skills
nemo-evaluator-plugin
Use when working on the Evaluator plugin CLI, jobs, SDK-backed specs, metric types, or plugin-owned Evaluator skills.
Installation
npx skills add promptingcompany/nv-skills --skill nemo-evaluator-plugin
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
INVOKE THIS SKILL when building evaluation pipelines for LangSmith. Covers three core component…
4.2K installsEvaluate models, datasets, and agents with the NeMo Evaluator plugin. Use for metric selection,…
1.8K installsHandles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, runn…
1.1K installsHandles LLM-as-judge and code evaluator workflows on Arize including creating/updating evaluato…
2.5K installsTechnology stack evaluation and comparison with TCO analysis, security assessment, and ecosyste…
904 installsEvaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) …
741 installsAlso in this package
Other skills from promptingcompany/nv-skills · top by installs.
npx skills add promptingcompany/nv-skills
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Also listed on
Alternate registries and mirrors of this skill.
Repository health
main
Skill metadata
Parsed from SKILL.md frontmatter.
More metadata
- owner
- nemo-platform
- maturity
- active
Package contents
Files included with this skill beyond the listing page.
-
skill md
SKILL.md5,225 B -
docs
SUMMARY.md146 B
History
- First seen on skills.sh
- First recorded snapshot · 39 installs
SKILL.md
Evaluator Plugin
Use this skill for evaluation tasks against a running NeMo Platform server. The plugin-backed CLI interface is nemo evaluator; the legacy generated nemo evaluation API command group is not the target surface for new guidance.
CLI Interface
Prerequisites
- all commands in this file assume that the shell's working dir is at the root of the Nvidia-NeMo/nemo-platform repo
- activate the Python virtual environment before invoking the
nemoCLI:source .venv/bin/activate
Check plugin status from the CLI:
nemo evaluator info
Metric Types
Explore Available Metrics
To view available metric names, run:
nemo evaluator metric-types
To view a specific metric schema, pass a metric name from the metric_types list above:
nemo evaluator metric-types <metric-name>
Inspect all the registered metric schema contracts:
nemo evaluator evaluate explain
Note: use
nemo evaluator evaluate explainas the source of truth for the current plugin input schema. It will return a large json schema response, so strongly prefernemo evaluator metric-typeswhen you only need metric names and corresponding schemas.
Evaluation Spec
Evaluation spec is a payload that is provided to CLI as an input to execute evaluation.
At a high level, a spec describes:
metrics: bundled Evaluator SDK metric configurationsdataset: inline rows to evaluate or platform FilesetRef that contains the datasetparams: optional Evaluator SDK execution parameterstarget: optional model or agent target for online evaluation
See the LLM-judge spec example at [assets/specs/llmasjudge.json](./assets/specs/llmasjudge.json).
Metric Bundle Payloads
The checked-in [spec examples](./assets/specs) use bundled SDK metrics. The fields under metrics[*].payload are generated by bundle_metric(metric, CloudpickleMetricBundlePackager()).
To see the pattern for configuring a pre-defined SDK metric, for example ExactMatchMetric, and converting it into bundled metric JSON, inspect buildmetricbundleexample() in [generateexamplespecs.py](./scripts/generateexample_specs.py) and run:
uv run --frozen python skills/nemo-evaluator-plugin/scripts/generate_example_specs.py
Run Evaluations
Run Using File Spec Reference
When using the nemo evaluator evaluate run command, results are saved into local temporary directories and the link is printed to stdout. Prefer the --spec-file named argument over inline shell JSON because metric bundles include serialized payloads. Examples of various specs are provided in the [assets/specs](./assets/specs/) directory.
Evaluate using exact-match metric
See the spec example at [assets/specs/exactmatchmetric.json](./assets/specs/exactmatchmetric.json).
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_metric.json
Evaluate using a benchmark metric set
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_benchmark.json
Evaluate using LLM-Judge metric
Uses an LLM to score responses. See the spec example at [assets/specs/llmasjudge.json](./assets/specs/llmasjudge.json).
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/llm_as_judge.json
Run Evaluation As A Durable Job
Use the nemo evaluator evaluate submit command to create a durable evaluation job. The response of this command returns a job handler object instead of the evaluation result.
nemo evaluator evaluate submit \
--spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_metric.json
The submit response includes the generated job's name field, for example nemo-evaluator-zlhn1ecd. Wait for the job to complete, then list and download the job results.
nemo jobs get-status <job-name>
nemo jobs get <job-name>
nemo jobs results list <job-name>
nemo jobs results download aggregate-scores --job <job-name> --output-file aggregate-scores.json
nemo jobs results download row-scores --job <job-name> --output-file row-scores.jsonl
Python SDK Interface
Evaluator Python SDK client is exposed as evaluator variable on NeMoPlatform instance:
from nemo_platform import NeMoPlatform
platform_client = NeMoPlatform(base_url="http://localhost:8080")
status = platform_client.evaluator.plugin_status()
See examples of using the plugin SDK interface in [pluginsdkexamples.py](./assets/examples/pluginsdkexamples.py).
Security
Make sure not to print any secrets to stdout since this can be collected as logs
Additional Resources
For LLM-judge setup notes, see [LLM Judge Notes](references/llm-judge.md).
For evaluator API key auth, see [Evaluator API Auth](references/api-auth.md).
For local and cluster troubleshooting, see [Evaluation Troubleshooting](references/troubleshooting.md).