arize-ai/phoenix

phoenix-evals

Build and run evaluators for AI/LLM applications using Phoenix.

All-time #7921 Trending #6301 First seen Jan 27, 2026
8-week activity · all time api

Installation

$ npx skills add arize-ai/phoenix --skill phoenix-evals

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from arize-ai/phoenix.

npx skills add arize-ai/phoenix

Browse all from arize-ai/phoenix

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Also listed on

Alternate registries and mirrors of this skill.

Repository health

Stars 11.4K
License LICENSE
Default branch main
Open issues 837
Status Active

Skill metadata

Parsed from SKILL.md frontmatter.

Version1.0.0
LicenseApache-2.0
CompatibilityRequires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.
More metadata
version
1.0.0
languages
Python, TypeScript

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 5,111 B
  • docs SUMMARY.md 81 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1,405 installs

SKILL.md

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

Task Files
Setup [setup-python](references/setup-python.md), [setup-typescript](references/setup-typescript.md)
Decide what to evaluate [evaluators-overview](references/evaluators-overview.md)
Choose a judge model [fundamentals-model-selection](references/fundamentals-model-selection.md)
Use pre-built evaluators [evaluators-pre-built](references/evaluators-pre-built.md)
Build code evaluator [evaluators-code-python](references/evaluators-code-python.md), [evaluators-code-typescript](references/evaluators-code-typescript.md)
Build LLM evaluator [evaluators-llm-python](references/evaluators-llm-python.md), [evaluators-llm-typescript](references/evaluators-llm-typescript.md), [evaluators-custom-templates](references/evaluators-custom-templates.md)
Batch evaluate DataFrame [evaluate-dataframe-python](references/evaluate-dataframe-python.md)
Run experiment [experiments-running-python](references/experiments-running-python.md), [experiments-running-typescript](references/experiments-running-typescript.md)
Run evals in a test runner (CI gate) [integrations-pytest](references/integrations-pytest.md), [integrations-vitest-jest](references/integrations-vitest-jest.md)
Create dataset [experiments-datasets-python](references/experiments-datasets-python.md), [experiments-datasets-typescript](references/experiments-datasets-typescript.md)
Generate synthetic data [experiments-synthetic-python](references/experiments-synthetic-python.md), [experiments-synthetic-typescript](references/experiments-synthetic-typescript.md)
Validate evaluator accuracy [validation](references/validation.md), [validation-evaluators-python](references/validation-evaluators-python.md), [validation-evaluators-typescript](references/validation-evaluators-typescript.md)
Sample traces for review [observe-sampling-python](references/observe-sampling-python.md), [observe-sampling-typescript](references/observe-sampling-typescript.md)
Analyze errors [error-analysis](references/error-analysis.md), [error-analysis-multi-turn](references/error-analysis-multi-turn.md), [axial-coding](references/axial-coding.md)
RAG evals [evaluators-rag](references/evaluators-rag.md)
Avoid common mistakes [common-mistakes-python](references/common-mistakes-python.md), [fundamentals-anti-patterns](references/fundamentals-anti-patterns.md)
Production [production-overview](references/production-overview.md), [production-guardrails](references/production-guardrails.md), [production-continuous](references/production-continuous.md)

Workflows

Starting Fresh: [observe-tracing-setup](references/observe-tracing-setup.md) → [error-analysis](references/error-analysis.md) → [axial-coding](references/axial-coding.md) → [evaluators-overview](references/evaluators-overview.md)

Building Evaluator: [fundamentals](references/fundamentals.md) → [common-mistakes-python](references/common-mistakes-python.md) → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

RAG Systems: [evaluators-rag](references/evaluators-rag.md) → evaluators-code- (retrieval) → evaluators-llm- (faithfulness)

Gating CI: evaluators-{code|llm}-{python|typescript} → integrations-{pytest|vitest-jest} → [production-continuous](references/production-continuous.md)

Production: [production-overview](references/production-overview.md) → [production-guardrails](references/production-guardrails.md) → [production-continuous](references/production-continuous.md)

Reference Categories

Prefix Description
fundamentals-* Types, scores, anti-patterns
observe-* Tracing, sampling
error-analysis-* Finding failures
axial-coding-* Categorizing failures
evaluators-* Code, LLM, RAG evaluators
experiments-* Datasets, running experiments
integrations-* Run evals from test runners (pytest, Vitest, Jest) as a CI gate
validation-* Validating evaluator accuracy against human labels
production-* CI/CD, monitoring

Key Principles

Principle Action
Error analysis first Can't automate what you haven't observed
Custom > generic Build from your failures
Code first Deterministic before LLM
Validate judges >80% TPR/TNR
Binary > Likert Pass/fail, not 1-5
Invariants gate, signals trend assert/expect hard invariants (CI red); log LLM-judge quality signals and gate the aggregate (acceptance criteria), not every case