94,993 packages matching “agent-evals”
@aether-agent/evals
v0.3.1 · 6 days ago
Evaluation harness for the Aether agent SDK
No known vulnerabilities
swarmkit-eval
v0.1.1 · 7 days ago
Evaluation infrastructure for the swarmkit ecosystem — (harness x model x task x arm x seed) agent evals with ground-truth scoring, cost-matched Pareto reporting, and scalable parallel execution.
No known vulnerabilities
@evalpal/sdk
v0.1.0 · 16 days ago
TypeScript SDK for EvalPal — AI agent evals, tracing, and Autopilot diagnostics. Integrations for Vercel AI SDK, LangChain, LlamaIndex, and OpenTelemetry.
No known vulnerabilities
@sanity/agent-evals
v0.0.6 · 3 months ago
Vitest-style evaluation framework for Sanity Agent
No known vulnerabilities
mcp-evals
v2.0.1 · 1 year ago
GitHub Action for evaluating MCP server tool calls using LLM-based scoring
No known vulnerabilities
@rulvar/evals
v1.182.0 · 7 hours ago
Rulvar evals: eval cases, golden outputs, rubric and judge graders, matrix sweeps, canary fingerprint.
No known vulnerabilities
promptfoo
v0.122.0 · 9 hours ago
LLM eval & testing toolkit
No known vulnerabilities
@mastra/evals
v1.7.0 · 3 hours ago
No description provided.
No known vulnerabilities
@vitest-evals/report-ui
v0.15.0 · 19 days ago
Local React report UI for vitest-evals JSON artifacts.
No known vulnerabilities
@arizeai/phoenix-evals
v2.2.0 · 1 day ago
A library for running evaluations for AI use cases
No known vulnerabilities
lastlight-evals
v0.9.10 · 22 hours ago
Eval harness for Last Light workflows — drives the real production workflows against a mocked GitHub and grades deterministically (SWE-bench compatible).
No known vulnerabilities
@azure/ai-projects
v2.4.0 · 5 hours ago
Azure AI Projects client library.
No known vulnerabilities
vitest-evals
v0.15.0 · 19 days ago
Harness-backed AI testing on top of Vitest.
No known vulnerabilities
@agent-assistant/telemetry
v0.4.35 · 2 months ago
Usage, cost, and response telemetry primitives for Agent Assistant
No known vulnerabilities
niceeval
v0.11.3 · 10 days ago
Agent-native eval tool — eval agents, services, functions, and coding-agent fixtures
No known vulnerabilities
@vitest-evals/core
v0.15.0 · 19 days ago
Shared primitives, schemas, and report collection helpers for vitest-evals.
No known vulnerabilities
@mono-agent/agent-evals
v0.4.0 · 1 month ago
Local-first end-to-end eval runner for agent responders and harnesses.
No known vulnerabilities
evalite
v0.19.0 · 9 months ago
Test your LLM-powered apps with a TypeScript-native, Vitest-based eval runner. No API key required.
No known vulnerabilities
neutral-evals
v1.0.0 · 9 months ago
Neutral Agent Evals
No known vulnerabilities
arc-skill-eval
v0.26.1 · 9 days ago
Pi-native library and CLI that runs Anthropic-standard skill evals (evals/evals.json) with LLM-judged + script assertions.
No known vulnerabilities