11,559 packages matching “eval-harness”
eval-harness
v0.2.18 · 1 month ago
Evaluation harness for configurable runner work packages
No known vulnerabilities
@plaited/agent-eval-harness
v1.0.1 · 1 month ago
General-purpose eval harness for running trials against CLI agents
No known vulnerabilities
@accelerate-data/promptfoo-eval-harness
v1.7.2 · 8 days ago
Promptfoo + OpenCode eval harness for agent behavior. Owns model/tier policy, provider wiring, package discovery, state export, and artifact guards. Consumers own eval YAML, prompts, fixtures, and assertions.
No known vulnerabilities
lastlight-evals
v0.9.16 · 1 hour ago
Eval harness for Last Light workflows — drives the real production workflows against a mocked GitHub and grades deterministically (SWE-bench compatible).
No known vulnerabilities
@tonyclaw/eval-harness-win32-x64
v0.2.19 · 1 month ago
Windows x64 native binary for @tonyclaw/eval-harness
No known vulnerabilities
@agent-relay/evals
v11.4.1 · 3 days ago
Agent Relay eval harness — scenario runner, broker harness, and scoring utilities for testing relay-connected agents
No known vulnerabilities
@reaatech/agent-eval-harness-mcp-server
v0.1.2 · 1 month ago
Three-layer MCP tool server (judge, suite, gate) for agent-eval-harness
No known vulnerabilities
@elisym/eval
v0.2.0 · 1 month ago
Eval harness for payment-enabled AI agents - deterministic assertions over traces and ledgers
No known vulnerabilities
langdrift
v0.4.0 · 1 month ago
Locale-aware eval harness for AI agent behavior
No known vulnerabilities
@reaatech/agent-eval-harness-judge
v0.2.0 · 1 month ago
Provider-agnostic LLM-as-judge with calibration and consensus for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-suite
v0.1.2 · 1 month ago
Orchestrated evaluation suite runner with results aggregation for agent-eval-harness
No known vulnerabilities
@looprun-ai/eval
v0.16.0 · 17 minutes ago
looprun eval harness: run a generated subject's cases against a model target (governed or ungoverned variant), dump per-case traces for the LLM judge, fold verdicts, and certify against the bar.
No known vulnerabilities
crewhaus
v0.5.1 · 1 day ago
CrewHaus — the meta-harness compiler for AI agents. Compile one crewhaus.yaml spec into a CLI agent, channel bot, RAG pipeline, multi-agent crew, eval harness, voice or browser agent, and more.
No known vulnerabilities
@odla-ai/ai
v0.14.0 · 7 hours ago
One general-purpose interface for AI inference across Anthropic (Claude), OpenAI, and Google (Gemini) — text, image, and audio — plus a tool-use agent engine and eval harness for Node 20+ and Cloudflare Workers.
No known vulnerabilities
@reaatech/agent-eval-harness-gate
v0.1.2 · 1 month ago
CI regression gates, threshold checks, and JUnit/GitHub integration for agent-eval-harness
No known vulnerabilities
undetermini
v0.2.4 · 28 days ago
Eval harness for non-deterministic (LLM) code — subjects, variants (provider × model × reasoning), weighted-assertion scoring, trial cache, SQLite runs, CLI + TUI.
No known vulnerabilities
@bondarewicz/dreamteam
v1.4.1 · 9 days ago
Dream Team — a Claude Code agent roster with a cross-provider, schema-enforced eval harness (Claude, Ollama, Codex). Bun-only.
No known vulnerabilities
@bhaskarauthor/eval-harness
v0.1.0-alpha.1 · 26 days ago
Lightweight eval harness for prompt/agent regression tests in CI for product teams
No known vulnerabilities
@reaatech/agent-eval-harness-observability
v0.1.1 · 1 month ago
OpenTelemetry observability (tracing, metrics, logging, dashboards) for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-types
v0.1.0 · 3 months ago
Shared domain types and Zod schemas for agent-eval-harness
No known vulnerabilities