97,770 packages matching “agent-eval”
agent-eval
v0.0.1 · 11 months ago
No description provided.
No known vulnerabilities
@vercel/agent-eval-playground
v0.1.3 · 5 months ago
Web-based playground for browsing agent-eval experiment results
No known vulnerabilities
@tangle-network/agent-eval
v0.143.0 · 3 hours ago
Evaluate and improve AI agents from runs, traces, judges, and feedback. Compare candidates, cluster failures, measure lift, and gate releases.
No known vulnerabilities
@reaatech/agent-eval-harness-types
v0.1.0 · 3 months ago
Shared domain types and Zod schemas for agent-eval-harness
No known vulnerabilities
@iris-eval/mcp-server
v0.4.4 · 1 month ago
The agent eval standard for MCP. Score every agent output for quality, safety, and cost.
No known vulnerabilities
fasteval
v0.1.0-alpha.0 · 1 month ago
Lightweight TypeScript agent eval tool — eval agents, services, functions, and coding-agent fixtures
No known vulnerabilities
@reaatech/agent-eval-harness-observability
v0.1.1 · 1 month ago
OpenTelemetry observability (tracing, metrics, logging, dashboards) for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-trajectory
v0.1.0 · 3 months ago
Trajectory loading, evaluation, and comparison for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-golden
v0.1.0 · 3 months ago
Golden trajectory management, comparison, and curation for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-cost
v0.1.0 · 3 months ago
Cost tracking, budget management, and reporting for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-latency
v0.1.0 · 3 months ago
Latency monitoring, SLA enforcement, and optimization analysis for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-mcp-server
v0.1.2 · 1 month ago
Three-layer MCP tool server (judge, suite, gate) for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-tool-use
v0.1.0 · 3 months ago
Tool-use validation (selection, schema compliance, result verification) for agent-eval-harness
No known vulnerabilities
@tangle-network/traces
v0.11.5 · 2 hours ago
Point it at your coding-agent session traces (Claude Code, Codex, OpenCode, Gemini, Pi, …) and get failure-mode + efficiency findings. CLI + SDK over the @tangle-network/agent-eval analyst suite — observe live sessions, run your own analysts, redact, and
No known vulnerabilities
@reaatech/agent-eval-harness-judge
v0.2.0 · 1 month ago
Provider-agnostic LLM-as-judge with calibration and consensus for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-suite
v0.1.2 · 1 month ago
Orchestrated evaluation suite runner with results aggregation for agent-eval-harness
No known vulnerabilities
@vercel/agent-eval
v1.4.0 · 11 days ago
Framework for testing AI coding agents in isolated sandboxes
No known vulnerabilities
@reaatech/agent-eval-harness-gate
v0.1.2 · 1 month ago
CI regression gates, threshold checks, and JUnit/GitHub integration for agent-eval-harness
No known vulnerabilities
@reaatech/agent-eval-harness-cli
v0.1.2 · 1 month ago
CLI interface for agent-eval-harness with eval, judge, compare, gate, golden, report, and serve commands
No known vulnerabilities
code-agent-eval
v0.0.1-alpha.11 · 25 days ago
TypeScript library for evaluating prompts against coding agents (Claude Code, Cursor, etc.) with multi-iteration testing and scoring
No known vulnerabilities