477,085 packages matching “judge-d”
autoevals
v0.3.0 · 2 months ago
Universal library for evaluating AI models
No known vulnerabilities
@pwtap/plugin-ai-judge
v0.1.2 · 27 days ago
AI/LLM judge matchers for Playwright — toPassRubric/toScoreAtLeast/toMatchImage over Ollama, OpenAI-compatible endpoints (OpenAI/OpenRouter/NVIDIA/Groq/…), and native Claude
No known vulnerabilities
@credence/judge
v0.5.0 · 10 days ago
Semantic conflict adjudication for Credence: declarative conflict families, cached verdicts, optional LLM.
No known vulnerabilities
agentic-test-runner
v0.5.0 · 28 days ago
LLM-judged CLI test runner — define test cases in YAML, execute commands, let an LLM judge PASS/FAIL. Outputs JSONL trace.
No known vulnerabilities
@qvac/bci-whispercpp
v0.7.2 · 2 days ago
Brain-Computer Interface (BCI) neural signal transcription addon for qvac, powered by whisper.cpp
No known vulnerabilities
@elsium-ai/testing
v0.21.0 · 16 days ago
Testing utilities, mock providers, fixtures, and eval framework for ElsiumAI
No known vulnerabilities
llm-governance-gateway
v0.8.0 · 25 days ago
Governed structured-output LLM pipeline: rate limit → spend caps (per-user + global circuit breaker) → cache → provider-chain failover → Zod validation → usage logging → LLM-judge, with a deterministic mock provider for CI.
No known vulnerabilities
@skill-harness/core
v0.9.0 · 1 day ago
skill-harness engine — spec, discover, run, LLM-judge grade, score, results (internal API)
No known vulnerabilities
@mutagent/evaluator
v0.2.0-alpha.6 · 14 days ago
mutagent-evaluator: a generic, subject-agnostic AI-agent auditor — reviewer, never executor. Audits any skill/agent against a generated subject profile and emits a 4-tab master-audit report.
No known vulnerabilities
@blazediff/agent
v0.11.0 · 2 days ago
Agentic visual regression for BlazeDiff. Auto-discovers routes, captures deterministic screenshots, runs CI checks.
No known vulnerabilities
@mastra/evals
v1.9.0 · 17 hours ago
No description provided.
No known vulnerabilities
deepthink-js
v2.0.1 · 7 days ago
SOTA NPM module for agentic processes using local or cloud LLMs.
No known vulnerabilities
devcouncil
v0.4.2 · 1 month ago
Gated orchestrator for AI-assisted software development
No known vulnerabilities
@forwardimpact/libharness
v3.0.0 · 1 month ago
Autonomous agent team harness — coordinate a lead and participant agents in one async session, with eval, benchmark, and trace tooling to prove the changes worked.
No known vulnerabilities
ads-mcp
v0.3.0 · 17 days ago
Render web and SwiftUI evidence, run explicit visual judgment, and trace ADS UI decisions.
No known vulnerabilities
@eslint/js
v10.0.1 · 6 months ago
ESLint JavaScript language implementation
No known vulnerabilities
@themoltnet/agent-runtime
v0.43.1 · 1 day ago
MoltNet agent runtime — coding-agent-agnostic Task execution loop
No known vulnerabilities
propline-cli
v0.24.1 · 55 minutes ago
Command-line interface for the PropLine player props betting odds API. Wraps the propline SDK with pretty-printed tables and JSON output.
No known vulnerabilities
@skill-harness/adapters
v0.9.0 · 1 day ago
skill-harness harness adapters — pi runner + claude-code judge routing (internal API)
No known vulnerabilities
@autodevops/verifier
v0.1.2 · 27 days ago
Standalone CLI tool for verification tasks
No known vulnerabilities